A method for predicting the proliferation or apoptosis of ovarian cancer cells in an organism intervened by bevacizumab

The machine learning model is used to construct significant feature extraction and prediction methods after intervention of ovarian cancer cells, which solves the problem of prediction in the existing technology, and achieves efficient and accurate prediction of ovarian cancer cell proliferation or apoptosis, improving research efficiency and accuracy.

CN116313048BActive Publication Date: 2025-07-04ANHUI MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310149625.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-16
Publication Date
2025-07-04
Estimated Expiration
2043-02-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the proliferation or apoptosis of ovarian cancer cells after bevacizumab intervention, and traditional dimensionality reduction algorithms cannot extract significant features, resulting in the time-consuming and low accuracy of the analysis by researchers.

Method used

Using machine learning model, the prediction methods for ovarian cancer cell proliferation or apoptosis are constructed through the Relief feature selection algorithm and the extreme random tree regression model, including data acquisition and preprocessing, significant feature extraction and prediction model establishment, and the data characteristics after bevacizumab intervention are used for accurate prediction.

Benefits of technology

It greatly shortens the analysis and prediction time, improves the accuracy and consistency of prediction, avoids the problems of knowledge limitations and artificial fatigue, provides accurate interpretation of significant characteristics, and improves the predictive ability of ovarian cancer cells after bevacizumab intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116313048B_ABST
    Figure CN116313048B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for predicting the proliferation or apoptosis of ovarian cancer cells in an organism intervened by bevacizumab, which solves the defect of being difficult to predict the proliferation / apoptosis of ovarian cancer cell data after intervention compared with the prior art. The present invention includes the following steps: acquisition and preprocessing of basic data of ovarian cancer cells in an organism; construction of a proliferation / apoptosis prediction model; prediction of the proliferation / apoptosis of ovarian cancer cells in an organism intervened by bevacizumab. Based on a machine learning model, the present invention constructs a significant feature extraction and prediction model for ovarian cancer cells in an organism after being intervened by bevacizumab, which can not only obtain the significant features affecting the proliferation / apoptosis of ovarian cancer cells in an organism after being intervened by bevacizumab, but also greatly shorten the time required for analysis and prediction, avoiding the problems of knowledge limitation and manual fatigue. At the same time, the predicted value obtained based on the present invention has good consistency with the actual value and a high accuracy rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biological big data analysis, and specifically relates to a method for predicting the proliferation or apoptosis of ovarian cancer cells in an organism intervened by bevacizumab. Background Art

[0002] Ovarian cancer is one of the common gynecological malignancies, and its fatality rate ranks second among gynecological cancers. Cytoreductive surgery and platinum-based chemotherapy remain the first-line treatment methods for newly diagnosed advanced ovarian cancer patients. Most patients have no signs of disease after first-line treatment, but approximately 70% of patients relapse within the next 3 years. Recurrent ovarian cancer is currently incurable, and continuous treatment is given at each recurrence, with the progression-free survival (hereinafter referred to as "PFS") gradually shortening, subsequent treatment of patients becoming increasingly difficult, and the survival prognosis becoming increasingly poor. There is an urgent need in clinical work to extend the PFS of ovarian cancer patients to improve their survival prognosis. With the addition of anti-angiogenic therapies such as bevacizumab as a first-line treatment regimen, the PFS of ovarian cancer patients has been extended to a certain extent, but it is only effective for some patients. In addition, bevacizumab treatment is costly and has potential drug toxicity effects, often bringing a heavy economic, physical, and psychological burden to patients. Considering the cost, potential toxicity, and limited clinical benefits of bevacizumab treatment, efforts are being made in clinical practice to identify the advantageous population for bevacizumab application and establish an accurate efficacy prediction model for bevacizumab.

[0003] For the establishment of a therapeutic efficacy prediction model, relevant clinical studies have a long cycle and often inevitably face risks such as loss to follow-up, dropout from the study group, and bias. Establishing animal models also has problems such as non-human origin, long cycle, and high cost. Currently, in view of the drawbacks of clinical research and animal experiments, an ovarian cancer organoid chip model is established using an organoid chip cell experiment platform, and on this basis, the establishment of relevant prediction models emerges as the times require. At the cellular level, the efficacy evaluation of ovarian cancer treated with bevacizumab can be disassembled into the cell proliferation or apoptosis of ovarian cancer cells after being intervened with bevacizumab. The assessment of cell proliferation or apoptosis depends on the researchers manually performing tumor attachment and volume measurement, manually counting the number of cells suspended in the culture medium and quantitatively evaluating the tumor volume, which is time-consuming and laborious, and the accuracy of judging the degree of cell proliferation / apoptosis highly depends on the knowledge, experience, and preference of the researchers. We can simulate the ovarian cancer tumor microenvironment in the body by adjusting the temperature, humidity environment, and the physicochemical environment of the culture medium during the culture process of the ovarian cancer organoid chip, and obtain data such as the changes in various indicators in the tumor microenvironment before and after being intervened with bevacizumab. These data are huge and complex, and it is difficult to conduct a full-feature analysis through manual statistics. Often, only one of these indicators can be selected to statistically analyze the changes and the correlation between tumor proliferation / apoptosis. These characteristic data are not equally important, and the data of many of these characteristics have little reference significance for cell proliferation / apoptosis. This leads to the situation that when researchers analyze, they often spend a lot of time on interfering features and easily ignore some meaningful characteristic data.

[0004] Then, how to use ovarian cancer cells as data carriers, based on the proliferation / apoptosis data of ovarian cancer cells in the body under the intervention of bevacizumab, and use big data analysis technology to extract the significant features that affect the proliferation / apoptosis of ovarian cancer cells in the body under the intervention of bevacizumab from the numerous characteristic data in the ovarian cancer tumor microenvironment.

[0005] There is currently a lack of effective methods to solve the problem of affecting the proliferation / apoptosis of ovarian cancer cells in the body after bevacizumab intervention. At present, the mainstream methods of reducing feature dimensionality, such as principal component analysis (PCA), linear discriminant analysis (LDA) and other traditional dimensionality reduction algorithms, while reducing the feature dimension, it is difficult to find the corresponding relationship between the feature data obtained by dimensionality reduction and the original features, and it is impossible to extract significant features. Establishing an efficient prediction model for the proliferation / apoptosis of ovarian cancer cells in the body after bevacizumab intervention will help reduce the intensive work of researchers in finding feature data, help improve the researchers' accurate interpretation of the indicator data after bevacizumab intervention of ovarian cancer cells in the body, and help solve the problem of features that are easily overlooked by researchers and have important reference significance for the results of bevacizumab intervention. According to the extracted significant features, the survival prediction model is further optimized to improve the prediction performance.

[0006] Therefore, how to use big data analysis technology to preview, analyze, and predict the proliferation / apoptosis of ovarian cancer cells in the body under the intervention of bevacizumab based on the body's ovarian cancer cell data has become a technical problem that needs to be solved urgently. Summary of the invention

[0007] The purpose of the present invention is to solve the defect in the prior art that it is difficult to predict the proliferation / apoptosis of ovarian cancer cells after data intervention, and to provide a method for predicting the proliferation or apoptosis of ovarian cancer cells in an organism intervened by bevacizumab to solve the above problem.

[0008] In order to achieve the above object, the technical solution of the present invention is as follows:

[0009] A method for predicting proliferation or apoptosis of ovarian cancer cells in an organism intervened by bevacizumab comprises the following steps:

[0010] 11) Acquisition and preprocessing of basic data of ovarian cancer cells in the body: Acquiring basic data of ovarian cancer cells in the body, including basic data of bevacizumab intervention, extracting and marking basic data of ovarian cancer cells before and after bevacizumab intervention and whether ovarian cancer cells proliferate / apoptosis or not as labels, performing data encoding, data cleaning, and standardized preprocessing operations on the basic data to obtain processed data, and dividing the data set for subsequent proliferation / apoptosis prediction model of ovarian cancer cells after bevacizumab intervention, as well as proliferation / apoptosis-related significant feature extraction and prediction model training and testing;

[0011] 12) Construction of proliferation / apoptosis prediction model: Construct a full-feature classification model, obtain the alternative significant feature set of ovarian cancer cell proliferation / apoptosis according to the Relief feature selection algorithm, extract the significant feature set of ovarian cancer proliferation / apoptosis based on the alternative significant feature set, and then establish a proliferation / apoptosis prediction model based on the extracted significant features;

[0012] 13) Prediction of the proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab: Construct a full-feature regression model for the proliferation / apoptosis duration of ovarian cancer cells after intervention. The model generalization ability index is measured by the coefficient of determination R 2 of the test set. Obtain the alternative significant feature set of the proliferation / apoptosis duration of ovarian cancer cells after intervention based on the full-feature regression prediction model, then extract the significant feature set of the proliferation / apoptosis duration of ovarian cancer cells after intervention based on the alternative significant feature set, and finally establish a prediction model for the proliferation / apoptosis duration data of ovarian cancer cells after intervention to realize the prediction of the proliferation or apoptosis of ovarian cancer cells in the body intervened by bevacizumab.

[0013] The acquisition and preprocessing of the basic data of the ovarian cancer cells in the body include the following steps:

[0014] 21) Acquisition of ovarian cancer cell data in the body: Collect and record the information data of ovarian cancer cells in different bodies, including three types of information data; The first type of data is the detection index data of total protein, albumin, globulin, albumin / globulin ratio, alkaline phosphatase, lactate dehydrogenase, creatinine, urea, potassium, sodium, chloride, bicarbonate, calcium, phosphorus, glucose, total cholesterol, triglyceride, high-density lipoprotein cholesterol, non-high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, very-low-density lipoprotein cholesterol, apolipoprotein A1, apolipoprotein B, lipoprotein a, and free fatty acids in the ovarian cancer tumor microenvironment when starting to intervene with bevacizumab; The second type of data is the detection index data of total protein, albumin, globulin, albumin / globulin ratio, alkaline phosphatase, lactate dehydrogenase, creatinine, urea, potassium, sodium, chloride, bicarbonate, calcium, phosphorus, glucose, total cholesterol, triglyceride, high-density lipoprotein cholesterol, non-high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, very-low-density lipoprotein cholesterol, apolipoprotein A1, apolipoprotein B, lipoprotein a, and free fatty acids in the ovarian cancer tumor microenvironment during tumor proliferation / apoptosis; The third type of data is the categorical variable data on whether the ovarian cancer cell population continues to proliferate or apoptose;

[0015] 22) Preprocessing of the ovarian cancer cell data in the body intervened by bevacizumab. The data of the ovarian cancer cells in the body intervened by bevacizumab are encoded, cleaned, normalized, and the dataset is divided. The specific steps are as follows:

[0016] 221) Encoding of data on ovarian cancer cells in the body intervened with bevacizumab: Non-numerical data in the data is encoded into numerical data for algorithm processing. The same type of data is encoded with one code to achieve data encoding. The operation object includes the classification variable of proliferation / apoptosis;

[0017] 222) Cleaning of data on ovarian cancer cells in the body intervened with bevacizumab:

[0018] When extracting significant features and making predictions based on the characteristic data of ovarian cancer cells before intervention, the characteristic variable data after intervention is deleted; the characteristic data with missing values exceeding half is deleted; abnormal characteristic data is found and deleted through visualization methods such as combined scatter plots, box plots, and variable distribution plots; for discrete data, median filling or mode filling operations are used, and for continuous data, mean filling is used;

[0019] 223) Normalization processing of data on ovarian cancer cells in the body intervened with bevacizumab: The data processed above is normalized to achieve data normalization processing;

[0020] 224) Dataset division of the prediction model: The processed data is randomly divided into a training set and a test set according to a ratio of 8:2. The ratio of proliferation / apoptosis of ovarian cancer cells in the test set and the training set is made balanced and reasonable for the subsequent training and testing of the proliferation / apoptosis prediction model of ovarian cancer cells in the body after intervention with bevacizumab;

[0021] 225) Extraction of significant features and dataset division of the prediction model for the duration of proliferation / apoptosis after bevacizumab intervention: Ensure that the distribution of the duration of proliferation / apoptosis of ovarian cancer cells in the body after intervention with bevacizumab in the test set and the training set is consistent for the subsequent training and testing of the model for extracting significant features and predicting the duration of proliferation / apoptosis of ovarian cancer cells in the body after intervention with bevacizumab. The specific steps are as follows:

[0022] 2251) The dataset of the model for extracting significant features and predicting the duration of proliferation / apoptosis after bevacizumab intervention is randomly divided into a significant feature training set and a significant feature test set according to a ratio of 8:2;

[0023] 2252) Plot the data distribution curves of the duration of cell proliferation / apoptosis after ovarian cancer intervention in the significant feature training set and the significant feature test set, and initially establish a full-feature regression model for the duration of cell proliferation / apoptosis after ovarian cancer drug use. Select polynomial regression and decision tree regression models without adjusting model hyperparameters. The model generalization ability index is measured by the coefficient of determination R 2 of the test set. Observe the model generalization ability and the duration data distribution curves on the training set and the test set. The calculation formula of the coefficient of determination R 2 is as follows,

[0024]

[0025] In the above formula, y i represents the true data of the proliferation / apoptosis duration of ovarian cancer cells in body i, represents the estimated value of the proliferation / apoptosis duration obtained by the model estimating the ovarian cancer cells in body i, represents the average value of the true proliferation / apoptosis duration data of ovarian cancer cells in the body, n is the total number of samples, and SSR, SSE, and SST represent the regression sum of squares, the residual sum of squares, and the total sum of squares of deviations, respectively;

[0026] 2253) If the generalization ability of the model is greater than 0.70 and the data distribution curves on the training set and the test set are approximately similar, then select this data set partitioning scheme; otherwise, go to step 2251).

[0027] The construction of the proliferation / apoptosis prediction model includes the following steps:

[0028] 31) Based on the basic data of ovarian cancer cells in the body, establish a full-feature classification model. The index value of the generalization ability of the full-feature classification model is measured by the classification accuracy rate of the test set. The definition of the classification accuracy rate is as follows:

[0029]

[0030] Among them, TP represents the number of samples correctly classified as positive examples, that is, the number of data correctly predicted as the proliferation of ovarian cancer cells, TN represents the number of samples correctly classified as negative examples, that is, the number of data correctly predicted as the apoptosis of ovarian cancer cells, P represents the number of all positive examples, that is, the number of all data of the proliferation of ovarian cancer cells, and N represents the number of all negative examples, that is, the number of all data of the apoptosis of ovarian cancer cells;

[0031] Select a support vector machine classifier model, a naive Bayes classifier model, or a decision tree classifier model, and perform multiple trainings and parameter tuning on different models; for the hyperparameters of each classification model, use the grid search method, the random search method, or the Bayesian optimization parameter tuning method to adjust the hyperparameters and strengthen the generalization performance of the model; after multiple trainings of multiple models, compare them and select the classification model with the best generalization performance as the full-feature classification model for the proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab in this step. The obtained optimal classification model is the extremely randomized tree classification model whose hyperparameters are determined by the random search method, and the optimal generalization performance is denoted as α;

[0032] 32) Obtain an alternative significant feature set for the proliferation / apoptosis of ovarian cancer cells in the body intervened with bevacizumab according to the Relief feature selection algorithm: Specify a threshold τ, and select features with feature weight values greater than τ, or specify the number k of features to be selected, and select k features into the alternative significant feature set;

[0033] 33) Extract a significant feature set for the proliferation / apoptosis of ovarian cancer cells in the body intervened with bevacizumab based on the alternative significant feature set;

[0034] Based on the alternative significant feature set with k features obtained by the Relief feature selection algorithm in 32), preliminarily sort the features of the alternative significant feature set in descending order according to the weight magnitudes of the k features;

[0035] Construct an alternative significant feature evaluation set. In the i-th selection, select the first i features from the alternative significant features, where i = 1, 2, …, k, and k is the number of features in the alternative feature set. Based on the full-feature classification model obtained in step 31), adjust the input features to the currently selected i features, perform multiple trainings and hyperparameter optimizations, etc., to obtain the generalization ability of the extremely randomized tree classification model, and denote the optimal generalization ability as β;

[0036] Taking α obtained in step 31) as a reference, if β ≥ 0.75 or β ≥ 0.9 * α, it indicates that significant features of the proliferation / apoptosis of ovarian cancer cells in the body intervened with bevacizumab have been successfully extracted. Otherwise, increase the value of k, go to step 32), and continue the iteration until significant features of the proliferation / apoptosis of ovarian cancer cells in the body intervened with bevacizumab are successfully extracted;

[0037] The above iteration termination condition of β ≥ 0.9 * α is regarded as the significance rate of introducing the full-feature subset. The significance rate of the feature subset here is defined as where ξ = 1.0×10 -6 , this parameter is used to prevent the error of the denominator being zero, that is, the ratio of the generalization ability value β of the best classification model based on the subset to the generalization ability index value α of the best classification model based on the full-feature data; the larger this index value is, the better. Close to 0 indicates that the feature subset is not a significant feature set, close to 1 indicates that the feature subset can be regarded as a significant feature set, and greater than 1 indicates that the feature subset can improve the prediction ability for the classification problem of the proliferation / apoptosis of ovarian cancer cells in the body intervened with bevacizumab; here, the threshold of this ratio is taken as 0.9;

[0038] Meanwhile, the iteration termination condition of β ≥ 0.75 indicates that when the generalization ability of the selected feature subset for the classification task of the survival prognosis of ovarian cancer with the proliferation / apoptosis of ovarian cancer cells in the body intervened with bevacizumab is greater than 0.75, it is regarded as extracting significant features. Too low indicates that it is necessary to return to step 32) and iterate again to extract significant features of ovarian cancer survival prognosis;

[0039] 34) Based on the prediction model of the proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab trained using all features in step 31), the generalization ability of the model is the classification accuracy rate on the test set. Using the significant features obtained in step 33), further training and hyperparameter optimization of the model are carried out to obtain the proliferation / apoptosis prediction model established based on the significant features.

[0040] The prediction of the proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab includes the following steps:

[0041] 41) Construct a full-feature regression model for the duration of proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab:

[0042] 411) Based on the selected dataset partitioning scheme, establish a full-feature regression model for the duration of proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab based on all the feature data of ovarian cancer cells. The generalization ability index of the model is measured by the coefficient of determination R 2 of the test set;

[0043] 412) Model training and hyperparameter optimization. For hyperparameter optimization, the grid search method, random search method or Bayesian optimization method is selected;

[0044] 413) Select the extremely randomized tree regression prediction model with the optimal generalization ability, and the generalization ability is denoted as γ;

[0045] 42) Based on the full-feature regression prediction model, obtain an alternative significant feature set for the duration of proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab:

[0046] The number of features k in the alternative significant feature set is default set to 30;

[0047] 421) According to the obtained extremely randomized tree regression prediction model, obtain the weights of all the features of ovarian cancer cells. The greater the weight, the greater the importance of the corresponding feature;

[0048] 422) According to the weights of the features, select the top k important features as the alternative significant feature set for the duration of proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab;

[0049] 423) Arrange the features in the alternative significant feature set in descending order according to the weights of the features, and the result is used as the final result of the alternative significant feature set;

[0050] 43) Based on the alternative significant feature set, extract the significant feature set for the duration of proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab:

[0051] 431) Based on the alternative significant feature set of the proliferation / apoptosis duration of ovarian cancer cells in the body affected by the intervention of bevacizumab, select the first i features, where i = 1, 2, 3, …, k;

[0052] 432) Based on these i features, establish an extremely randomized tree regression prediction model for the proliferation / apoptosis duration of ovarian cancer cells in the body affected by the intervention of bevacizumab, perform model training and hyperparameter optimization, and use the coefficient of determination R of the test set as the model generalization ability index 2 to measure, and select the optimal value of the generalization ability as the prediction ability index value η of these i features for the proliferation / apoptosis duration after the intervention of ovarian cancer cells i ;

[0053] 433) For η = {η1, η2, …, η k )), select the maximum value among them as η max , and obtain the corresponding number of features numbers;

[0054] 434) When η max ≥ 0.9*γ, or η max ≥ 0.75, then output the first numbers features in the alternative significant feature set as the significant features; otherwise, increase the value of k and go back to step 42);

[0055] 44) Based on the extracted significant features, establish a prediction model for the proliferation duration data of ovarian cancer cells in the body affected by the intervention of bevacizumab:

[0056] Based on the significant features of the proliferation / apoptosis duration of ovarian cancer cells in the body affected by the intervention of bevacizumab extracted in step 43), establish a prediction model for the proliferation / apoptosis duration of ovarian cancer cells in the body affected by the intervention of bevacizumab. Use the coefficient of determination R on the test set as the model generalization ability index 2 to measure. The model selects the extremely randomized tree regression model. Based on the model training results and hyperparameter results of the full-feature regression prediction model extracted using all features in step 41), further perform training and hyperparameter optimization to obtain a prediction model for the proliferation / apoptosis duration data of ovarian cancer cells in the body affected by the intervention of bevacizumab based on the extracted significant features, and output the prediction results of the proliferation or apoptosis duration of ovarian cancer cells in the body intervened by bevacizumab.

[0057] It also includes the analysis step of risk factors, that is, according to the prediction results of the proliferation or apoptosis of ovarian cancer cells in the body intervened by bevacizumab, obtain the risk factors affecting the proliferation of ovarian cancer cells in the body after the intervention of bevacizumab. The specific steps are as follows:

[0058] Based on the significant features of the proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab and the significant features affecting the duration of proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab, taking the intersection of these two sets of significant features can obtain the common significant influencing features; taking the union of these two sets of significant features gives all the significant features; comprehensively analyzing the features that have significant effects on both the progress direction of proliferation / apoptosis and the duration of proliferation / apoptosis in the intersection and the significant features that affect at least one of the progress direction and duration in the union, the risk factors affecting the proliferation of ovarian cancer cells in the body after bevacizumab intervention are obtained.

[0059] It also includes a data analysis method for the proliferation / apoptosis of ovarian cancer cells in the body intervened by drugs, and the steps are as follows:

[0060] Collect the index data of ovarian cancer cells in the body and extract significant features, and construct a prediction model for the proliferation / apoptosis characteristics data of ovarian cancer cells in the body after drug intervention based on the significant features; perform corresponding intersection and union set operations on the significant features of each important data to obtain the set of significant features of various data that affect the proliferation / apoptosis of ovarian cancer cells in the body after drug intervention; the obtained set of significant features and the prediction model obtained in this extraction process provide a drug resistance trend prediction for the drug intervention of ovarian cancer cells in the body.

[0061] Beneficial effects

[0062] A prediction method for the proliferation or apoptosis of ovarian cancer cells in the body intervened by bevacizumab according to the present invention, compared with the prior art, constructs a significant feature extraction and prediction model for ovarian cancer cells in the body after bevacizumab intervention based on a machine learning model. It can not only obtain the significant features that affect the proliferation / apoptosis of ovarian cancer cells in the body after bevacizumab intervention, but also greatly shorten the time required for analysis and prediction, avoiding the problems of knowledge limitation and manual fatigue. At the same time, the predicted value based on the present invention has good consistency with the actual value and a high accuracy rate.

[0063] In practical applications, on the one hand, the present invention can provide researchers with the relevant significant features of the proliferation / apoptosis of ovarian cancer cells in the body after bevacizumab intervention, facilitating researchers to accurately observe the significant feature data among dozens or even hundreds of feature data; on the other hand, the present invention can improve the prediction ability of the proliferation / apoptosis of ovarian cancer cells in the body after bevacizumab intervention, and the results have key supporting significance for studying the changes in lipid metabolism during the bevacizumab intervention process. Brief description of the drawings

[0064] Figure 1 It is the method sequence diagram of the present invention;

[0065] Figure 2It is the flowchart of the method implementation involved in the present invention;

[0066] Figure 3a It is the ROC curve graph of the extremely randomized tree model adopted by the present invention;

[0067] Figure 3b It is the calibration curve graph of the extremely randomized tree model adopted by the present invention;

[0068] Figure 4a It is the comparison graph of the ROC curves between the method of the present invention and the prior art method;

[0069] Figure 4b It is the comparison graph of the calibration curves between the method of the present invention and the prior art method;

[0070] Figure 5 It is the nomogram reflecting the relationship between significant features and proliferation time of the present invention. Detailed implementation manners

[0071] To have a further understanding and recognition of the structural features and achieved effects of the present invention, the following is a detailed description with preferred embodiments and accompanying drawings:

[0072] First, the definition of significant features is given: Let the initial full feature set {a1, a2, L, a n}, denoted as U, the predicted target feature is y, the significant feature set D is a feature subset selected from the full feature set U that contains all important information, and the significant features are all the features in the significant feature set. If no domain knowledge is used as a prior assumption and the significant feature set D is solved by traversal, then all possible subsets need to be traversed, and the number of traversals is That is, 2 n -1. When the number of features n is large, the problem of combinatorial explosion will be encountered, which is not feasible in practical applications.

[0073] Suppose several features are selected from the full feature set U to obtain the alternative significant feature set T, and the optimal value of the generalization ability of the model about the target feature y established based on the full feature set U is λ u , and the optimal value of the generalization ability of the model about the target feature y established based on the alternative significant feature set T is λ t . The model generalization ability index value mentioned here is processed as follows, that is, the larger the model generalization ability index value, the greater the generalization ability of the model. Further, the model generalization ability index value is uniformly processed into a number greater than or equal to 0. Define the significance rate θ, The introduction of ξ prevents the phenomenon that the denominator takes a value of 0. Introduce the significance rate threshold θ threshold , when θ ≥ θ threshold, consider the alternative set of significant features as the desired set of significant features. Based on the above description, solving for significant features is an iterative process, where the optimal value of the model generalization ability can be taken as the value of the generalization ability with the best performance obtained after several model selections and model trainings, and the significance rate threshold θ threshold It is recommended to take the value of 0.9, and the significance rate threshold θ threshold Taking different values can obtain significant feature sets with different significance levels. Based on the above significance evaluation index and related definitions of significant features, it is possible to achieve rapid extraction of significant features while ensuring the significance rate of significant features, and efficiently solve the problem of extracting significant features.

[0074] Such as Figure 1 and Figure 2 As shown, a method for predicting the proliferation or apoptosis of ovarian cancer cells in an organism intervened by bevacizumab according to the present invention includes the following steps:

[0075] The first step, acquisition and preprocessing of basic data of ovarian cancer cells in the organism:

[0076] (1) Acquisition of data of ovarian cancer cells in the organism, the specific steps are as follows:

[0077] Collect and record information data of ovarian cancer cells from different organisms, mainly including three categories of information data; the first category of data, detection index data of total protein, albumin, globulin, albumin / globulin ratio, alkaline phosphatase, lactate dehydrogenase, creatinine, urea, potassium, sodium, chloride, bicarbonate, calcium, phosphorus, glucose, total cholesterol, triglyceride, high-density lipoprotein cholesterol, non-high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, very-low-density lipoprotein cholesterol, apolipoprotein A1, apolipoprotein B, lipoprotein a, free fatty acids in the ovarian cancer tumor microenvironment when starting to intervene with bevacizumab; the second category of data, detection index data of total protein, albumin, globulin, albumin / globulin ratio, alkaline phosphatase, lactate dehydrogenase, creatinine, urea, potassium, sodium, chloride, bicarbonate, calcium, phosphorus, glucose, total cholesterol, triglyceride, high-density lipoprotein cholesterol, non-high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, very-low-density lipoprotein cholesterol, apolipoprotein A1, apolipoprotein B, lipoprotein a, free fatty acids in the ovarian cancer tumor microenvironment during tumor proliferation or apoptosis; the third category of data, categorical variable data on whether the ovarian cancer cell population continues to proliferate or apoptose;

[0078] (2) Preprocessing of data of ovarian cancer cells in the organism intervened by bevacizumab, mainly including data encoding, data cleaning, data normalization processing and dataset partitioning of data of ovarian cancer cells in the organism intervened by bevacizumab, the specific steps are as follows:

[0079] (2-1) Encoding of ovarian cancer cell data in the body intervened with bevacizumab: Non-numerical data in the data is encoded into numerical data that can be processed by the algorithm. The same type of data is encoded into one code to achieve data encoding. The main operation objects include classification variables of proliferation or apoptosis;

[0080] (2-2) Cleaning of ovarian cancer cell data in the body intervened with bevacizumab: Delete irrelevant attribute columns. For example, when extracting significant features and making predictions based on the ovarian cancer cell characteristic data before intervention, delete the characteristic variable data after intervention; Delete meaningless attribute columns with a large number of missing values, specifically delete characteristic data with more than half of the missing values; Delete abnormal data in the attributes. Through visualization methods such as combined scatter plots, box plots, and variable distribution plots, find and delete abnormal characteristic data; For characteristic variable data with fewer missing values, perform filling operations. The specific missing value filling operations are as follows. For discrete data, mainly use median filling or mode filling operations. For continuous data, mainly use mean filling;

[0081] (2-3) Normalization processing of ovarian cancer cell data in the body intervened with bevacizumab: To eliminate the influence of dimension and singular sample data, normalize the data processed above to achieve data normalization processing: First, normalize the characteristic data to the interval [0,1]. If the generalization ability of the subsequent model is not ideal enough, then standardize the data or normalize it to the interval [-1,1]. For the convenience of comparison, all characteristic data after missing value filling can also be directly used in the subsequent processing (that is, omit the data normalization operation), observe the generalization ability of the model, and then select the appropriate data preprocessing method;

[0082] (2-4) Dataset division of the prediction model: Randomly divide the processed data into a training set and a test set according to a ratio of 8:2, and ensure that the proportion of proliferation or apoptosis of ovarian cancer cells in the test set and the training set is balanced and reasonable, for the subsequent training and testing of the proliferation or apoptosis prediction model of ovarian cancer cells in the body intervened with bevacizumab;

[0083] (2-5) Dataset division of the significant feature extraction and prediction model for the proliferation or apoptosis duration after bevacizumab intervention: Ensure that the distribution of the proliferation or apoptosis time length of ovarian cancer cells in the body intervened with bevacizumab in the test set and the training set is the same, for the subsequent training and testing of the significant feature extraction and prediction model that affects the proliferation or apoptosis duration of ovarian cancer cells in the body intervened with bevacizumab. The specific steps are as follows:

[0084] (2-5-1) Randomly divide the dataset into a training set and a test set according to a ratio of 8:2;

[0085] (2-5-2) Plot the data distribution curves of the cell proliferation or apoptosis duration after ovarian cancer intervention on the training set and the test set, and preliminarily establish a full-feature regression model for the cell proliferation or apoptosis duration after ovarian cancer medication. Select different regression models, such as polynomial regression, decision tree regression and other regression models, without adjusting the model hyperparameters. The model generalization ability index is measured by the coefficient of determination R 2 of the test set, observe the model generalization ability and the duration data distribution curves on the training set and the test set. The calculation formula of the coefficient of determination R 2 is as follows.

[0086]

[0087] In the above formula, y i represents the true data of the proliferation or apoptosis duration of ovarian cancer cells in organism i, represents the estimated value of the proliferation or apoptosis duration obtained by the model for estimating ovarian cancer cells in organism i, represents the average value of the true proliferation or apoptosis duration data of ovarian cancer cells in the organism. n is the total number of samples. SSR, SSE, and SST represent the regression sum of squares, the residual sum of squares, and the total sum of squares of deviations respectively;

[0088] (2-5-3) If the model generalization ability is greater than 0.70 and the data distribution curves on the training set and the test set are approximately similar, then select this data set partitioning scheme. Otherwise, go to step (2-5-1).

[0089] Second step, construction of the proliferation or apoptosis prediction model:

[0090] (1) Based on the basic data of ovarian cancer cells in the organism, establish a full-feature classification model. The index value of the model generalization ability of the full-feature classification model is measured by the classification accuracy rate of the test set. The definition of the classification accuracy rate is as follows:

[0091]

[0092] Among them, TP represents the number of cases correctly classified as positive examples, that is, the number of data correctly predicted as ovarian cancer cell proliferation. TN represents the number of cases correctly classified as negative examples, that is, the number of data correctly predicted as ovarian cancer cell apoptosis. P represents the number of all positive examples, that is, the number of all ovarian cancer cell proliferation data. N represents the number of all negative examples, that is, the number of all ovarian cancer cell apoptosis data;

[0093] Use the random search method to determine the extremely randomized tree (ExtraTree) classification model with hyperparameters as the full-feature classification model for the proliferation or apoptosis of ovarian cancer cells in the organism intervened with bevacizumab in this step, and record the optimal generalization performance as α;

[0094] (2) Select the alternative significant feature set of the proliferation or apoptosis of ovarian cancer cells in the body intervened with bevacizumab according to the Relief feature selection algorithm:

[0095] The specific steps of the Relief feature selection algorithm are as follows:

[0096] (2-1) Given the training set {(x1,y1),(x2,y2),L,(x m ,y m )}, randomly select a sample x i , and then find the nearest neighbor sample x i from the samples of the same class as the sample x i,nh , called the near hit, and find the nearest neighbor sample x i from the samples of different classes from the sample x i,nm , called the near miss;

[0097] (2-2) Update the component δ corresponding to the relevant statistic for the attribute j according to the following rules j : If the distance between the sample x i and the near hit x i,nh on a certain feature j is less than the distance between the sample x i and the near miss x i,nm , then the feature j is beneficial to distinguishing the nearest neighbors of the same class and different classes, and the weight of the feature j is increased; conversely, if the distance between the sample x i and the near hit x i,nh on a certain feature j is greater than the distance between the sample x i and the near miss x i,nm , the feature j has a negative effect on distinguishing the nearest neighbors of the same class and different classes, and the weight of the feature j is reduced;

[0098] The calculation formula for implementing the above steps is as follows, where i in the following formula indicates the subscript of the randomly selected sample:

[0099]

[0100] where represents the value of the sample x a on the attribute j, depends on the type of the attribute j.

[0101] For the discrete attribute j, the calculation formula is as follows

[0102]

[0103] For the continuous attribute j, the calculation formula is as follows

[0104]

[0105] (2 - 3) Repeat the above process m times, average the obtained results, and finally obtain the relevant statistical component of each attribute. The value of the relevant statistical component characterizes the weight of the corresponding feature. The greater the weight of the feature, the stronger the classification ability of the feature; conversely, the weaker the classification ability of the feature.

[0106] (2 - 4) Specify a threshold τ, and select the features whose feature weight values are greater than τ, or specify the number k of features to be selected, and select k features.

[0107] (3) Extract the significant feature set of the proliferation or apoptosis of ovarian cancer cells in the body intervened by bevacizumab based on the alternative significant feature set;

[0108] In the previous step, based on the Relief feature selection algorithm, an alternative significant feature set with k features is obtained, and the features of the alternative significant feature set are initially sorted in descending order according to the weight of the k features;

[0109] Construct an alternative significant feature evaluation set. Select the first i features in the alternative significant features for the i - th time, where i = 1, 2, 3, …, k, and k is the number of features in the alternative feature set. Based on the full - feature classification model obtained in step (1), adjust the input features to the currently selected i features, perform multiple trainings and hyperparameter optimizations, etc., to obtain the generalization ability of the ExtraTree classification model, and record the optimal generalization ability as β;

[0110] Taking α obtained in step (1) as a reference, if β ≥ 0.75 or β ≥ 0.9 * α, it means that the significant features of the proliferation or apoptosis of ovarian cancer cells in the body intervened by bevacizumab are successfully extracted. Conversely, increase the value of k, go to step (2), and continue the iteration until the significant features of the proliferation or apoptosis of ovarian cancer cells in the body intervened by bevacizumab are successfully extracted;

[0111] The above iteration termination condition of β ≥ 0.9 * α is regarded as the significant rate of introducing the full - feature subset. The significant rate of the feature subset here is defined as where ξ = 1.0×10 -6, this parameter is used to prevent the error of a zero denominator, which is the ratio of the generalization ability value β of the best classification model based on the subset to the generalization ability index value α of the best classification model based on the full-feature data; the larger this index value, the better. Being close to 0 indicates that the feature subset is not a significant feature set, being close to 1 indicates that the feature subset can be regarded as a significant feature set, and being greater than 1 indicates that the feature subset can improve the prediction ability for the classification problem of the proliferation or apoptosis of ovarian cancer cells in the body intervened with bevacizumab; here, it is recommended to take the threshold of this ratio as 0.9; meanwhile, the β≥0.75 iteration termination condition indicates that when the generalization ability of the selected feature subset for the classification task of the survival prognosis of ovarian cancer with the proliferation or apoptosis of ovarian cancer cells in the body intervened with bevacizumab is greater than 0.75, it can be regarded as extracting significant features. If it is too low, it may indicate that it is necessary to return to step (2) and iterate again to extract the significant features of the survival prognosis of ovarian cancer;

[0112] (4) Based on the prediction model for the proliferation or apoptosis of ovarian cancer cells in the body intervened with bevacizumab extracted, the generalization ability of the model uses the classification accuracy rate on the test set. On the basis of the previous step, further training and hyperparameter optimization of the model are carried out to obtain the prediction model for the survival prognosis of ovarian cancer based on significant features.

[0113] The third step is to construct a model for extracting significant features and predicting the duration of cell proliferation or apoptosis of ovarian cancer cells in the body intervened with bevacizumab:

[0114] (1) Construct a full-feature regression model for the duration of cell proliferation or apoptosis of ovarian cancer cells in the body intervened with bevacizumab:

[0115] (1-1) Based on the selected dataset division scheme, establish a full-feature regression model for the duration of cell proliferation or apoptosis of ovarian cancer cells in the body intervened with bevacizumab based on all feature data of ovarian cancer cells. The generalization ability index of the model is measured by the coefficient of determination R 2 of the test set;

[0116] (1-2) Model training and hyperparameter optimization. Hyperparameter optimization methods can include grid search, random search, and Bayesian optimization, etc.;

[0117] (1-3) Select the extreme random tree (ExtraTree) regression prediction model with the optimal generalization ability of the model, and record the corresponding generalization ability as γ;

[0118] (2) Based on the full-feature regression prediction model, obtain an alternative significant feature set for the duration of cell proliferation or apoptosis of ovarian cancer cells in the body intervened with bevacizumab:

[0119] The number k of features in the alternative significant feature set is defaulted to 30;

[0120] (2-1) According to the obtained ExtraTree regression prediction model, the weights of all features of the ovarian cancer cells in the body are obtained. The larger the weight, the greater the importance of the corresponding feature;

[0121] (2-2) According to the weights of the features, select the top k important features as the alternative significant feature set that affects the proliferation or apoptosis duration of ovarian cancer cells in the body intervened by bevacizumab;

[0122] (2-3) Arrange the features in the alternative significant feature set in descending order according to the weights of the features, and the result is used as the final result of the alternative significant feature set;

[0123] (3) Extract the significant feature set that affects the proliferation or apoptosis duration of ovarian cancer cells in the body intervened by bevacizumab based on the alternative significant feature set:

[0124] (3-1) Based on the alternative significant feature set that affects the proliferation or apoptosis duration of ovarian cancer cells in the body intervened by bevacizumab, select the top i features, where i = 1, 2, 3, …, k;

[0125] (3-2) Based on these i features, establish an ExtraTree regression prediction model for the proliferation or apoptosis duration of ovarian cancer cells in the body intervened by bevacizumab, perform model training and hyperparameter optimization. The model generalization ability index is determined by the coefficient of determination R of the test set 2 Measure, and select the optimal value of the generalization ability as the prediction ability index value η of these i features for the proliferation or apoptosis duration after ovarian cancer cell intervention i ;

[0126] (3-3) For η = {η1, η2, L, η k )), select the largest value among them as η max , and obtain the corresponding number of features numbers;

[0127] (3-4) When η max ≥0.9*γ, or η max ≥0.75, then output the top numbers features in the alternative significant feature set as the significant features. Otherwise, increase the value of k and go back to step (2);

[0128] (4) Establish a prediction model for the proliferation duration data of ovarian cancer cells in the body intervened by bevacizumab based on the extracted significant features:

[0129] Based on the significant features of the proliferation or apoptosis duration of ovarian cancer cells in the body intervened by bevacizumab extracted above, establish a prediction model for the proliferation or apoptosis duration of ovarian cancer cells in the body intervened by bevacizumab. The model generalization ability index is determined by the coefficient of determination R on the test set2 Measure. The model selects the ExtraTree regression model. Based on the above model training results and hyperparameter results, further training and hyperparameter optimization are carried out to obtain a prediction model for the data of the duration of proliferation or apoptosis of ovarian cancer cells in the body affected by the intervention of bevacizumab based on the extracted significant features.

[0130] (5) The steps for obtaining the risk factors affecting the proliferation of ovarian cancer cells in the body after bevacizumab intervention are as follows:

[0131] Based on the significant features of the proliferation or apoptosis of ovarian cancer cells in the body intervened by bevacizumab and the significant features of the duration of proliferation or apoptosis of ovarian cancer cells in the body affected by the intervention of bevacizumab, taking the intersection of these two sets of significant features can obtain the common significant influencing features; taking the union of these two sets of significant features can obtain all significant features; based on the above steps, the risk factors affecting the proliferation of ovarian cancer cells in the body after bevacizumab intervention can be obtained.

[0132] To explore the significant influencing features of the proliferation or apoptosis data of ovarian cancer cells in the body intervened by other drugs, based on the above description and the idea of the steps, relevant index data of ovarian cancer cells in the body can be collected and significant features can be extracted, and a prediction model for the proliferation or apoptosis characteristic data of ovarian cancer cells in the body after drug intervention based on the significant features can be constructed; based on the significant features of each important data, the intersection and union set operations of the corresponding significant features can be carried out to obtain a set of significant features of various data affecting the proliferation or apoptosis of ovarian cancer cells in the body after drug intervention; the set of extracted significant features and the prediction model obtained in the extraction process are very helpful for providing a more accurate and efficient drug resistance trend prediction for the drug intervention of ovarian cancer cells in the body.

[0133] The method proposed by the present invention will be described below by taking a certain actual ovarian cancer data set as an example:

[0134] The selected dataset contains 138 cases of ovarian cancer cells in the body. The basic information such as total protein, albumin, globulin, albumin / globulin ratio, alkaline phosphatase, lactate dehydrogenase, creatinine, urea, potassium, sodium, chloride, bicarbonate, calcium, phosphorus, glucose, total cholesterol, triglyceride, high-density lipoprotein cholesterol, non-high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, very-low-density lipoprotein cholesterol, apolipoprotein A1, apolipoprotein B, lipoprotein a, free fatty acids, and pathological type in the ovarian cancer tumor microenvironment was collected as research characteristics. The proliferation or apoptosis of ovarian cancer and the duration of proliferation or apoptosis after bevacizumab intervention were used as research objects. According to the prior knowledge obtained from basic research, in the case of using bevacizumab to treat ovarian cancer, abnormal lipid metabolism will inhibit the effect of bevacizumab treatment, shorten the progression-free survival period, and abnormal lipid metabolism occurs after the tumor continues to proliferate after bevacizumab intervention. The main factors affecting the proliferation or apoptosis of ovarian cancer after intervention extracted by the model adopted in the present invention include non-high-density lipoprotein cholesterol, potassium, free fatty acids, high-density lipoprotein cholesterol, plasma thrombin time determination, fibrinogen content, apolipoprotein A1, etc. Thus, it can be confirmed that blood lipid content has a certain impact on tumor proliferation / apoptosis. Based on the extracted prediction model for the proliferation or apoptosis of ovarian cancer after bevacizumab intervention, the generalization ability can reach 83.33%.

[0135] As Figure 3a 、 Figure 3b shown, Figure 3 shows the ROC curve and calibration curve of the Extra Tree model adopted in the present invention. As Figure 4a and Figure 4b shown, it is a comparison of the ROC curves and calibration curves of several models with better performance in the tested methods. Among them, the AUC value of the Extra Tree model adopted in the present invention is the largest, reaching 0.86, and its performance is significantly better than other methods. The most important characteristics affecting the duration of proliferation or apoptosis of ovarian cancer cells in the body after bevacizumab intervention include glycated serum protein (fructosamine), high-density lipoprotein cholesterol, plasma prothrombin time activity, age, very-low-density lipoprotein cholesterol, lipoprotein, chloride, calcium, uric acid, creatinine, direct bilirubin, sodium, etc., indicating that some raw materials and products in the lipid metabolism process do have an impact on the length of the progression-free survival period. Based on the extracted significant characteristics of the data on the proliferation duration of ovarian cancer cells in the body after bevacizumab intervention, the generalization ability of the prediction model reaches 87.85%. Through Figure 5The nomogram shown can more intuitively reflect the relationship between the selected significant features and the proliferation duration data. The total score is 364. If the defined duration is less than 210, it is considered excellent, and the odds ratio is 2.08. Based on the significant features of the proliferation or apoptosis of ovarian cancer cells in the body intervened by bevacizumab and the significant features affecting the duration of proliferation or apoptosis of ovarian cancer cells in the body intervened by bevacizumab, by performing the intersection and union operations of the sets, the common influencing features and all influencing features affecting the proliferation or apoptosis and the duration of proliferation or apoptosis of ovarian cancer cells in the body can be obtained respectively, so as to better understand the risk factors of various index data. The present invention avoids the problems of long time consumption and low precision in traditional manual analysis, applies machine learning to medicine, and provides a reliable and high-precision new tool for cancer-related research.

[0136] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting the proliferation or apoptosis of ovarian cancer cells in an organism intervened by bevacizumab, characterized in that, Including the following steps: 11) Acquisition and preprocessing of basic data of ovarian cancer cells in the body: Obtain the basic data of ovarian cancer cells in the body, including the basic data intervened by bevacizumab, extract and label the basic data of ovarian cancer cells before and after intervention with bevacizumab, and whether ovarian cancer cells proliferate / apoptose as labels. Perform data encoding, data cleaning, and normalization preprocessing operations on the basic data to obtain processed data, and divide the data set; 12) Construct a full-feature regression model: 121) Based on the selected dataset partitioning scheme, a full-feature regression model for the proliferation / apoptosis duration of ovarian cancer cells in the body affected by the application of bevacizumab is established based on all the characteristic data of ovarian cancer cells in the body. The model generalization ability index is measured by the coefficient of determination R of the test set 2 Measure; 122) Model training and hyperparameter optimization; 123) Select the extremely randomized tree regression prediction model with the best model generalization ability; 13) Construct the number k of features in the alternative significant feature set; 131) According to the above-mentioned extremely randomized tree regression prediction model, obtain the weights of all features of ovarian cancer cells in the body. The greater the weight, the greater the importance of the corresponding feature; 132) According to the weights of the features, select the top k important features as the alternative significant feature set affecting the proliferation / apoptosis duration of ovarian cancer cells in the body intervened by bevacizumab; 133) Sort the features in the alternative significant feature set in descending order according to the weights of the features, and the result is used as the final result of the alternative significant feature set; 14) Construct an alternative significant feature set: 141) Based on the alternative significant feature set affecting the proliferation / apoptosis duration of ovarian cancer cells in the body intervened by bevacizumab, select the first i features, where i = 1, 2, 3,..., k; 142) Based on these i features, an extreme random tree regression prediction model for the proliferation / apoptosis duration of ovarian cancer cells in the body affected by the intervention of bevacizumab is established, model training and hyperparameter optimization are carried out, and the model generalization ability index is measured by the coefficient of determination R of the test set 2 . The optimal value of the generalization ability is selected as the prediction ability index value η of these i features for the proliferation / apoptosis duration after the intervention of ovarian cancer cells i ; 143) For η = {η1, η2, …, η k ), select the largest value among them as η max , and obtain the corresponding number of features numbers; 144) When η max ≥ 0.9 * γ, or η max ≥ 0.75, then output the first numbers features in the alternative significant feature set as the significant features; otherwise, increase the value of k and go to step 13); 15) Based on the significant features of the duration of ovarian cancer cell proliferation / apoptosis in the body affected by bevacizumab extraction in step 14), a prediction model for the duration of ovarian cancer cell proliferation / apoptosis in the body affected by bevacizumab is established. The model generalization ability index is determined by the coefficient of determination R on the test set 2 The model uses the extremely randomized tree regression model. Based on the model training results and hyperparameter results of the full-feature regression prediction model extracted using all features in step 12), further training and hyperparameter optimization are carried out to obtain a prediction model for the data of the duration of ovarian cancer cell proliferation / apoptosis in the body affected by bevacizumab based on the extracted significant features, and the prediction results of the duration of ovarian cancer cell proliferation or apoptosis in the body affected by bevacizumab are output.

2. The predictive method for the proliferation or apoptosis of ovarian cancer cells in an organism intervened with bevacizumab according to claim 1, characterized in that The acquisition and preprocessing of the basic data of ovarian cancer cells in the body include the following steps: 21) Acquisition of ovarian cancer cell data in the body: Collect and record the information data of ovarian cancer cells of different bodies, including three types of information data; The first type of data is the detection index data of total protein, albumin, globulin, albumin / globulin ratio, alkaline phosphatase, lactate dehydrogenase, creatinine, urea, potassium, sodium, chloride, bicarbonate, calcium, phosphorus, glucose, total cholesterol, triglyceride, high-density lipoprotein cholesterol, non-high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, very-low-density lipoprotein cholesterol, apolipoprotein A1, apolipoprotein B, lipoprotein a, and free fatty acids in the ovarian cancer tumor microenvironment when starting to intervene with bevacizumab; The second type of data is the detection index data of total protein, albumin, globulin, albumin / globulin ratio, alkaline phosphatase, lactate dehydrogenase, creatinine, urea, potassium, sodium, chloride, bicarbonate, calcium, phosphorus, glucose, total cholesterol, triglyceride, high-density lipoprotein cholesterol, non-high-density lipoprotein cholesterol, low-density lipoprotein cholesterol, very-low-density lipoprotein cholesterol, apolipoprotein A1, apolipoprotein B, lipoprotein a, and free fatty acids in the ovarian cancer tumor microenvironment during tumor proliferation / apoptosis; The third type of data is the categorical variable data of whether the ovarian cancer cell population continues to proliferate or apoptose; 22) Preprocessing of the data of ovarian cancer cells in the body intervened by bevacizumab. Perform data encoding, data cleaning, data normalization processing, and data set division on the data of ovarian cancer cells in the body intervened by bevacizumab. The specific steps are as follows: Coding of data of ovarian cancer cells in the body intervened by bevacizumab: Non-numerical data in the data is coded into numerical data for algorithm processing, and the same type of data is coded into one code to achieve data coding. The operation object includes the classification variable of proliferation / apoptosis; 222) Cleaning of data of ovarian cancer cells in the body intervened by bevacizumab: When extracting significant features and making predictions based on the feature data of ovarian cancer cells before intervention, the feature variable data after intervention is deleted; the feature data with more than half of the missing values is deleted; the abnormal feature data is found and deleted through the visualization methods of combined scatter plots, box plots, and variable distribution plots; for discrete data, median filling or mode filling operations are used, and for continuous data, mean filling is used; 223) Normalization processing of data of ovarian cancer cells in the body intervened by bevacizumab: The data processed above is normalized to achieve data normalization processing; 224) Dataset division of the prediction model: The processed data is randomly divided into a training set and a test set according to a ratio of 8:

2. The proliferation / apoptosis ratio of ovarian cancer cells in the test set and the training set is balanced and reasonable, and is used for the training and testing of the proliferation / apoptosis prediction model of ovarian cancer cells in the body after intervention with bevacizumab; 225) Extraction of significant features of the proliferation / apoptosis duration after bevacizumab intervention and dataset division of the prediction model: Ensure that the distribution of the proliferation / apoptosis time of ovarian cancer cells in the body after intervention with bevacizumab in the test set and the training set is consistent, and is used for the training and testing of the model for extracting significant features and predicting the proliferation / apoptosis duration of ovarian cancer cells in the body after intervention with bevacizumab. The specific steps are as follows: 2251) The dataset of the model for extracting significant features of the proliferation / apoptosis duration after bevacizumab intervention and predicting is randomly divided into a significant feature training set and a significant feature test set according to a ratio of 8:2; Plot the data distribution curves of the cell proliferation / apoptosis duration after ovarian cancer intervention on the significant feature training set and the significant feature test set, and preliminarily establish a full-feature regression model for the cell proliferation / apoptosis duration after ovarian cancer medication. Select the polynomial regression and decision tree regression models without adjusting the model hyperparameters. The model generalization ability index is determined by the coefficient of determination R 2 of the test set. Observe the model generalization ability and the duration data distribution curves on the training set and the test set. The calculation formula of the coefficient of determination R 2 is as follows In the above formula, y i represents the true data of the proliferation / apoptosis duration of ovarian cancer cells in the body i, represents the estimated value of the proliferation / apoptosis duration obtained by the model for estimating ovarian cancer cells in the body i, represents the average value of the true proliferation / apoptosis duration data of ovarian cancer cells in the body, n is the total number of samples, and SSR, SSE, and SST represent the regression sum of squares, the residual sum of squares, and the total sum of squares of deviations, respectively; 2253) If the generalization ability of the model is greater than 0.70, and the data distribution curves on the training set and the test set are approximately similar, then select this dataset division scheme; otherwise, go to step 2251).

3. The predictive method for the proliferation or apoptosis of ovarian cancer cells in an organism intervened by bevacizumab according to claim 1, characterized in that, It also includes the analysis steps of risk factors, that is, according to the prediction results of the proliferation or apoptosis of ovarian cancer cells in the body intervened by bevacizumab, the risk factors affecting the proliferation of ovarian cancer cells in the body after intervention with bevacizumab are obtained. The specific steps are as follows: Based on the significant features of the proliferation / apoptosis of ovarian cancer cells in the body intervened by bevacizumab and the significant features affecting the proliferation / apoptosis duration of ovarian cancer cells in the body intervened by bevacizumab extracted, the intersection of these two sets of significant features is taken to obtain the common significant influencing features; The union of these two sets of significant features is taken to obtain all significant features; comprehensively analyze the features in the intersection that have significant effects on both the progress direction of proliferation / apoptosis and the proliferation / apoptosis duration, and the significant features in the union that affect at least one of the progress direction and duration, to obtain the risk factors affecting the proliferation of ovarian cancer cells in the body after intervention with bevacizumab.

4. A prediction method for the proliferation or apoptosis of ovarian cancer cells in an organism intervened by bevacizumab according to claim 1, characterized in that, It also includes the data analysis method based on the proliferation / apoptosis data of ovarian cancer cells intervened by drugs. The steps are as follows: Collect the index data of ovarian cancer cells in the body, extract significant features, and construct a prediction model for the proliferation / apoptosis characteristic data of ovarian cancer cells in the body after drug intervention; perform union and intersection operations on the significant features of each important data to obtain a significant feature set of various data on the proliferation / apoptosis of ovarian cancer cells in the body after drug intervention; the obtained significant feature set and the prediction model obtained in the extraction process provide a drug resistance trend prediction for drug intervention in ovarian cancer cells in the body.

Citation Information

Patent Citations

  • Method for predicting growth trend of tumor

    CN110428905A

  • Ovarian cancer prognosis risk model based on polyunsaturated fatty acid related gene and preparation method and application thereof

    CN114807374A