A method for determining the IC50 value of cytotoxicity of selenium compounds to cancer cells based on gene expression
The model established through the FCBF algorithm and Gauss-Newton iterative algorithm solves the problems of long time and large influence of human factors in determining the IC50 value of the toxicity of selenium compounds on cancer cells in the existing technology, realizes fast and accurate IC50 value determination, and guides the dosage of selenium compounds in animal experiments.
Patent Information
- Application Number
- CN202210534900.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-05-17
AI Technical Summary
The existing methods for determining the IC50 value of the toxicity of selenium compounds to cancer cells are time-consuming, cumbersome, and subject to significant human influence, making it difficult to efficiently and accurately determine the IC50 value of a large number of cancer cells.
The FCBF algorithm was used to reduce the dimensionality of cancer cell gene expression data, and a Gauss-Newton iterative algorithm model was established. The IC50 value of the toxicity of selenium compounds on cancer cells was determined through cancer cell gene expression.
It achieves the rapid and accurate determination of the IC50 value of selenium compounds against cancer cells, guides the dosage of selenium compounds in tumor-bearing animals, reduces pharmacodynamic risks, and improves experimental efficiency and accuracy.
Smart Images

Figure QLYQS_1 
Figure QLYQS_2 
Figure QLYQS_4
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of detection and evaluation of the sensitivity of selenium compounds to cancer cells, and particularly to a method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on gene expression. Background Art
[0002] Selenium, as an essential trace element for the human body, plays a vital role in the body, and selenium compounds are increasingly being recognized as potential anti-tumor drugs. The half-inhibitory concentration (IC50) value of cancer cells is often used as an important indicator to detect the sensitivity of cells to selenium compounds. The IC50 value can further guide the dosage of selenium compounds in tumor-bearing animals, reducing adverse consequences such as the inability to observe pharmacodynamics due to too low a dose or the direct death of animals due to too high a dose. This is conducive to observing the inhibitory effect of selenium compounds on tumors at the animal level, and then developing selenium compounds that are effective in treating tumors. The IC50 value of the toxicity of selenium compounds to cancer cells provides an important reference for the dosage of selenium compounds in tumor-bearing animals. However, the existing conventional experimental methods for determining the IC50 value of the toxicity of selenium compounds to cancer cells (such as the MTT method) have long detection times, cumbersome processes, and are greatly affected by human factors. Faced with a large number of individual cancer cell toxicity IC50 values, the detection will consume a lot of time and energy of scientific research staff. Therefore, developing a method that can replace experimental cell experiments to determine the IC50 value of the toxicity of selenium compounds to cancer cells will significantly reduce unnecessary experimental time, avoid and reduce experimental costs, and more efficiently and quickly determine the IC50 value of selenium compounds to cancer cells, thereby providing an important reference for the dosage of animal experiments.
[0003] Therefore, there is a need in the art to develop a simple and rapid method for determining the IC50 value of the toxicity of selenium compounds to cancer cells. Summary of the Invention
[0004] The object of the present invention is to provide a method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression, wherein the method can accurately, efficiently and quickly determine the IC50 value of selenium compounds to cancer cells.
[0005] In a first aspect, the present invention provides a method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression, the method comprising the steps of:
[0006] (1) The FCBF algorithm is used to reduce the dimensionality of cancer cell gene expression data to obtain the reduced dimensionality cancer cell gene expression dataset F′;
[0007] (2) dividing the cancer cell gene expression dataset F′ obtained after dimensionality reduction in step (1) into a training sample subset F′ and a test sample subset F′;
[0008] (3) Gauss-Newton iterative algorithm is performed using the F′ training sample subset and the IC50 values of the toxicity of selenium compounds to different cancer cells to obtain a determination model for the IC50 value of the toxicity of selenium compounds to cancer cells based on the gene expression of cancer cells; through the determination model, the IC50 value of the toxicity of selenium compounds to cancer cells is determined according to the gene expression of cancer cells.
[0009] Preferably, the method comprises a predictive method.
[0010] Preferably, the method is carried out under in vitro cell culture conditions.
[0011] Preferably, the method comprises an in vitro assay method and / or an auxiliary assay method.
[0012] Preferably, the methods include non-diagnostic and / or non-therapeutic methods.
[0013] Preferably, the cancer comprises a human cancer or a non-human mammalian cancer.
[0014] Preferably, the cancer includes one or more of lung cancer, liver cancer, colon cancer, breast cancer and immune system cancer.
[0015] Preferably, the lung cancer is selected from the group consisting of non-small cell lung cancer, small cell lung cancer, or a combination thereof.
[0016] Preferably, the lung cancer includes lung adenocarcinoma and non-small cell lung cancer.
[0017] Preferably, the immune system cancer comprises leukemia.
[0018] Preferably, the immune system cancer includes one or more of acute promyelocytic leukemia and monocytic leukemia.
[0019] Preferably, the cancer cells include one or more of A549 cells, NCI-H522 cells, HEPG2 cells, HEP38217 cells, HCT116 cells, MCF-7 cells, HL-60 cells and THP-1 cells.
[0020] Preferably, the selenium compound is selected from the following group of selenium compounds or their analogs: selenomethionine (Se-Met), selenocysteine (Se-cys), selenomethylselenocysteine hydrochloride (D-Se-cys), sodium selenite (Sod-sem), or a combination thereof.
[0021] Preferably, the selenium compound analogue includes a pharmaceutically acceptable salt of a selenium compound or a prodrug of a selenium compound.
[0022] Preferably, the pharmaceutically acceptable salt of the selenium compound includes a salt formed by a selenium compound and hydrochloric acid, mucic acid, D-glucuronic acid, hydrobromic acid, hydrofluoric acid, hydroiodic acid, sulfuric acid, nitric acid, phosphoric acid, formic acid, acetic acid, trifluoroacetic acid, propionic acid, oxalic acid, malonic acid, succinic acid, fumaric acid, maleic acid, lactic acid, malic acid, tartaric acid, citric acid, picric acid, methanesulfonic acid, phenylmethanesulfonic acid, benzenesulfonic acid, aspartic acid or glutamic acid.
[0023] Preferably, the structure of the selenium compound analog is similar to the structure of a selenium compound selected from the group consisting of selenomethionine (Se-Met), selenocysteine (Se-cys), selenomethylselenocysteine hydrochloride (D-Se-cys), sodium selenite (Sod-sem), or a combination thereof.
[0024] Preferably, the skeleton structure of the selenium compound analogue is the same as or similar to the skeleton structure of the selenium compound.
[0025] Preferably, the gene expression comprises mRNA expression.
[0026] Preferably, the gene expression includes gene expression level.
[0027] Preferably, the cancer cell gene expression comprises the cancer cell gene expression in the CCLE database.
[0028] Preferably, the cancer cell gene expression includes the mRNA expression level of the cancer cells described in Table 1 of the specification.
[0029] Preferably, the obtained cancer cell gene expression dataset F′ after dimensionality reduction includes the cancer cell gene expression dataset F′ after dimensionality reduction described in Table 2 of the specification.
[0030] Preferably, the step (1) comprises:
[0031] (1.1) Organize the cancer cell gene expression data and express it in the form of the following matrix:
[0032]
[0033] There are i samples and j genes in total. The i-th row is represented by Si, which represents the sample with cancer cell category Ci, where Ci = {C1, C2, ..., C M The jth column is denoted by Fj, which represents the gene expression data of cancer cells. Then Fij is the expression data of the jth gene in the i-th sample, where i = 1, ..., M, and j = 1, 2, ..., N.
[0034] (1.2) The correlation between the gene expression data Fj of the cancer cells in the established matrix and the cancer cell category Ci is measured. The correlation is represented by CSU′i:
[0035]
[0036] (1.3) Arrange in descending order to obtain the gene expression sequence F′:
[0037] F′=F′-{F j |CSU′ i =0, i=1,…,M}
[0038] Among them, F′ on the left side of the equation represents the new feature subset;
[0039] (1.4) Delete all features with CSU′i=0 in F′, select a feature from F′ to become the first element in the feature subset, store F1′ in the set, and start from the second feature in the original F′ to search and delete the approximate redundant cover formed by all features of F1′ to obtain a new feature subset:
[0040] F′=F′-{F j |CSU′ i >CSU′ j And CSU′ i ≤ISU′ i,j, i=1,…,M, j=1,…,N)
[0041] Where F′ on the left side of the equation represents the new feature subset;
[0042] (1.5) Set F1′ to the next remaining feature in the new feature subset F′ and repeat step (1.4) until the last feature of F′;
[0043] (1.6) The cancer cell gene expression dataset F′ after feature selection is obtained, that is, the cancer cell gene expression dataset F′ after dimensionality reduction is obtained.
[0044] Preferably, in step (2), the cancer cell gene expression dataset F′ obtained in step (1) after dimensionality reduction is divided into an F′ training sample subset and an F′ test sample subset according to a ratio of 6:2.
[0045] Preferably, the step (3) comprises:
[0046] (3.1) Assume is the regression coefficient to be estimated b=(b0,b1,...,b M ) T The initial value of , the nonlinear regression model is:
[0047] y i =f(x i ,b)+ε i , (i=1, 2, ..., M)
[0048] Where y is the IC50 value of the toxicity of selenium compounds to cancer cells, x is the gene expression value in the F′ training sample subset, ε is the error, and b is the regression coefficient to be estimated;
[0049] (3.2) f(x i , b) at the initial value g 0 Taylor expansion is done here:
[0050]
[0051] k represents the position distance with regression coefficient;
[0052] (3.3) Substituting the formula obtained in step (3.2) into the nonlinear regression model in step (3.1) yields:
[0053]
[0054] Let Y = y i -f(x i , g 0 ), but
[0055]
[0056] (3.4) is expressed as a matrix The relationship between the gene expression of cancer cells and the IC50 of selenium compounds on cancer cells is: Y = BX + E, let:
[0057]
[0058] (3.5) The Gauss-Newton iteration method is used to find the least squares fit and estimate the corrected regression coefficient B:
[0059] B=(X T X) -1 X T Y
[0060] Let g 0 is the first iteration value, then g 1 =g 0 +b 0 ;
[0061] (3.6) Set the residual sum of squares to:
[0062]
[0063] Where s is the number of iterations;
[0064] (3.7) Repeat step (3.5) and, under a given tolerance value ε, When the accuracy requirement is met, the iteration is stopped, g s To output the result;
[0065] (3.8) Obtain a model for determining the IC50 value of selenium compounds' toxicity to cancer cells based on cancer cell gene expression: y = g s X+ε,
[0066] Where y is the IC50 value of the toxicity of selenium compounds to cancer cells, X is the gene expression value of cancer cells, and ε is the error;
[0067] The IC50 value of the toxicity of selenium compounds to cancer cells is determined based on the gene expression of cancer cells using the assay model.
[0068] Preferably, the step (3) further includes detecting the reliability of the measurement model, wherein detecting the reliability of the measurement model includes:
[0069] Substitute the cancer cell gene expression data of the F′ test sample subset into the determination model y=g s In X+ε, the IC50 value of the toxicity of selenium compounds to cancer cells is obtained and compared with the actual experimental value to test the reliability of the model.
[0070] A second aspect of the present invention provides a device, comprising:
[0071] processor; and
[0072] The memory stores computer instructions. When the computer instructions are executed by the processor, the processor executes the method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression as described in the first aspect of the present invention.
[0073] In a third aspect, the present invention provides a non-transitory computer storage medium storing a computer program. When the computer program is executed by one or more processors, the processors execute the method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression as described in the first aspect of the present invention.
[0074] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features specifically described below (such as embodiments) can be combined with each other to form new or preferred technical solutions. DETAILED DESCRIPTION
[0075] The present invention develops a method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression. The method overcomes the shortcomings of conventional experimental methods (such as the MTT method) for determining the IC50 value of the toxicity of selenium compounds to cancer cells, such as long detection time, complicated and cumbersome process, and great influence of human factors. In addition, the detection of the IC50 values of the toxicity of a large number of individual cancer cells will consume a lot of time and energy of scientific research personnel. In the method described, after the FCBF algorithm is used to reduce the dimensionality of cancer cell gene expression data, a Gauss-Newton iterative algorithm is used to establish a correlation between the reduced dimensionality gene expression data and the IC50 value of the toxicity of selenium compounds to cancer cells, thereby obtaining a determination model for the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression. Through the determination model, the IC50 value of the toxicity of selenium compounds to cancer cells is determined based on the gene expression data, thereby accurately, efficiently and quickly determining the IC50 value of the selenium compound to cancer cells. The determined IC50 value can then be used to guide the animal dosage of selenium compounds for tumor-bearing animals, thereby reducing adverse experimental consequences such as the inability to observe pharmacodynamics due to too low an animal dosage or the direct death of the animal due to too high an administration dose, thereby facilitating the investigation of the inhibitory effect of selenium compounds and their analogues on tumors at the animal level, thereby developing selenium compounds that are effective in treating tumors.
[0076] the term
[0077] As used herein, the terms "include," "comprise," and "contain" are used interchangeably to encompass not only open definitions but also semi-closed and closed definitions. In other words, the terms encompass "consisting of," "consisting essentially of."
[0078] As used herein, the term “FCBF algorithm” stands for Fast Correlation-Based Filter Solution.
[0079] As used herein, the term "CCLE" stands for Cancer Cell Line Encyclopedia.
[0080] As used herein, "IC50 value" and "IC 50 The term "half-inhibitory concentration" is used interchangeably to refer to the 50% inhibitory concentration, which is the concentration of an inhibitor (such as selenium compounds and their analogs that are toxic to cancer cells) that achieves 50% inhibition.
[0081] As used herein, the English term “Gauss-Newton iteration method” is Gauss-Newton iteration method.
[0082] As used herein, the English name of “Taylor expansion” is Taylor expansion.
[0083] method
[0084] The present invention provides a method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression, the method comprising the steps of:
[0085] (1) The FCBF algorithm is used to reduce the dimensionality of cancer cell gene expression data to obtain the reduced dimensionality cancer cell gene expression dataset F′;
[0086] (2) dividing the cancer cell gene expression dataset F′ obtained after dimensionality reduction in step (1) into a training sample subset F′ and a test sample subset F′;
[0087] (3) Gauss-Newton iterative algorithm is performed using the F′ training sample subset and the IC50 values of the toxicity of selenium compounds to different cancer cells to obtain a determination model for the IC50 value of the toxicity of selenium compounds to cancer cells based on the gene expression of cancer cells; through the determination model, the IC50 value of the toxicity of selenium compounds to cancer cells is determined according to the gene expression of cancer cells.
[0088] The method of the present invention for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression can be a prediction method, an in vitro determination method and / or an auxiliary determination method.
[0089] Preferably, the method is carried out under in vitro cell culture conditions.
[0090] The cancer described in the present invention may include (but is not limited to) one or more of lung cancer, liver cancer, colon cancer, breast cancer and immune system cancer.
[0091] In a preferred embodiment of the present invention, the lung cancer is selected from the group consisting of non-small cell lung cancer, small cell lung cancer, or a combination thereof.
[0092] In a preferred embodiment of the present invention, the lung cancer includes lung adenocarcinoma and non-small cell lung cancer.
[0093] In a preferred embodiment of the present invention, the immune system cancer includes leukemia.
[0094] In a preferred embodiment of the present invention, the immune system cancer includes one or more of acute promyelocytic leukemia and monocytic leukemia.
[0095] The cancer cells described in the present invention may include (but are not limited to) one or more of A549 cells, NCI-H522 cells, HEPG2 cells, HEP38217 cells, HCT116 cells, MCF-7 cells, HL-60 cells and THP-1 cells.
[0096] The selenium compounds described in the present invention include (but are not limited to) selenium compounds or their analogs selected from the following group: selenomethionine (Se-Met), selenocysteine (Se-cys), selenomethylselenocysteine hydrochloride (D-Se-cys), sodium selenite (Sod-sem), or a combination thereof.
[0097] Preferably, the selenium compound analogue includes a pharmaceutically acceptable salt of a selenium compound or a prodrug of a selenium compound.
[0098] The pharmaceutically acceptable salts of the selenium compounds of the present invention are not particularly limited and may include (but are not limited to) salts formed by the selenium compounds and hydrochloric acid, mucic acid, D-glucuronic acid, hydrobromic acid, hydrofluoric acid, hydroiodic acid, sulfuric acid, nitric acid, phosphoric acid, formic acid, acetic acid, trifluoroacetic acid, propionic acid, oxalic acid, malonic acid, succinic acid, fumaric acid, maleic acid, lactic acid, malic acid, tartaric acid, citric acid, picric acid, methanesulfonic acid, phenylmethanesulfonic acid, benzenesulfonic acid, aspartic acid or glutamic acid.
[0099] In a preferred embodiment of the present invention, the structure of the selenium compound analog is similar to the structure of a selenium compound selected from the group consisting of selenomethionine (Se-Met), selenocysteine (Se-cys), selenomethylselenocysteine hydrochloride (D-Se-cys), sodium selenite (Sod-sem), or a combination thereof.
[0100] Preferably, the skeleton structure of the selenium compound analogue is the same as or similar to the skeleton structure of the selenium compound.
[0101] In the method of the present invention, the gene expression may include mRNA expression
[0102] In the method described in the present invention, the gene expression may include gene expression level.
[0103] In a preferred embodiment of the present invention, the cancer cell gene expression includes the cancer cell gene expression in the CCLE database.
[0104] Typically, the cancer cell gene expression includes the mRNA expression level of the cancer cells described in Table 1 of the specification.
[0105] Typically, the obtained cancer cell gene expression dataset F′ after dimensionality reduction includes the cancer cell gene expression dataset F′ after dimensionality reduction described in Table 2 of the specification.
[0106] In a preferred embodiment of the present invention, step (1) comprises:
[0107] (1.1) Organize the cancer cell gene expression data and express it in the form of the following matrix:
[0108]
[0109] There are i samples and j genes in total. The i-th row is represented by Si, which represents the sample with cancer cell category Ci, where Ci = {C1, C2, ..., C M The jth column is denoted by Fj, which represents the gene expression data of cancer cells. Then Fij is the expression data of the jth gene in the i-th sample, where i = 1, ..., M, and j = 1, 2, ..., N.
[0110] (1.2) The correlation between the gene expression data Fj of the cancer cells in the established matrix and the cancer cell category Ci is measured. The correlation is represented by CSU′i:
[0111]
[0112] (1.3) Arrange in descending order to obtain the gene expression sequence F′:
[0113] F′=F′-{F j |CSU′ i =0, i=1,…,M}
[0114] Among them, F′ on the left side of the equation represents the new feature subset;
[0115] (1.4) Delete all features with CSU′i=0 in F′, select a feature from F′ to become the first element in the feature subset, store F1′ in the set, and start from the second feature in the original F′ to search and delete the approximate redundant cover formed by all features of F1′ to obtain a new feature subset:
[0116] F′=F′-{F j |CSU′ i >CSU′ j And CSU′ i ≤ISU′ i,j ,i=1,…,M,j=1,…,N}
[0117] Where F′ on the left side of the equation represents the new feature subset;
[0118] (1.5) Set F1′ to the next remaining feature in the new feature subset F′ and repeat step (1.4) until the last feature of F′;
[0119] (1.6) The cancer cell gene expression dataset F′ after feature selection is obtained, that is, the cancer cell gene expression dataset F′ after dimensionality reduction is obtained.
[0120] In a preferred embodiment of the present invention, in step (2), the cancer cell gene expression dataset F′ obtained in step (1) after dimensionality reduction is divided into an F′ training sample subset and an F′ test sample subset according to a ratio of 6:2.
[0121] In a preferred embodiment of the present invention, the step (3) comprises:
[0122] (3.1) Assume is the regression coefficient to be estimated b=(b0,b1,...,b M ) T The initial value of , the nonlinear regression model is:
[0123] y i =f(x i , b)=+ε i , (i=1, 2, ..., M)
[0124] Where y is the IC50 value of the toxicity of selenium compounds to cancer cells, x is the gene expression value in the F′ training sample subset, ε is the error, and b is the regression coefficient to be estimated;
[0125] (3.2) f(x i , b) at the initial value g 0 Taylor expansion is done here:
[0126]
[0127] k represents the position distance with regression coefficient;
[0128] (3.3) Substituting the formula obtained in step (3.2) into the nonlinear regression model in step (3.1) yields:
[0129]
[0130] Let Y = y i -f(x i , g 0 ), but
[0131]
[0132] (3.4) is expressed as a matrix The relationship between the gene expression of cancer cells and the IC50 of selenium compounds on cancer cells is: Y = BX + E, let:
[0133]
[0134] (3.5) The Gauss-Newton iteration method is used to find the least squares fit and estimate the corrected regression coefficient B:
[0135] B=(X T X) -1 X T Y
[0136] Let g 0 is the first iteration value, then g 1 =g 0 +b 0 ;
[0137] (3.6) Set the residual sum of squares to:
[0138]
[0139] Where s is the number of iterations;
[0140] (3.7) Repeat step (3.5) and, under a given tolerance value ε, When the accuracy requirement is met, the iteration is stopped, g s To output the result;
[0141] (3.8) Obtain a model for determining the IC50 value of selenium compounds' toxicity to cancer cells based on cancer cell gene expression: y = g s X+ε,
[0142] Where y is the IC50 value of the toxicity of selenium compounds to cancer cells, X is the gene expression value of cancer cells, and ε is the error;
[0143] The IC50 value of the toxicity of selenium compounds to cancer cells is determined based on the gene expression of cancer cells using the assay model.
[0144] In a preferred embodiment of the present invention, the step (3) further includes detecting the reliability of the measurement model, and the detecting the reliability of the measurement model includes:
[0145] Substitute the cancer cell gene expression data of the F′ test sample subset into the determination model y=g s In X+ε, the IC50 value of the toxicity of selenium compounds to cancer cells is obtained and compared with the actual experimental value to test the reliability of the model.
[0146] Device
[0147] The present invention also provides a device, comprising:
[0148] processor; and
[0149] The memory stores computer instructions. When the computer instructions are executed by the processor, the processor executes the method of the present invention for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression.
[0150] Non-transitory computer storage media
[0151] The present invention also provides a non-transient computer storage medium storing a computer program. When the computer program is executed by one or more processors, the processors execute the method of the present invention for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression.
[0152] The main excellent technical effects of the present invention include:
[0153] The present invention has developed a method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression. The method can determine the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression data. Gene expression data has the advantages of small sample size and high dimensionality. The method described in the present invention can quickly, efficiently and accurately determine the IC50 value of the toxicity of selenium compounds to cancer cells based on gene expression data.
[0154] The present invention will be further described below in conjunction with specific examples. It should be understood that the following specific examples are based on the present technical solution and provide detailed implementation methods and specific operating processes, but the scope of protection of the present invention is not limited to these examples.
[0155] Example 1
[0156] 1. Determination of IC50 value of toxicity of selenium compounds to cancer cells based on cancer cell gene expression
[0157] The method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression comprises the following steps:
[0158] (1) The mRNA expression data (gene expression data) of lung cancer cells (A549, NCI-H522), liver cancer cells (HEPG2, HEP38217), colon cancer cells (HCT116), breast cancer cells (MCF-7) and immune system cancer cells (HL-60, THP-1) in the CCLE database are shown in Table 1. The first row represents gene j, j = 1, 2, ..., N, with a total of 16,483 genes. The first column represents different cancer cells and their numbers i in the CCLE database, i = 1, 2, ..., M, with a total of 8 different cancer cell lines.
[0159] Table 1 mRNA expression data of different cancer cells in the CCLE database
[0160]
[0161]
[0162] The FCBF algorithm is used to perform dimensionality reduction on the cancer cell mRNA expression data (gene expression data) in the CCLE database to obtain the reduced dimensionality cancer cell mRNA expression data (gene expression data) set F′. Specifically, the following steps are performed:
[0163] (1.1) The mRNA expression data (gene expression data) of cancer cells in the CCLE database were sorted and expressed in the form of the following matrix:
[0164]
[0165] There are i samples and j genes in total. The i-th row is represented by Si, which represents the sample with cancer cell category Ci, where Ci = {C1, C2, ..., C M The jth column is denoted by Fj, which represents the mRNA expression data (gene expression data) of cancer cells. Then Fij is the mRNA expression data (gene expression data) of the jth gene in the i-th sample, i = 1, ..., M, j = 1, 2, ..., N;
[0166] (1.2) The correlation between the mRNA expression data (gene expression data) Fj of cancer cells in the established matrix and the cancer cell category Ci is measured. The correlation is represented by CSU′i:
[0167]
[0168] (1.3) Arrange in descending order to obtain the new gene expression sequence F′:
[0169] F′=F′-{F j |CSU′ i =0, i=1,…,M}
[0170] Where F′ on the left side of the equation represents the new feature subset;
[0171] (1.4) Delete all features with CSU′i=0 in F′, select a feature from F′ to become the first element in the feature subset, store F1′ in the set, and start from the second feature in the original F′ to search and delete the approximate redundant cover formed by all features of F1′ to obtain a new feature subset:
[0172] F′=F′-{F j |CSU′ i >CSU′ j And CSU′ i ≤ISU′ i ,j,i=1,…,M,j=1,…,N}, where F′ on the left side of the equation represents the new feature subset;
[0173] (1.5) Set F1′ to the next remaining feature in the new feature subset F′ and repeat step (1.4) until the last feature of F′;
[0174] (1.6) Obtain the mRNA expression data (gene expression data) set F′ of cancer cells after feature selection, that is, obtain the mRNA expression data (gene expression data) set F′ of cancer cells after dimensionality reduction.
[0175] Input the mRNA expression data (gene expression data) of different cancer cells in Table 1, and use the FCBF algorithm to perform dimensionality reduction processing on the mRNA expression data (gene expression data) of different cancer cells to obtain the reduced dimensionality mRNA expression data (gene expression data) set F' of different cancer cells (Table 2). Table 2 Use the FCBF algorithm to perform dimensionality reduction processing on the mRNA expression data (gene expression data) of different cancer cells to obtain the reduced dimensionality mRNA expression data (gene expression data) set F' of different cancer cells
[0176]
[0177] (2) The mRNA expression data (gene expression data) set F′ of cancer cells obtained in step (1) after dimensionality reduction is divided into an F′ training sample subset and an F′ test sample subset according to a ratio of 6:2.
[0178] (3) The IC50 values of different selenium compounds against different cancer cells are shown in Table 3. The first row represents four different selenium compounds: selenomethionine (Se-Met), selenocysteine (Se-cys), selenomethylselenocysteine hydrochloride (D-Se-cys) and sodium selenite (Sod-sem). The first column is the name of the different cancer cell lines. The unit of IC50 value is μM.
[0179] Table 3 IC50 values (μM) of different selenium compounds against different cancer cells
[0180] Se-Met(μM) Se-cys(μM) D-Se-cys(μM) Sod-sem(μM) ACH-000002(HL-60) 1708.5 49.2 8013.8 180.5 ACH-000019(MCF-7) 6251.4 890.3 43.7 307.6 ACH-000146(THP-1) 3251.4 98.6 8082 207.4 ACH-000343(NCI-H522) 689.5 3.7 648.9 3.3 ACH-000625(HEP38217) 643.3 63.6 2271 16.6 ACH-000681(A549) 781.7 4.7 1431.4 3.4 ACH-000739(HEPG2) 570 3.2 2100 6.4 ACH-000971(HCT116) 308.2 2.3 2184.1 5.7
[0181] The Gauss-Newton iterative algorithm was used to determine the IC50 values of selenocyanide compounds against different cancer cells using the F′ training sample subset and the IC50 values of selenocyanide compounds against different cancer cells. A determination model for determining the IC50 values of selenocyanide compounds against cancer cells based on the mRNA expression level (gene expression level) of cancer cells was obtained. This was specifically achieved by the following steps:
[0182] (3.1) Assume is the regression coefficient to be estimated b=(b0,b1,...,b M ) T The initial value of , the nonlinear regression model is:
[0183] y i =f(x i ,b)+ε i (i=1, 2, ..., M)
[0184] Where y is the IC50 value of the toxicity of selenium compounds to cancer cells, x is the mRNA expression level (gene expression level) value in the F′ training sample subset, ε is the error, and b is the regression coefficient to be estimated;
[0185] (3.2) f(x i , b) at the initial value g 0 Taylor expansion is done here:
[0186]
[0187] k represents the position distance with regression coefficient;
[0188] (3.3) Substituting the formula obtained in step (3.2) into the nonlinear regression model in step (3.1) yields:
[0189]
[0190] Let Y = y i -f(x i , g 0 ), but
[0191]
[0192] (3.4) is expressed as a matrix The relationship between the mRNA expression level (gene expression level) of cancer cells and the IC50 of selenium compounds' toxicity to cancer cells is: Y = BX + E. Let:
[0193]
[0194] (3.5) The Gauss-Newton iteration method is used to find the least squares fit and estimate the corrected regression coefficient B:
[0195] B=(X T X) -1 X T Y
[0196] Let g 0 is the first iteration value, then g 1 =g 0 +b 0 ;
[0197] (3.6) Set the residual sum of squares to:
[0198]
[0199] Where s is the number of iterations;
[0200] (3.7) Repeat step (3.5) and, under a given tolerance value ε, When the accuracy requirement is met, the iteration is stopped, g s To output the result;
[0201] (3.8) Obtain a determination model for the IC50 value of selenium compound toxicity to cancer cells based on the mRNA expression level (gene expression level) of cancer cells: y = g s X+ε;
[0202] Where y is the IC50 value of the toxicity of selenium compounds to cancer cells, X is the mRNA expression level (gene expression level) of cancer cells, and ε is the error;
[0203] Substitute the cancer cell mRNA expression (gene expression) data of the F′ test sample subset into the measurement model y=g s In X+ε, the IC50 value of the toxicity of selenium compounds to cancer cells is obtained and compared with the actual experimental value to test the reliability of the model;
[0204] Through the measurement model, the IC50 value of the toxicity of selenium compounds to cancer cells can be determined based on the mRNA expression level (gene expression level) of cancer cells.
[0205] 2. Validation of IC50 values of selenium compounds for cancer cell toxicity based on cancer cell gene expression
[0206] From the mRNA expression data (gene expression data) set F′ of cancer cells, the mRNA expression data (gene expression data) of liver cancer cell HEPG2 or immune system cancer cell HL-60 of the F′ training sample subset and the toxicity IC50 value of selenomethionine (Se-Met) on liver cancer cell HEPG2 or immune system cancer cell HL-60 were studied to obtain the fitting determination model y=g s X+ε; Substitute the mRNA expression data (gene expression data) of the liver cancer cell HEPG2 or immune system cancer cell HL-60 from the F′ test sample subset in the cancer cell mRNA expression data (gene expression data) set F′ into the obtained measurement model y=g s In X+ε, the error between the IC50 value calculated by the model and the IC50 value detected by the actual cell experiment is p<0.05, indicating that the method of determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression in Example 1 can accurately determine the IC50 value of the toxicity of selenium compounds to cancer cells, and can replace the actual cell experiment to accurately determine the IC50 value of the toxicity of selenium compounds to cancer cells.
[0207] From the mRNA expression data (gene expression data) set F' of cancer cells, the mRNA expression data (gene expression data) of colon cancer cells HCT116 or breast cancer cells MCF-7 of the F' training sample subset were selected and the toxicity IC50 value of sodium selenite (Sod-sem) on colon cancer cells HCT116 or breast cancer cells MCF-7 was studied to obtain the fitting determination model y=g s X+ε; Substitute the mRNA expression data (gene expression data) of colon cancer cell HCT116 or breast cancer cell MCF-7 of the F′ test sample subset in the cancer cell mRNA expression data (gene expression data) set F′ into the obtained measurement model y=g s In X+ε, the error between the IC50 value calculated by the model and the IC50 value detected by the actual cell experiment is p<0.01, indicating that the method of determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression in Example 1 can accurately determine the IC50 value of the toxicity of selenium compounds to cancer cells, and can replace the actual cell experiment to accurately determine the IC50 value of the toxicity of selenium compounds to cancer cells.
[0208] The above-described embodiments are merely descriptions of preferred embodiments of the present invention. Obviously, the specific implementation of the present invention is not limited to the above-described embodiments. As long as various non-substantial improvements are made using the method concepts and technical solutions of the present invention, or the concepts and technical solutions of the present invention are directly applied to other occasions without improvement, they are all within the scope of protection of the present invention.
Claims
1. A method for determining the IC50 value of toxicity of selenium compounds to cancer cells based on cancer cell gene expression, characterized in that: The method comprises the steps of: (1) The FCBF algorithm is used to reduce the dimensionality of cancer cell gene expression data to obtain the reduced dimensionality cancer cell gene expression dataset F′; (2) dividing the cancer cell gene expression dataset F′ obtained after dimensionality reduction in step (1) into a training sample subset F′ and a test sample subset F′; (3) Gauss-Newton iterative algorithm is performed using the F′ training sample subset and the IC50 values of the toxicity of selenium compounds to different cancer cells to obtain a determination model for the IC50 value of the toxicity of selenium compounds to cancer cells based on the gene expression of cancer cells; through the determination model, the IC50 value of the toxicity of selenium compounds to cancer cells is determined according to the gene expression of cancer cells.
2. The method according to claim 1, wherein The step (1) comprises: (1.1) Organize the cancer cell gene expression data and express it in the form of the following matrix: There are i samples and j genes in total. The i-th row is represented by Si, which represents the sample with cancer cell category Ci, where Ci = {C1, C2, ..., C M The jth column is denoted by Fj, which represents the gene expression data of cancer cells. Then Fij is the expression data of the jth gene in the i-th sample, where i = 1, ..., M, and j = 1, 2, ..., N. (1.2) The correlation between the gene expression data Fj of the cancer cells in the established matrix and the cancer cell category Ci is measured. The correlation is represented by CSU′i: (1.3) Arrange in descending order to obtain the gene expression sequence F′: F′=F′-{F j |CSU′ i =0,i=1,…,M} Among them, F′ on the left side of the equation represents the new feature subset; (1.4) Delete all features with CSU′i=0 in F′, select a feature from F′ to become the first element in the feature subset, store F1′ in the set, and start from the second feature in the original F′ to search and delete the approximate redundant cover formed by all features of F1′ to obtain a new feature subset: F′=F′-{F j |CSU′ i >CSU′ j And CSU′ i ≤ISU′ i,j ,i=1,…,M,j=1,…,N} Among them, F′ on the left side of the equation represents the new feature subset; (1.5) Set F1′ to the next remaining feature in the new feature subset F′ and repeat step (1.4) until the last feature of F′; (1.6) The cancer cell gene expression dataset F′ after feature selection is obtained, that is, the cancer cell gene expression dataset F′ after dimensionality reduction is obtained.
3. The method according to claim 1, wherein In the step (2), the cancer cell gene expression dataset F′ obtained in the step (1) after dimensionality reduction is divided into an F′ training sample subset and an F′ test sample subset according to a ratio of 6:
2.
4. The method according to claim 1, wherein The step (3) comprises: (3.1) Assume is the regression coefficient to be estimated b=(b0,b1,…,b M ) T The initial value of , the nonlinear regression model is: y i =f(x i ,b)+ε i ,(i=1,2,...,M) Where y is the IC50 value of the toxicity of selenium compounds to cancer cells, x is the gene expression value in the F′ training sample subset, ε is the error, and b is the regression coefficient to be estimated; (3.2) f(x i , b) at the initial value g 0 Taylor expansion is done here: k represents the position distance with regression coefficient; (3.3) Substituting the formula obtained in step (3.2) into the nonlinear regression model in step (3.1) yields: make but (3.4) is expressed as a matrix The relationship between the gene expression of cancer cells and the IC50 of selenium compounds on cancer cells is: Y = BX + E, let: (3.5) The Gauss-Newton iteration method is used to find the least squares fit and estimate the corrected regression coefficient B: B=(X T X) -1 X T Y Let g 0 is the first iteration value, then g 1 =g 0 +b 0 ; (3.6) Set the residual sum of squares to: Where s is the number of iterations; (3.7) Repeat step (3.5) and, under a given tolerance value ε, When the accuracy requirement is met, the iteration is stopped, g s To output the result; (3.8) Obtain a model for determining the IC50 value of selenium compounds' toxicity to cancer cells based on cancer cell gene expression: y = g s X+ε, Where y is the IC50 value of the toxicity of selenium compounds to cancer cells, X is the gene expression value of cancer cells, and ε is the error; The IC50 value of the toxicity of selenium compounds to cancer cells is determined based on the gene expression of cancer cells using the assay model.
5. The method according to claim 1, wherein The methods described include in vitro assays and / or auxiliary assays.
6. The method according to claim 1, wherein The cancer cells include one or more of A549 cells, NCI-H522 cells, HEPG2 cells, HEP38217 cells, HCT116 cells, MCF-7 cells, HL-60 cells and THP-1 cells.
7. The method according to claim 1, wherein The selenium compound is selected from the following group of selenium compounds or their analogs: selenomethionine (Se-Met), selenocysteine (Se-cys), selenomethylselenocysteine hydrochloride (D-Se-cys), sodium selenite (Sod-sem), or a combination thereof.
8. The method according to claim 1, wherein The gene expression includes mRNA expression.
9. A device, characterized in that: The device comprises: processor; and A memory storing computer instructions, which, when executed by the processor, causes the processor to execute the method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression as claimed in claim 1.
10. A non-transitory computer storage medium storing a computer program, which, when executed by one or more processors, causes the processors to execute the method for determining the IC50 value of the toxicity of selenium compounds to cancer cells based on cancer cell gene expression as claimed in claim 1.
Citation Information
Patent Citations
Drug IC50 deep learning model prediction method based on molecular structure and gene expression
CN114373550A