Gene data classification method and device based on fuzzy rough set and incremental learning
By constructing the optimal feature subset of gene expression data based on fuzzy rough set and incremental learning, the classification model is trained, and the problems of low accuracy and poor stability of gene data classification in traditional methods are solved, achieving more efficient and accurate gene data classification.
Patent Information
- Application Number
- CN202411531163.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-30
AI Technical Summary
Traditional gene expression data classification methods are prone to low classification accuracy and poor stability due to the complex structure, high dimensions, small sample size, fast update speed, many noise interferences, and many redundant attributes.
A gene data classification method based on fuzzy rough set and incremental learning is adopted to train the sample data set by obtaining gene expression data, construct an optimal feature subset, and use this subset to train the gene data classification model until the difference between the predicted and the real classification results converges.
While ensuring the modeling speed of gene classification models, it improves the accuracy of gene data classification, reduces the model's dependence on redundant features, and improves classification efficiency and accuracy.
Smart Images

Figure CN119541657B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of electronic digital data processing, and in particular to a gene data classification method and device based on fuzzy rough sets and incremental learning. Background Art
[0002] Genetic data classification refers to the process of grouping and categorizing genetic data according to certain standards or rules. Genetic data classification not only helps humans better understand the structure and function of genetic data, but also provides important data support for scientific research, disease diagnosis, drug development, etc.
[0003] As a potential application scenario, gene data classification can obtain the genomic data of tumor patients' samples, and then identify specific cancer subtypes based on the gene expression profiles in the genomic data. For example, breast cancer can be divided into three subtypes based on gene expression profiles: Luminal A, Luminal B, HER2+ and triple-negative breast cancer. Using gene data classification, the specific breast cancer subtype of the patient can be determined based on the different gene expression profiles of different patients.
[0004] However, due to the characteristics of gene expression data such as complex structure, high dimensionality, small number of effective samples, fast update speed, much noise interference, and many redundant attributes, traditional gene expression data classification methods are prone to low classification accuracy and poor stability. Summary of the invention
[0005] In view of this, the embodiments of the present application provide a gene data classification method and device based on fuzzy rough sets and incremental learning, so as to improve the accuracy of gene data classification while ensuring the modeling speed of the gene classification model.
[0006] In a first aspect, an embodiment of the present application provides a gene data classification method based on fuzzy rough sets and incremental learning, wherein the method comprises:
[0007] Acquiring gene expression data to be classified, and inputting the gene expression data to be classified into a target gene data classification model;
[0008] The target gene data classification model outputs a target classification result based on the gene expression data to be classified;
[0009] The target gene data classification model is pre-trained through the following steps:
[0010] Acquire a gene expression data training sample data set, wherein the gene expression data training sample data set includes a plurality of gene expression data;
[0011] Based on the training sample data set, construct an optimal feature subset of the gene expression data training sample data set; wherein the optimal feature subset includes a plurality of high-quality gene expression feature vectors, each of which is a gene expression feature vector whose importance is greater than a preset importance threshold value and is determined by using a preset fuzzy rough set model, wherein the importance is used to describe the contribution of the gene expression data to the gene classification result;
[0012] Inputting the optimal feature subset into a pre-constructed gene data classification model, and training the gene data classification model until the target data difference between the predicted gene classification result output by the gene data classification model and the true classification result converges; wherein the predicted gene classification result is the predicted classification result output by the gene data classification model for each gene expression feature vector in the feature subset, and the true classification result is the true classification result of each gene expression feature vector in the feature subset;
[0013] The gene data classification model when the target data difference converges is determined as the target gene data classification model.
[0014] Optionally, in some possible embodiments, the step of obtaining a gene expression data training sample dataset includes:
[0015] Acquire newly added gene expression data as an incremental training sample data set, and determine the gene expression data that has been used for training the gene data classification model as a stock training sample data set;
[0016] Using a K-means algorithm to calculate the cosine similarity between the gene expression data in the incremental training sample data set and the gene expression data in the stock training sample data set, and adding mask information to the corresponding loss function according to the calculated similarity;
[0017] Constructing a preset total cross entropy loss function based on the loss function after adding mask information, and training the gene data classification model until the preset total cross entropy loss function converges;
[0018] Among them, the preset total cross entropy loss function is constructed based on the following formula:
[0019] L total =L ce +λ·L kd +γ·L pro
[0020] Among them, L total is the preset total cross loss function, L ceis: the cross loss function between the predicted classification result and the actual classification result output by the gene data classification model based on the stock training sample data set, L pro is the cross loss function between the predicted classification result and the actual classification result output by the gene data classification model based on the incremental training sample data set, L kd is the distillation loss function, λ is the distillation loss function L kd The weight of γ is the cross loss function L between the predicted classification result and the actual classification result output by the gene data classification model based on the incremental training sample data set. pro The weight of
[0021] Among them, L kd is the distillation loss function, where the distillation loss function L kd Satisfies the following formula:
[0022] L kd =||F t -F t-1 ||
[0023] Among them, F t is the number of features of gene expression data selected during the training of the gene data classification model in the current round, F t-1 It is the number of features of the gene expression data selected during the previous round of training of the gene data classification model.
[0024] Optionally, in some possible embodiments, the optimal feature subset is obtained by:
[0025] Based on each of the gene expression data, a preset subset space is constructed, and a fuzzy strategy is calculated based on the preset subset space using a Gaussian membership function;
[0026] Based on the fuzzy strategy, combining the fuzzy attributes and the fuzzy decision attributes of each of the gene expression data, determining an average value of the mutual information of the fuzzy attributes and the fuzzy decision attributes;
[0027] From the gene expression data training sample data set, gene expression feature vectors whose average value of the mutual information meets a preset importance screening condition are selected to form the optimal feature subset.
[0028] Optionally, in some possible embodiments,
[0029] The method of constructing a preset subset space based on each of the gene expression data, and calculating a fuzzy strategy based on the preset subset space using a Gaussian membership function, comprises:
[0030] Constructing Fuzzy Decision System FDS<U,C,D> , where U is a non-empty priority object set, C is a fuzzy attribute of the gene expression data, and D is a fuzzy decision attribute of the gene expression data;
[0031] And the relative fuzzy similarity relationship between the gene expression data sample x and the gene expression data sample y in the preset subset space B is calculated according to the following formula:
[0032]
[0033] Among them, d R (x, y) refers to the relative distance between gene expression data sample x and gene expression data sample y, δ refers to the parameter of Gaussian membership function, which controls the preset subset space The relative fuzzy neighborhood particle size between the gene expression data is determined based on the following formula: Among them, the relatively fuzzy neighborhood particle Used to specify the feature refinement between the gene expression data;
[0034]
[0035] Among them, ε is the fuzzy neighborhood parameter;
[0036] The fuzzy decision of the gene expression data is determined based on the following formula
[0037]
[0038] Among them, D i It refers to the i-th fuzzy decision equivalence class in the gene expression data sample set.
[0039] Optionally, in some possible embodiments,
[0040] The average value of the mutual information between the fuzzy attribute and the fuzzy decision attribute includes: the average value of the fuzzy neighborhood mutual information of the simulated decision about the preset subset space I R (D; B), the average value of the fuzzy neighborhood relative dependence mutual information RDI (D; B) of the fuzzy decision on the preset subset space; the fuzzy strategy D is combined with the fuzzy attributes and fuzzy decision attributes of each of the gene expression data to determine the average value of the mutual information of the fuzzy attributes and the fuzzy decision attributes, including:
[0041] Based on the relative fuzzy neighborhood particles and the fuzzy decision, the fuzzy neighborhood relative upper approximation of the simulated decision about the preset subset space B is calculated according to the following formula: The simulation decision is about the relative lower approximation of the simulation neighborhood of the preset subset space B
[0042]
[0043] Based on the following formula, the fuzzy neighborhood relative positive domain POS of decision D about the preset subset space B is calculated: B (D) Fuzzy neighborhood relative dependence
[0044]
[0045] Wherein, |U| is the number of gene expression data in the gene expression data training sample data set; D j (x i ) is for the i-th gene expression data sample x i The jth fuzzy decision of
[0046] For each gene expression feature vector in the gene expression data training sample data set, the average value of the fuzzy neighborhood mutual information of the simulation decision about the preset subset space is calculated based on the following formula: R (D; B):
[0047]
[0048] in, is the number of samples of the relative fuzzy neighborhood particles in the gene expression data training sample dataset, |D(x i )| is the number of samples in the fuzzy decision equivalence class;
[0049] And based on the following formula, the average value of the fuzzy neighborhood relative dependence mutual information RDI (D; B) of the fuzzy decision on the preset subset space is calculated:
[0050]
[0051] Optionally, in some possible embodiments, selecting the gene expression feature vectors whose average value of the mutual information meets a preset importance screening condition from the gene expression data training sample data set to form the optimal feature subset includes:
[0052] According to the following formula, the feature importance Sig(b,B,D) of each gene expression feature vector is calculated:
[0053] Sig(b,B,D)=RDI(D;B∪{b})-RDI(D;B)
[0054] Among them, b is the remaining gene expression feature vector that is not selected into the preset subset space. If the feature importance Sig(b, B, D) of the remaining gene expression feature vector is greater than the preset importance threshold, the mutual information of the remaining gene expression feature vector meets the preset importance screening condition.
[0055] In a second aspect, an embodiment of the present application provides a gene data classification device based on fuzzy rough sets and incremental learning, wherein the device comprises:
[0056] A data acquisition module, used to acquire gene expression data to be classified, and input the gene expression data to be classified into a target gene data classification model;
[0057] A classification module, configured to output a target classification result based on the gene expression data to be classified by the target gene data classification model;
[0058] Wherein, the target gene data classification model is pre-trained by the following steps:
[0059] Acquire a gene expression data training sample data set, wherein the gene expression data training sample data set includes a plurality of gene expression data;
[0060] Based on the training sample data set, construct an optimal feature subset of the gene expression data training sample data set; wherein the optimal feature subset includes a plurality of high-quality gene expression feature vectors, each of which is a gene expression feature vector whose importance is greater than a preset importance threshold value and is determined by using a preset fuzzy rough set model, wherein the importance is used to describe the contribution of the gene expression data to the gene classification result;
[0061] Inputting the optimal feature subset into a pre-constructed gene data classification model, and training the gene data classification model until the target data difference between the predicted gene classification result output by the gene data classification model and the true classification result converges; wherein the predicted gene classification result is the predicted classification result output by the gene data classification model for each gene expression feature vector in the feature subset, and the true classification result is the true classification result of each gene expression feature vector in the feature subset;
[0062] The gene data classification model when the target data difference converges is determined as the target gene data classification model.
[0063] In a third aspect, an embodiment of the present application provides an electronic device, wherein the electronic device includes:
[0064] Processor; and
[0065] Memory for storing programs,
[0066] The program includes instructions, which, when executed by the processor, cause the processor to execute the genetic data classification method based on fuzzy rough sets and incremental learning described in the first aspect.
[0067] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the genetic data classification method based on fuzzy rough sets and incremental learning described in the first aspect.
[0068] Beneficial effects of this application:
[0069] The embodiment of the present application provides a gene data classification method and device based on fuzzy rough sets and incremental learning, wherein the gene data classification method inputs the acquired gene expression data to be classified into a target gene data classification model, and the target gene data classification model outputs a target classification result based on the input gene expression data to be classified. Since the target gene data classification model is obtained in advance by: acquiring a gene expression data training sample data set, and then determining a gene expression feature vector with an importance greater than a preset importance threshold with the help of a preset model rough set model, using the gene expression feature vector with an importance greater than the preset importance threshold, training the pre-constructed gene data classification model, until the target data difference between the predicted gene classification result output by the gene data classification model and the actual classification result converges, and then determining the gene data classification model when the target data difference converges as the target gene data classification model, etc.
[0070] In the process, by adopting the preset fuzzy rough set model, the gene expression data feature vectors in the gene expression data training sample data set can be screened, and only the gene expression feature vectors with importance greater than the preset importance threshold are retained to participate in the training of the gene data classification model, which can effectively reduce the interference of gene expression data with poor contribution to the gene classification results on the model training. It is achieved that the data dimension that the gene data classification model needs to process is reduced while retaining important classification information, which helps to improve the classification efficiency and accuracy of the gene data classification model. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Further details, features and advantages of the present application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0072] Figure 1 A schematic diagram of a process of a gene data classification method based on fuzzy rough sets and incremental learning provided in an embodiment of the present application is shown;
[0073] Figure 2A schematic diagram of a process for training a gene expression data classification model provided in an embodiment of the present application is shown;
[0074] Figure 3 Another schematic diagram of a process for training a gene expression data classification model provided in an embodiment of the present application is shown;
[0075] Figure 4 A schematic diagram of the structure of a gene data classification device based on fuzzy rough sets and incremental learning provided in an embodiment of the present application is shown;
[0076] Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement an embodiment of the present application is shown. DETAILED DESCRIPTION
[0077] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not intended to limit the scope of protection of the present application.
[0078] It should be understood that the various steps described in the method implementation of the present application can be performed in different orders and / or performed in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.
[0079] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0080] It should be noted that the modifications of "one" and "plurality" mentioned in the present application are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0081] In the field of biological sciences, genes are the basic units responsible for carrying genetic information in organisms. They are all the nucleotide sequences required to produce a polypeptide chain or functional RNA. Genes control the life activities and genetic shape of organisms by expressing corresponding functional products, such as RNA or protein. Gene expression data contains information about gene activity and can reflect the current physiological state of cells, such as whether cells are in a normal state or a deteriorated state, whether drugs are effective against tumor cells, etc. Therefore, in the field of biological sciences, gene function and gene expression regulation information can be obtained by analyzing gene expression data.
[0082] With the development of DNA microarray technology, a huge amount of gene expression data has been generated. As described in the previous paragraph, gene expression data contains rich information about gene activity. Analyzing the patterns hidden in gene expression data is of great significance for understanding and inferring biological gene functions and studying gene regulation mechanisms. Therefore, gene expression data classification (hereinafter referred to as gene expression data) is the focus and hot spot of current research in the field of biological sciences. Because gene data has the characteristics of complex structure, high dimensionality, few samples and fast update, high noise, and many redundant attributes. Therefore, the use of traditional data analysis methods may face problems such as long time consumption and low classification accuracy.
[0083] In view of this, the present application provides a genetic data classification method, device and electronic device based on fuzzy rough sets and incremental learning, which is committed to improving the accuracy of genetic data classification. In the first aspect, the present application provides a genetic data classification method based on fuzzy rough sets and incremental learning, which is applied to any electronic device with genetic data classification function, including but not limited to personal mobile terminals, computers or servers, etc. Figure 1 As shown, the method comprises the following steps:
[0084] S11, obtaining gene expression data to be classified, and inputting the gene expression data to be classified into a target gene data classification model;
[0085] S12. The target gene data classification model outputs a target classification result based on the gene expression data to be classified.
[0086] In some embodiments, Figure 2 As shown, the target gene data classification model is pre-trained through the following steps:
[0087] S21, obtaining a gene expression data training sample data set, wherein the gene expression data training sample data set includes a plurality of gene expression data;
[0088] S22, based on the training sample data set, constructing an optimal feature subset of the gene expression data training sample data set; wherein the optimal feature subset includes a plurality of high-quality gene expression feature vectors, each of which is a gene expression feature vector whose importance is greater than a preset importance threshold value and is determined by using a preset fuzzy rough set model, wherein the importance is used to describe the contribution of the gene expression data to the gene classification result;
[0089] S23, inputting the optimal feature subset into a pre-constructed gene data classification model, and training the gene data classification model until the target data difference between the predicted gene classification result output by the gene data classification model and the true classification result converges; wherein the predicted gene classification result is the predicted classification result output by the gene data classification model for each gene expression feature vector in the feature subset, and the true classification result is the true classification result of each gene expression feature vector in the feature subset;
[0090] S24, determining the gene data classification model when the target data difference converges as the target gene data classification model.
[0091] The gene data classification method inputs the acquired gene expression data to be classified into a target gene data classification model, and the target gene data classification model outputs a target classification result based on the input gene expression data to be classified. Since the target gene data classification model is obtained in advance by: obtaining a gene expression data training sample data set, and then using a preset model rough set model to determine a gene expression feature vector with an importance greater than a preset importance threshold, using the gene expression feature vector with an importance greater than the preset importance threshold, training the pre-constructed gene data classification model until the target data difference between the predicted gene classification result output by the gene data classification model and the actual classification result converges, and then determining the gene data classification model when the target data difference converges as the target gene data classification model, and the like.
[0092] During the training of the gene data classification model, by adopting a preset fuzzy rough set model, the gene expression data feature vectors in the gene expression data training sample data set can be screened, and only the gene expression feature vectors with an importance greater than a preset importance threshold are retained to participate in the training of the gene data classification model, which can effectively reduce the interference of gene expression data with a poor contribution to the gene classification results on the model training. The use of the embodiment of the present application realizes the reduction of the data dimension required to be processed by the gene data classification model while retaining important classification information, which helps to improve the classification efficiency and accuracy of the gene data classification model.
[0093] The following will describe the above steps S11-S12 and S21-S24 in detail with reference to specific implementation examples:
[0094] In an embodiment of the present application, the gene expression data to be classified obtained in step S11 may be existing gene expression data, but the specific type corresponding to the gene expression data is temporarily unknown. Among them, the gene expression data may be data that has been collected and stored in a specified database based on existing gene expression data collection means, or it may be data temporarily collected in real time based on existing gene expression data. Exemplarily, a certain gene expression data is gene expression data just collected from a patient, and it is necessary to determine the specific subtype of the patient's illness with the aid of the target gene data classification model provided in an embodiment of the present application, then the gene expression data just collected can be determined as gene expression data to be classified. The state of specific gene expression data can be flexibly changed according to the actual state, and this application is not strictly limited.
[0095] In an embodiment of the present application, during the execution of step S11, a pre-trained target gene data classification model can be called, and the acquired gene expression data to be identified is used as input and passed into the target gene data classification model. The target gene data classification model performs gene classification based on each data feature of the gene expression data to be classified, finds out the classification results associated with each data feature, and finally determines the classification result with a correlation degree greater than a preset correlation degree threshold as the target classification result.
[0096] Among them, what type of gene classification the target gene data classification model can perform depends on what kind of training samples are used for training during the training phase. Exemplarily, the target gene data classification model aims to determine which breast cancer subtype the breast cancer patient suffers from based on the gene expression data of the breast cancer patient. In the model training phase, the training sample data taken is the gene expression data of the known breast cancer subtype. By using supervised training, it is possible to accurately extract data features related to breast cancer subtypes in the gene expression data, and determine which breast cancer subtype the patient belongs to based on the degree of correlation between the data features and the breast cancer subtype.
[0097] In some embodiments, the above-mentioned gene expression data to be classified can be obtained by DNA microarray technology, or each gene expression data in the gene expression data training sample data set can be obtained. Among them, the working principle of DNA microarray technology is based on the principle of base complementarity, that is, A is paired with T, and C is paired with G. By fixing a large number of DNA or RNA probes on a solid surface, such as a glass sheet or a silicon wafer, these probes can be hybridized with the mRNA or cDNA in the labeled sample. Then, a laser scanner or a fluorescence microscope is used to detect the fluorescent signal after hybridization, so as to obtain the expression data of each gene. In this way, DNA microarray technology can also convert the fluorescent signal corresponding to each gene expression data obtained into a corresponding electrical signal, so as to obtain the electrical signal sequence corresponding to each gene expression data. As an embodiment, the electrical signal corresponding to the gene expression data can be understood as the data sequence corresponding to each gene expression data.
[0098] On this basis, when executing step S21, the first gene expression feature vector can be obtained by preprocessing the data sequence corresponding to the gene expression data. As an implementation method, the preprocessing types include: data filtering, data normalization and other processing. Specifically, in order to overcome the influence of different data dimensions on the processing results, thereby ensuring the validity of the analysis results, in an embodiment of the present application, when executing step S21, the gene expression data can be subjected to maximum and minimum normalization processing, and the gene expression data can be processed into a data distribution space of [0, 1]. Exemplarily, the following maximum and minimum normalization (or simply referred to as normalization processing) processing formula can be used to process the gene expression data into a data distribution space of [0, 1]:
[0099]
[0100] Among them, x i is the actual observed value, which can be understood as a gene feature actually obtained by DNA microarray technology. minx is the minimum value of all observed values, that is, the minimum value of the current batch of gene expression data obtained by DNA array technology. maxx is the maximum value of all observed values, that is, the maximum value of the current batch of gene expression data obtained by DNA array technology.
[0101] In the embodiment of the present application, each feature vector obtained after the normalization process of a gene feature actually obtained by the DNA microarray technology is the first gene expression feature vector corresponding to each gene expression data. On this basis, step S22 is executed to construct a feature subset of the gene expression data training sample data set. The purpose of executing step S22 is to screen the gene expression data in the training sample data set and obtain gene expression data that is helpful for the training of the gene data classification model.
[0102] In the embodiment of the present application, gene expression data is specifically screened by a preset fuzzy rough set model. The preset fuzzy rough set selection model is specifically a mathematical algorithm model for data processing, which selects data features by using fuzzy rough set theory, thereby obtaining data features that are important and contribute to model training from a large number of data features, and then integrating the obtained data features that are important and contribute to model training into a feature set, which is the optimal feature subset.
[0103] Specifically, in some embodiments, the mathematical algorithm flow corresponding to the preset fuzzy rough set selection model can be as follows:
[0104] 1) Based on each gene expression data, a preset subset space is constructed. The preset subset space can be regarded as the beginning of a storage space, which is specifically used to integrate the gene expression data after screening.
[0105] 2) Based on the preset subset space, the fuzzy strategy D is calculated using the Gaussian membership function.
[0106] 3) Based on the fuzzy strategy D, the fuzzy attributes and fuzzy decision attributes of each gene expression data are combined to determine the average value of the mutual information of the fuzzy attributes and the fuzzy decision attributes.
[0107] 4) Selecting, from the gene expression data training sample data set, the gene expression feature vectors whose average value of the mutual information satisfies the preset importance screening condition, and constructing the optimal feature subset based on the selected feature vectors that satisfy the preset importance screening condition.
[0108] In the present application embodiment, the specific expression of gene expression data can be a data matrix, wherein different rows in the matrix represent different experimental conditions or sample data, and columns represent different genes. In this way, each element in the matrix represents the expression level of a specific gene under specific experimental conditions. As another embodiment, the specific expression of gene expression data can be an image form, which is specifically visualized in the form of a heat map, and the color depth represents the height of gene expression level. As another embodiment, the expression of gene expression data can be quantitative data, and the numerical value of gene expression data level is directly given in the quantitative data, and different numerical values come from: after reverse transcription and fluorescent labeling of mRNA samples, hybridization with probes on the chip and detection of fluorescence intensity are obtained.
[0109] In the embodiment of the present application, the preset subset space is constructed in order to construct a public gene expression data space. The preset subset space is a public gene expression data space set based on the actual gene classification requirements. For example, taking the above-mentioned breast cancer gene classification as an example, since there are many types of cancers that can be involved in gene expression data, how to screen breast cancer gene expression data from many cancers, a public gene expression data space can be constructed based on the gene expression data characteristics of breast cancer genes, and various gene attributes of gene expression data are limited for the public gene expression data space, wherein various gene attributes are specifically determined based on the gene attributes of breast cancer.
[0110] In an embodiment of the present application, the model attribute and the fuzzy decision attribute of the gene expression data are two attribute values of the gene expression data, wherein the fuzzy attribute refers to that the description of the gene expression level is not an absolute "yes" or "no", but has a certain ambiguity. For example, suppose that the expression level of a gene in cancer cells is 200 units and the expression level in normal cells is 50 units. The expression level here is not an absolute "high" or "low", but a range. At this time, a fuzzy set can be defined, such as "high expression", and its membership function can be described as the degree of gene expression level above 150 units.
[0111] Fuzzy decision attributes refer to attributes used for classification or decision-making in gene expression data analysis, and these attributes are also fuzzy. For example, in cancer classification, it is necessary to determine whether a cell sample belongs to a certain type of cancer based on gene expression data. The decision attribute "cancer type" here may be fuzzy. For example, the gene expression pattern of a cell sample may have the characteristics of two cancer types at the same time, so its membership to "cancer type A" and "cancer type B" may not be 100%, but 70% and 30% respectively.
[0112] Based on this, the fuzzy strategy D is calculated using the Gaussian membership function, and the Gaussian distribution (or normal distribution) corresponding to the gene attribute of the breast cancer gene can be used to describe the degree to which the specific gene expression data belongs to the public gene expression data space. Simple understanding is to judge the degree of association between each gene expression data and the public gene expression data space constructed based on the gene attribute of the breast cancer gene. If it is strongly associated, it indicates that the gene expression data has a high probability of carrying a breast cancer gene feature fragment, and if it is weakly associated, it indicates that the probability of carrying a breast cancer gene feature fragment in the gene expression data is relatively small. Because the Gaussian membership function has continuity and smoothness, it can make the continuous features described in the gene expression data extremely effective.
[0113] As an implementation mode, the above-mentioned construction of a preset subset space based on each gene expression data and calculation of a fuzzy strategy based on the preset subspace using a Gaussian membership function can be specifically implemented through the following steps:
[0114] Step 1: Construct fuzzy decision system FDS<U,C,D> , where U is a non-empty priority object set, C is a fuzzy attribute of the gene expression data, and D is a fuzzy decision attribute of the gene expression data. The fuzzy decision system FDS (Fuzzy Decision System) specifically simulates human thinking and decision-making processes, and uses the following formula to describe and process the uncertainty and fuzziness in the gene expression data, thereby converting the input gene expression data into a fuzzy set and providing fuzzy information for subsequent gene data classification.
[0115] Step 2: Calculate the relative fuzzy similarity between the gene expression data sample x and the gene expression data sample y in the preset subset space B according to the following formula:
[0116]
[0117] Among them, d R (x, y) refers to the relative distance between gene expression data sample x and gene expression data sample y, δ refers to the parameter of Gaussian membership function, It is used to overcome the defect of traditional fuzzy rough approximation that is sensitive to the existing data distribution. The value of is the relative fuzzy approximation induced by the preset subset space B on U, where,
[0118] Step 3: Control the preset subset space The relative fuzzy neighborhood particle size between the gene expression data is determined based on the following formula: Among them, the relatively fuzzy neighborhood particle Used to specify the feature refinement between the gene expression data;
[0119]
[0120] Among them, ε is the fuzzy neighborhood parameter. Among them, the fuzzy neighborhood parameter can more effectively find attributes that do not need to make fuzzy decisions from the perspective of relevance through non-additive tests and nonlinear operators, thereby helping to improve the efficiency of attribute reduction.
[0121] Step 4: Assume that the gene expression data sample set is divided into r fuzzy decision equivalence classes, that is, U / D = {D1, D2, K, D r},in The fuzzy decision can be calculated based on the following formula
[0122]
[0123] Among them, fuzzy decision Specifically expressed: Sample x is equivalent to fuzzy decision class D i The membership degree, D i It refers to the i-th fuzzy decision equivalence class in the gene expression data sample set. Fuzzy decision equivalence class refers to the process of classifying gene samples with similar membership into one class according to fuzzy decision attributes in the gene expression data set.
[0124] By selecting the embodiments of the present application, based on the neighborhood information in the gene expression data, the local relationship between genes can be better understood when selecting the gene expression feature vector, which helps to more comprehensively and accurately understand the data features existing in the gene expression data, thereby helping the subsequent gene data classification model to better classify genes based on the data features.
[0125] In the embodiment of the present application, the average value of the mutual information of the fuzzy attribute and the fuzzy decision attribute includes: the average value of the fuzzy neighborhood mutual information I of the simulated decision about the preset subset space B R (D; B), the average value of the fuzzy neighborhood relative dependence mutual information RDI (D; B) of the fuzzy decision on the preset subset space B; Based on this, the above-mentioned fuzzy strategy D, combined with the fuzzy attributes and fuzzy decision attributes of each gene expression data, determines the average value of the mutual information of the fuzzy attributes and the fuzzy decision attributes, including:
[0126] Based on the relative fuzzy neighborhood particles and the fuzzy decision, the fuzzy neighborhood relative upper approximation of the simulated decision about the preset subset space B is calculated according to the following formula: The simulation decision is about the relative lower approximation of the simulation neighborhood of the preset subset space B
[0127]
[0128] Based on the following formula, the fuzzy neighborhood relative positive domain POS of decision D about the preset subset space B is calculated: B (D) Fuzzy neighborhood relative dependence
[0129]
[0130] Wherein, |U| is the number of gene expression data in the gene expression data training sample data set; D j (x i ) is for the i-th gene expression data sample x iThe jth fuzzy decision of .
[0131] For each gene expression feature vector in the gene expression data training sample data set, the average value of the fuzzy neighborhood mutual information of the simulation decision about the preset subset space is calculated based on the following formula: R (D; B):
[0132]
[0133] in, is the number of samples of the relative fuzzy neighborhood particles in the gene expression data training sample dataset, |D(x i )| is the number of samples in the fuzzy decision equivalence class;
[0134] And based on the following formula, the average value of the fuzzy neighborhood relative dependence mutual information RDI (D; B) of the fuzzy decision on the preset subset space is calculated:
[0135]
[0136] By selecting the embodiment of the present application, by calculating the average value of the relative dependence mutual information between the gene expression data, the key characteristic genes in the gene expression data can be identified. Based on the key characteristic genes, other redundant characteristic genes can be removed, thereby reducing the dimension of the gene expression feature vector and reducing the amount of data required for gene data classification. It helps to improve the classification efficiency of gene data classification while ensuring the classification accuracy of gene data classification.
[0137] On this basis, in some embodiments, the gene expression feature vectors whose average value of the mutual information satisfies a preset importance screening condition are selected from the gene expression data training sample data set to form the optimal feature subset, including:
[0138] According to the following formula, the feature importance Sig(b,B,D) of each gene expression feature vector is calculated:
[0139] Sig(b,B,D)=RDI(D;B∪{b})-RDI(D;B)
[0140] Among them, b is the remaining gene expression feature vector that is not selected into the preset subset space. If the feature importance Sig(b, B, D) of the remaining gene expression feature vector is greater than the preset importance threshold, the mutual information of the remaining gene expression feature vector meets the preset importance screening condition.
[0141] As an implementation method, the preset importance threshold can be flexibly set according to actual experience. As a preferred embodiment, the preset importance threshold can be 0, that is, when no gene expression feature vector satisfies Sig(b,B,D)>0, or all feature vectors have been selected, then the feature vectors selected into the preset subset space B constitute the optimal feature subset.
[0142] By selecting an embodiment of the present application, the relative dependency mutual information between the gene expression data is calculated, and then the relative dependency mutual information is used to determine whether each gene expression data belongs to the preset subset space B, thereby deleting redundant feature vectors that have a low contribution to model training, reducing the dimension of the data that needs to be processed by the gene data classification model, and improving the processing efficiency of the gene data classification model.
[0143] In an embodiment of the present application, by adopting the above-mentioned fuzzy rough set feature selection algorithm model, feature screening can be performed on a large amount of complex, high-dimensional gene expression data, and feature vectors whose importance is greater than a preset importance threshold are screened out, that is, feature vectors with a higher contribution to model training are screened out, thereby reducing the impact of redundant feature vectors on the data processing accuracy of the gene data classification model.
[0144] On this basis, the selected optimal feature subset is input into a pre-constructed gene data classification model. In some embodiments, the gene data classification model is pre-constructed in the following manner:
[0145] A random forest classifier model is constructed based on the bagging method and the decision tree as the base learner. Then, the Bootstrap resampling technique is used to repeatedly extract N samples from the feature subset with replacement to form a new training sample data set, and the decision tree is trained. Among them, the number of training sample data in the new training sample data set should be 2 / 3 of the number of training sample data in the original gene expression data training sample data set. Then, the same sample data extraction method is used to repeatedly extract M new training sample data sets.
[0146] Then, a subtree of a decision tree is generated using each new training sample data set, and the F in the gene expression data is selected during the generation of each subtree of the random forest model. t As an implementation method, the description attribute information of gene expression data includes: gene sequence, genome location, gene expression data and protein interaction data.
[0147] Among them, the gene sequence is the DNA sequence or RNA sequence of the gene, which is used to analyze the structure and function of the gene. The genomic position refers to the specific location of the gene in the genome, which helps to study the regulatory mechanism of the gene. Gene expression data refers to the expression level of the gene in different tissues, cell types or physiological states. Protein interaction data refers to the interaction relationship between proteins, which is mainly used to study protein function and specific signal transduction processes.
[0148] In the process of subtree generation, the descriptive information with the largest information gain rate is found as the classification attribute, and each node is split until the sample data in all leaf nodes belong to the same category. At this time, a decision tree is generated, and the generated decision trees constitute a decision tree set.
[0149] Then, during the execution of step S23 or step S12, each decision tree in the generated decision tree set is used to predict the input gene expression data, and the mode of the prediction results is output as the predicted gene classification result of the gene expression data.
[0150] In the whole process of generating decision tree, the number of descriptive attributes F is selected t is random, and the corresponding expression is:
[0151]
[0152] Among them, X i is the gene expression data, specifically the feature vector in the optimal feature subset. L is the number of descriptive attributes of the training sample data. rand(a,b) means randomly generating a number in the interval (a,b). p The value range of
[0153] On the basis of the random forest model constructed as above, the data input into the random forest model is divided into stock gene expression data and newly added incremental gene expression data. With the expansion of gene data sample data and the growth of data volume, the classification model suffers from "catastrophic forgetting" and increased training time, both of which limit the performance of the gene data classification model and the accuracy of classification. Based on this, in the process of executing step S21, stock gene expression data and incremental gene expression data can be obtained respectively. Specifically, as an implementation method, the descriptive attribute information of the gene expression data is the characteristics of each gene on the gene expression data. Step S21 can be implemented by the following steps:
[0154] Acquire newly added gene expression data as an incremental training sample data set, and determine the gene expression data that has been used for training the gene data classification model as a stock training sample data set;
[0155] Using a K-means algorithm to calculate the cosine similarity between the gene expression data in the incremental training sample data set and the gene expression data in the stock training sample data set, and adding mask information to the corresponding loss function according to the calculated similarity;
[0156] Constructing a preset total cross entropy loss function based on the loss function after adding mask information, and training the gene data classification model until the preset total cross entropy loss function converges;
[0157] Among them, the preset total cross entropy loss function is constructed based on the following formula:
[0158] L total =L ce +λ·L kd +γ·L pro
[0159] Among them, L total is the preset total cross loss function, L ce is: the cross loss function between the predicted classification result and the actual classification result output by the gene data classification model based on the stock training sample data set, L pro is the cross loss function between the predicted classification result and the actual classification result output by the gene data classification model based on the incremental training sample data set, L kd is the distillation loss function, λ is the distillation loss function L kd The weight of γ is the cross loss function L between the predicted classification result and the actual classification result output by the gene data classification model based on the incremental training sample data set. pro The weight of
[0160] Among them, the distillation loss function L kd Satisfies the following formula:
[0161] L kd =||F t -F t-1 ||
[0162] Among them, F t is the number of features of gene expression data selected during the training of the gene data classification model in the current round, F t-1 It is the number of features of the gene expression data selected during the previous round of training of the gene data classification model.
[0163] Specifically, assume that the dataset of the i-1th classification task is The number of categories of gene features contained in this dataset is C i-1For each stock gene expression data ο, assuming that the number of samples of the stock gene expression data is N o , then the feature vector subset corresponding to the stock gene expression data can be recorded as: in, And 1≤j≤N o .in, refers to the parameters of the feature extractor.
[0164] Then, the K-means algorithm is used to divide the feature vector subset R into K representative gene expression data feature vector clusters S = {S1, S2, K, S k}, the square error can be minimized by the following formula:
[0165]
[0166] Among them, μ i For cluster S i The mean vector of is calculated as follows:
[0167]
[0168] Thus, the feature vector set corresponding to the stock gene expression data ο is P = {μ1, μ2, ..., μ k}, where the value of K determines the number of prototype sampled gene expression data corresponding to each stock gene expression data. The prototype sample is input into the current random forest classifier and then recorded as:
[0169]
[0170] Then, based on the predicted classification results output by the random forest classifier and the true label of the feature vector, the following cross entropy loss function is constructed:
[0171]
[0172] In order to maximize the use of the selected stock gene expression data and avoid confusion between the features of the incremental gene expression data and the stock gene expression data, Figure 3 As shown, in the stage of processing the incremental gene expression data, for the gene expression feature extracted by the current feature extractor, the prototype gene expression data corresponding to the feature is L2 normalized, and the cosine similarity between the stock gene expression data and the incremental gene expression data is calculated according to the following formula:
[0173]
[0174] Among them, for the i-th incremental gene expression data, the feature representation learned by the model based on the incremental gene expression data is recorded as For the i-th stock gene expression data, the features learned by the model based on the incremental gene expression data are recorded as || || represents the modulo operation.
[0175] At this time, a cosine similarity threshold σ is set. If the cosine similarity calculated by the above formula is greater than σ, it can be considered that the incremental gene expression data has extremely similar stock gene expression data. At this time, a mask is added to the distillation loss function. Specifically, a mask is added to the distillation loss function value to further enhance the distillation effect.
[0176] If the calculated cosine similarity is less than σ, the effect of increasing the cross entropy is achieved by adding a mask to the worse loss function, specifically adding a mask to the cross entropy loss function value. Based on this, in the process of training the gene data classification model based on the incremental gene expression data in the incremental training sample data set, the target data difference between the predicted classification result output by the gene data classification model and the true label can be determined by the following preset total loss function:
[0177] L total =L ce +λ·L kd +γ·L pro
[0178] Where λ is the distillation loss function L kd The weight of γ is the cross entropy loss function L between the predicted classification results and the actual classification results output by the gene data classification model based on the incremental training sample data set. pro The weight of .
[0179] In this way, when the incremental gene expression data is input into the gene data classification model and the gene data classification model is trained, the parameters of the gene data classification model corresponding to the minimum value of the preset total cross entropy loss function are continuously sought, and the parameters of the gene data classification model are optimized, so that the gene data classification model has the ability to classify new gene expression data on the basis of retaining the performance obtained by training the stock gene expression data. In this way, when new gene expression data is input into the gene data classification model, the parameters of the model can be updated and optimized based on the processing flow of the incremental training sample data set, so as to continuously improve the data processing accuracy of the gene data classification model with the help of new gene expression data.
[0180] In a second aspect, the present application embodiment provides a gene data classification device based on fuzzy rough sets and incremental learning, wherein: Figure 4 As shown, the device 40 includes:
[0181] The data acquisition module 401 is used to acquire gene expression data to be classified, and input the gene expression data to be classified into a target gene data classification model;
[0182] A classification module 402, configured to output a target classification result based on the gene expression data to be classified by the target gene data classification model;
[0183] Wherein, the target gene data classification model is pre-trained by the following steps:
[0184] Acquire a gene expression data training sample data set, wherein the gene expression data training sample data set includes a plurality of gene expression data;
[0185] Based on the training sample data set, construct an optimal feature subset of the gene expression data training sample data set; wherein the optimal feature subset includes a plurality of high-quality gene expression feature vectors, each of which is a gene expression feature vector whose importance is greater than a preset importance threshold value and is determined by using a preset fuzzy rough set model, wherein the importance is used to describe the contribution of the gene expression data to the gene classification result;
[0186] Inputting the optimal feature subset into a pre-constructed gene data classification model, and training the gene data classification model until the target data difference between the predicted gene classification result output by the gene data classification model and the true classification result converges; wherein the predicted gene classification result is the predicted classification result output by the gene data classification model for each gene expression feature vector in the feature subset, and the true classification result is the true classification result of each gene expression feature vector in the feature subset;
[0187] The gene data classification model when the target data difference converges is determined as the target gene data classification model.
[0188] Among them, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with the relevant laws and regulations and do not violate public order and good morals.
[0189] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0190] In a third aspect, the exemplary embodiments of the present application further provide an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to enable the electronic device to perform a method according to an embodiment of the present application when executed by the at least one processor.
[0191] The exemplary embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present application.
[0192] The exemplary embodiments of the present application further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to enable the computer to execute the method according to the embodiment of the present application.
[0193] refer to Figure 5 , the structural block diagram of the electronic device 500 that can be used as the server or client of the present application will now be described, which is an example of hardware devices that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples, and are not intended to limit the implementation of the present application described and / or required herein.
[0194] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0195] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 may be any type of device capable of inputting information to the electronic device 500, and the input unit 506 may receive input digital or character information, and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 507 may be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 508 may include, but is not limited to, a disk, an optical disk. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0196] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the aforementioned gene data classification method based on fuzzy rough sets and incremental learning may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. In some embodiments, the computing unit 501 may be configured to perform the aforementioned gene data classification method based on fuzzy rough sets and incremental learning by any other appropriate means (e.g., by means of firmware).
[0197] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0198] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0199] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0200] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0201] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0202] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.
Claims
1. A genetic data classification method based on fuzzy rough sets and incremental learning, characterized in that: The method comprises: Acquiring gene expression data to be classified, and inputting the gene expression data to be classified into a target gene data classification model; The target gene data classification model outputs a target classification result based on the gene expression data to be classified; The target gene data classification model is pre-trained through the following steps: Acquire a gene expression data training sample data set, wherein the gene expression data training sample data set includes a plurality of gene expression data; Based on the training sample data set, construct an optimal feature subset of the gene expression data training sample data set; wherein the optimal feature subset includes a plurality of high-quality gene expression feature vectors, each of which is a gene expression feature vector whose importance is greater than a preset importance threshold value and is determined by using a preset fuzzy rough set model, wherein the importance is used to describe the contribution of the gene expression data to the gene classification result; Inputting the optimal feature subset into a pre-constructed gene data classification model, and training the gene data classification model until the target data difference between the predicted gene classification result output by the gene data classification model and the true classification result converges; wherein the predicted gene classification result is the predicted classification result output by the gene data classification model for each gene expression feature vector in the feature subset, and the true classification result is the true classification result of each gene expression feature vector in the feature subset; Determining the gene data classification model when the target data difference converges as the target gene data classification model; The step of obtaining a gene expression data training sample dataset comprises: Acquire newly added gene expression data as an incremental training sample data set, and determine the gene expression data that has been used for training the gene data classification model as a stock training sample data set; Using a K-means algorithm to calculate the cosine similarity between the gene expression data in the incremental training sample data set and the gene expression data in the stock training sample data set, and adding mask information to the corresponding loss function according to the calculated similarity; Constructing a preset total cross entropy loss function based on the loss function after adding mask information, and training the gene data classification model until the preset total cross entropy loss function converges; Among them, the preset total cross entropy loss function is constructed based on the following formula: L total =L ce +λ·L kd +γ·L pro Among them, L total is the preset total cross loss function, L ce is: the cross loss function between the predicted classification result and the actual classification result output by the gene data classification model based on the stock training sample data set, L pro is the cross loss function between the predicted classification result and the actual classification result output by the gene data classification model based on the incremental training sample data set, L kd is the distillation loss function, λ is the distillation loss function L kd The weight of γ is the cross loss function L between the predicted classification result and the actual classification result output by the gene data classification model based on the incremental training sample data set. pro The weight of Among them, L kd is the distillation loss function, where the distillation loss function L kd Satisfies the following formula: L kd =||F t -F t-1 || Among them, F t is the number of features of gene expression data selected during the training of the gene data classification model in the current round, F t-1 The number of features of the gene expression data selected during the previous round of gene data classification model training; The optimal feature subset is obtained in the following way: Based on each of the gene expression data, a preset subset space is constructed, and a fuzzy strategy is calculated based on the preset subset space using a Gaussian membership function; Based on the fuzzy strategy, combining the fuzzy attributes and the fuzzy decision attributes of each of the gene expression data, determining an average value of the mutual information of the fuzzy attributes and the fuzzy decision attributes; From the gene expression data training sample data set, gene expression feature vectors whose average value of the mutual information meets a preset importance screening condition are selected to form the optimal feature subset.
2. The method according to claim 1, characterized in that The method of constructing a preset subset space based on each of the gene expression data, and calculating a fuzzy strategy based on the preset subset space using a Gaussian membership function, comprises: Constructing Fuzzy Decision System FDS<U,C,D> , where U is a non-empty priority object set, C is a fuzzy attribute of the gene expression data, and D is a fuzzy decision attribute of the gene expression data; And the relative fuzzy similarity relationship between the gene expression data sample x and the gene expression data sample y in the preset subset space B is calculated according to the following formula: Among them, d R (x, y) refers to the relative distance between gene expression data sample x and gene expression data sample y, δ refers to the parameter of Gaussian membership function, which controls the preset subset space The relative fuzzy neighborhood particle size between the gene expression data is determined based on the following formula: Among them, the relatively fuzzy neighborhood particle Used to specify the feature refinement between the gene expression data; Among them, ε is the fuzzy neighborhood parameter; The fuzzy decision of the gene expression data is determined based on the following formula Among them, D i It refers to the i-th fuzzy decision equivalence class in the gene expression data sample set, and r is the number of fuzzy decision equivalence classes.
3. The method according to claim 2, characterized in that The average value of the mutual information between the fuzzy attribute and the fuzzy decision attribute includes: the average value of the fuzzy neighborhood mutual information of the simulated decision about the preset subset space I R (D; B), the average value of the fuzzy neighborhood relative dependence mutual information RDI (D; B) of the fuzzy decision on the preset subset space; the fuzzy strategy D is combined with the fuzzy attributes and fuzzy decision attributes of each of the gene expression data to determine the average value of the mutual information of the fuzzy attributes and the fuzzy decision attributes, including: Based on the relative fuzzy neighborhood particles and the fuzzy decision, the fuzzy neighborhood relative upper approximation of the simulated decision about the preset subset space B is calculated according to the following formula: The simulation decision is about the relative lower approximation of the simulation neighborhood of the preset subset space B Based on the following formula, the fuzzy neighborhood relative positive domain POS of decision D about the preset subset space B is calculated: B (D) Fuzzy neighborhood relative dependence Wherein, |U| is the number of gene expression data in the gene expression data training sample data set; D j (x i ) is for the i-th gene expression data sample x i The jth fuzzy decision of For each gene expression feature vector in the gene expression data training sample data set, the average value of the fuzzy neighborhood mutual information of the simulation decision about the preset subset space is calculated based on the following formula: R (D; B): in, is the number of samples of the relative fuzzy neighborhood particles in the gene expression data training sample dataset, |D(x u )| is the number of samples in the fuzzy decision equivalence class; And based on the following formula, the average value of the fuzzy neighborhood relative dependence mutual information RDI (D; B) of the fuzzy decision on the preset subset space is calculated:
4. The method according to claim 3, characterized in that The step of selecting the gene expression feature vectors whose average value of the mutual information satisfies a preset importance screening condition from the gene expression data training sample data set to form the optimal feature subset comprises: According to the following formula, the feature importance Sig(b,B,D) of each gene expression feature vector is calculated: Sig(b,B,D)=RDI(D;B∪{b})-RDI(D;B) Among them, b is the remaining gene expression feature vector that is not selected into the preset subset space. If the feature importance Sig(b, B, D) of the remaining gene expression feature vector is greater than the preset importance threshold, the mutual information of the remaining gene expression feature vector meets the preset importance screening condition.
5. A gene data classification device based on fuzzy rough sets and incremental learning, characterized in that: The device comprises: A data acquisition module, used to acquire gene expression data to be classified, and input the gene expression data to be classified into a target gene data classification model; A classification module, configured to output a target classification result based on the gene expression data to be classified by the target gene data classification model; Wherein, the target gene data classification model is pre-trained by the following steps: Acquire a gene expression data training sample data set, wherein the gene expression data training sample data set includes a plurality of gene expression data; Based on the training sample data set, construct an optimal feature subset of the gene expression data training sample data set; wherein the optimal feature subset includes a plurality of high-quality gene expression feature vectors, each of which is a gene expression feature vector whose importance is greater than a preset importance threshold value and is determined by using a preset fuzzy rough set model, wherein the importance is used to describe the contribution of the gene expression data to the gene classification result; Inputting the optimal feature subset into a pre-constructed gene data classification model, and training the gene data classification model until the target data difference between the predicted gene classification result output by the gene data classification model and the true classification result converges; wherein the predicted gene classification result is the predicted classification result output by the gene data classification model for each gene expression feature vector in the feature subset, and the true classification result is the true classification result of each gene expression feature vector in the feature subset; Determining the gene data classification model when the target data difference converges as the target gene data classification model; The step of obtaining a gene expression data training sample dataset comprises: Acquire newly added gene expression data as an incremental training sample data set, and determine the gene expression data that has been used for training the gene data classification model as a stock training sample data set; Using a K-means algorithm to calculate the cosine similarity between the gene expression data in the incremental training sample data set and the gene expression data in the stock training sample data set, and adding mask information to the corresponding loss function according to the calculated similarity; Constructing a preset total cross entropy loss function based on the loss function after adding mask information, and training the gene data classification model until the preset total cross entropy loss function converges; Among them, the preset total cross entropy loss function is constructed based on the following formula: L total =L ce +λ·L kd +γ·L pro Among them, L total is the preset total cross loss function, L ce is: the cross loss function between the predicted classification result and the actual classification result output by the gene data classification model based on the stock training sample data set, L pro is the cross loss function between the predicted classification result and the actual classification result output by the gene data classification model based on the incremental training sample data set, L kd is the distillation loss function, λ is the distillation loss function L kd The weight of γ is the cross loss function L between the predicted classification result and the actual classification result output by the gene data classification model based on the incremental training sample data set. pro The weight of Among them, L kd is the distillation loss function, where the distillation loss function L kd Satisfies the following formula: L kd =||F t -F t-1 || Among them, F t is the number of features of gene expression data selected during the training of the gene data classification model in the current round, F t-1 The number of features of the gene expression data selected during the previous round of gene data classification model training; The optimal feature subset is obtained in the following way: Based on each of the gene expression data, a preset subset space is constructed, and a fuzzy strategy is calculated based on the preset subset space using a Gaussian membership function; Based on the fuzzy strategy, combining the fuzzy attributes and the fuzzy decision attributes of each of the gene expression data, determining an average value of the mutual information of the fuzzy attributes and the fuzzy decision attributes; From the gene expression data training sample data set, gene expression feature vectors whose average value of the mutual information meets a preset importance screening condition are selected to form the optimal feature subset.
6. An electronic device, characterized in that: The electronic device comprises: Processor; and Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Gene classification method and system based on clustering and random forest algorithms
CN108846259A
Brain metastasis tumor prognostic index reduction and classification method based on rough set optimization
CN111582370A