Storage medium storing a gene selection method program

By combining gene selection methods with cosine optimization algorithm and other strategies, the problem of removing redundant gene features in microarray gene expression data is solved, achieving more efficient and accurate gene selection, significantly reducing detection costs.

CN117238379BActive Publication Date: 2025-06-27WENZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311331114.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-17
Publication Date
2025-06-27
Estimated Expiration
2040-09-17

AI Technical Summary

Technical Problem

When processing microarray gene expression data, it is difficult to effectively remove redundant gene features, resulting in overfitting of detection algorithms, long training time, and may lead to incorrect detection results, and even delay patient treatment.

Method used

Combining cosine optimization algorithm, sea squirt strategy, moth to flame strategy and reverse learning strategy, a gene selection method is designed to optimize the gene selection process through binary encoding and fitness value updates to reduce the computational burden and noise of irrelevant genes.

Benefits of technology

It significantly reduces detection costs, simplifies gene expression testing, improves the accuracy and efficiency of gene selection, and reduces the possibility of wrong detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117238379B_ABST
    Figure CN117238379B_ABST
Patent Text Reader

Abstract

The present invention provides a storage medium storing a gene selection method program for gene feature selection. It includes obtaining a training set and a test set from a gene data microarray data set and determining an initial population; performing binary encoding on each individual of the current population using a transformation function; calculating the fitness value of the current population and updating relevant parameters in the salp swarm and moth-flame algorithms; setting relevant parameters of the sine-cosine optimization algorithm and updating the population using the sine-cosine optimization algorithm iteration formula; updating the population obtained by the sine-cosine optimization algorithm successively through the salp swarm, moth-flame and opposition-based learning strategies to obtain three populations; selecting the next-generation population through greedy selection; if the maximum number of iterations is reached, end the loop and output the optimal solution, otherwise continue to iterate until the iterative calculation ends. The present invention can more accurately and efficiently screen out the gene features that contribute the most to the category from genes, reducing the detection cost.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application is a divisional application of Chinese Patent No. 202010982171.4 with the invention title of "Gene Selection Method and Device", the full text of which is hereby incorporated by reference in its entirety. Technical Field

[0002] The present invention relates to gene selection technology in the field of data preprocessing, and in particular to a gene selection method. Background Art

[0003] In today's highly competitive world, humans are affected by various diseases, especially cancer, leukemia, etc. Medical tests help to identify the symptoms and causes of various diseases. With the rapid development of technologies related to the biomedical and health fields, a large amount of bioinformatics and clinical medicine data, especially molecular biology experiment data and gene data, has been growing at an unprecedented rate. At present, although humans have studied the occurrence and development processes of various diseases at the molecular level and discovered a large number of disease-causing genes, people have little understanding of their occurrence and regulatory mechanisms. The analysis of microarray gene expression data and protein expression data can be used to grasp physiological activity information at the molecular level.

[0004] However, the number of genes in microarray data is in the thousands, and some of the gene features may be irrelevant to the mining task or there may be redundancy among the features. In fact, only a small number of genes are truly relevant to sample classification. These redundant gene features may lead to overfitting of the detection algorithm's model and excessive training time, resulting in incorrect detection results, and even causing delays and losses of patients' lives.

[0005] In recent years, researchers in related fields have analyzed microarray data and used different types of machine learning algorithms and statistical methods, such as artificial neural networks and evolutionary algorithms, etc., which have been used to analyze gene expression data. However, due to the high dimensionality of gene data and the presence of a large amount of noise, more and more intelligent algorithms have become more important in microarray data analysis. Mining the minimum gene subset greatly reduces the computational burden and "noise" caused by irrelevant genes, and can even extract simple detection rules, enabling accurate detection without the need for any classifier. Moreover, it simplifies the gene expression test, including only a few genes instead of thousands of genes, which can significantly reduce the cost of detection. It requires further study of the possible biological relationships between these few genes and disease development and treatment. Gene selection methods based on particle swarm algorithms and bat algorithms have achieved quite good classification results.

[0006] The Sine Cosine Algorithm (SCA) is a newly emerging heuristic swarm intelligence algorithm that uses two mathematical formulas of sine and cosine functions to continuously explore and develop in the entire search space. However, in the process of gene screening, SCA still has a high room for improvement in terms of the convergence speed and convergence accuracy of the optimal solution. In this case, it is difficult to maintain an effective balance between exploration and exploitation.

[0007] Therefore, it is necessary to provide a gene selection method to achieve more accurate and efficient denoising of gene expression data and reduce the detection cost. Summary of the Invention

[0008] Based on an in-depth study of the characteristics of gene microarray data and aiming at the existing problems, the present invention designs a gene selection method to achieve more accurate and efficient denoising of gene expression data.

[0009] Specifically, according to one aspect of the present invention, an embodiment of the present invention provides a gene selection method, and the method includes the following steps:

[0010] Step S1: Obtain a training set and a test set from a gene data microarray data set, and determine an initial population;

[0011] Step S2: Perform binary encoding on each feature value of each individual in the current population using a transfer function;

[0012] Step S3: Calculate the fitness value of the current population and update the relevant parameters in the salp swarm and moth-flame optimization strategies;

[0013] Step S4: Set the relevant parameters of the sine cosine optimization algorithm and update the population using the sine cosine optimization algorithm iteration formula;

[0014] Step S5: Update the population obtained by the sine cosine algorithm successively through the salp swarm, moth-flame optimization, and opposition-based learning strategies to obtain three populations;

[0015] Step S6: Select the next generation population through greedy selection;

[0016] Step S7: If the maximum number of iterations is reached, end the loop and output the optimal solution; otherwise, continue to iterate until the iterative calculation ends.

[0017] According to another aspect of the present invention, in step S1, based on the training sample set obtained by feature extraction, an initial training sample population is set i = 1, 2,..., N, j = 1, 2,..., D, t = 0, where N is the number of training sample individuals, D is the feature value of each training sample, and X tDenote the population obtained in the \(t\)-th iteration. Denote the \(j\)-th eigenvalue of the \(i\)-th individual in the \(t\)-th iteration, where \(t\) represents the current iteration number and its value range is \([0, 1000]\).

[0018] According to another aspect of the present invention, in step S2, each eigenvalue of each individual in the population \(X\) t is simulated into a binary coding value through formula (1) and formula (2);

[0019]

[0020]

[0021] wherein, denotes the \(j\)-th eigenvalue of the \(i\)-th individual generated in the \(t\)-th iteration, \(r\) is a random number in \([0, 1]\), denotes the \(j\)-th binary coding value of the \(i\)-th individual generated in the \(t\)-th iteration, and Sig both denote the sigmoid function.

[0022] According to another aspect of the present invention, in step S3, formula (3) and formula (4) are used to calculate the fitness value of the population \(X\) t and update the optimal solution used in the salp swarm algorithm and the flame \(F\) involved in the moth-flame optimization algorithm t and the moth \(M\) t , where the flame \(F\) t is the population recombined in ascending order of the fitness values obtained from the above population \(X\), and the moth \(M\) t is \(X\) t ; t ;

[0023]

[0024]

[0025] wherein, Fitness i denotes the fitness value of the \(i\)-th individual, acc i denotes the classification accuracy rate, \(w\) A denotes the classification accuracy weight, \(w\) F denotes the feature selection number weight, \(R\) refers to the number of '1's in each binary individual value, that is, the length of the feature subset of the gene data; \(D\) is the dimension of the individual, that is, the total number of attributes in the gene dataset, \(cc\) represents the number of correctly classified samples in the sample, and \(uc\) represents the number of misclassified samples.

[0026] According to another aspect of the present invention, in the step S4, relevant parameters r1, r2, r3, and r4 of the sine-cosine optimization algorithm are set, and a new population is obtained by updating using formula (5):

[0027]

[0028] wherein, r1 is a linearly decreasing function in [0, 2], r2 is a random number in [0, 2π], and r3 and r4 are random numbers in [0, 1]. represents the j-th eigenvalue of the i-th individual generated in the (t + 1)-th iteration. is the j-th binary coding value of the i-th individual in the t-th iteration generated by formulas (1) and (2). represents the j-th binary coding value of the individual corresponding to the minimum fitness value in the binary coding population obtained by using formulas (3) and (4) in the t-th iteration.

[0029] According to another aspect of the present invention, in the step S5, the populations updated by the sine-cosine optimization algorithm are respectively updated through salp swarm, moth-flame, and opposition-based learning strategies to obtain three populations. The specific steps include:

[0030] First, the salp swarm update strategy transposes the population X t+1 obtained by formula (5), denoted as (X t+1 ). Specifically, when i < N / 2, the first half of the transposed population is updated using formula (6); when i > N / 2 and i < N + 1, the second half of the transposed population is updated using formula (7). Finally, the transposed populations are combined and transposed again to obtain a new population S T , where N is the same as above, and N is the number of training sample individuals. t+1 wherein,

[0031]

[0032]

[0033] where, t and t max are the current iteration number and the maximum iteration number respectively, c2 and c3 are random numbers in [0, 1], represents the j-th binary coding value of the individual corresponding to the minimum fitness value in the binary coding population obtained by using formulas (3) and (4) in the t-th iteration, ub j is the upper bound value of the j-th dimension, lb j is the lower bound value of the j-th dimension, is the transposed value of the i-th individual of the current population X t+1 in the j-th dimension in the (t + 1)-th iteration, Denote the current population X at the (t + 1)-th iteration t+1 The transposed value of the (i - 1)-th individual in the j-th dimension of is the transposed value of the i-th individual in the j-th dimension obtained by using the salp swarm optimization update strategy at the (t + 1)-th iteration;

[0034] Secondly, the moth - flame optimization update strategy adopts the navigation method of moths, taking the flame as the "wind vane" for moths to search in the search space, and updates the current position in a spiral manner. The population M is updated using formulas (8) - (10). t+1 ;

[0035]

[0036]

[0037]

[0038] Among them, is the j-th dimension value of the i-th moth individual at the (t + 1)-th iteration, is the j-th dimension value of the i-th flame individual at the (t + 1)-th iteration, is the distance between the flame and the moth at the (t + 1)-th iteration, b is a constant coefficient, k is a random number from - 1 to 1, n represents the maximum number of flames, t represents the current iteration number, t max represents the maximum number of iterations, l represents the current number of flames, and round represents rounding;

[0039] Finally, the opposition - based learning strategy is an opposition - based solution symmetric to the original solution; the opposition population O of the current population is obtained using formula (11). t+1 ;

[0040]

[0041] Among them, ub j is the upper bound value of the j-th dimension, lb j is the lower bound value of the j-th dimension, is the j-th dimension value of the i-th individual at the (t + 1)-th iteration.

[0042] According to another aspect of the present invention, in step S6, the three populations S t+1 , M t+1 and O t+1 obtained in step S5 are used. The fitness values are calculated according to formulas (3) and (4), sorted from small to large, and the first N individuals with small fitness values are selected as the next - generation population X t+1 , where N is the same as above and is the number of training sample individuals;

[0043] According to another aspect of the present invention, in step S7, if the maximum number of iterations is reached, the loop ends and the optimal solution is output; otherwise, the number of iterations is incremented by 1 and the process returns to step S2.

[0044] An embodiment of the present invention further provides a gene selection device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the foregoing gene selection method are implemented.

[0045] The present invention also provides a method for denoising gene expression data, which is characterized in that the gene expression data irrelevant to sample classification is removed by using the foregoing gene selection method.

[0046] Implementing the embodiments of the present invention has the following beneficial effects:

[0047] In view of the characteristics of gene microarray data, the salp swarm algorithm, moth-flame optimization algorithm and opposition-based learning strategy are combined into the SCA algorithm, which greatly reduces the computational burden and "noise" caused by irrelevant genes, and can even extract simple detection rules. At the same time, the gene expression test is simplified, and the detection cost can be significantly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, obtaining other drawings based on these drawings still belongs to the scope of the present invention.

[0049] Figure 1 It is a flowchart of the gene selection method provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION

[0050] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.

[0051] According to a preferred embodiment of the present invention, as Figure 1 shown, a gene selection (screening) method is provided, and the method includes the following steps:

[0052] Step S1: Obtain a training set and a test set from a gene data microarray dataset, and determine an initial population;

[0053] Step S2: Binary-encode each feature value of each individual in the current population using a transfer function;

[0054] Step S3: Calculate the fitness value of the current population and update the relevant parameters in the salp swarm and moth-flame optimization strategies;

[0055] Step S4: Set the relevant parameters of the sine-cosine optimization algorithm and update the population using the sine-cosine optimization algorithm iteration formula;

[0056] Step S5: Update the population obtained by the sine-cosine algorithm successively through the salp swarm, moth-flame and opposition-based learning strategies to obtain three populations;

[0057] Step S6: Select the next-generation population through greedy selection;

[0058] Step S7: If the maximum number of iterations is reached, end the loop and output the optimal solution; otherwise, continue the iteration until the iteration calculation ends.

[0059] Advantageously, in view of the characteristics of gene data, through the salp swarm strategy, moth-flame strategy and opposition-based learning strategy, combined with the sine-cosine optimization algorithm, the computational burden and "noise" caused by irrelevant genes are greatly reduced, the gene expression test is simplified, and the detection cost can be significantly reduced.

[0060] According to another preferred embodiment of the present invention, as Figure 1 shown, in Embodiment 1 of the present invention, a gene selection method is provided, and the method includes the following steps.

[0061] Currently, microarray data is usually obtained by DNA microarray technology. The analysis of microarray gene expression data and protein expression data can be used to grasp the physiological activity information at the molecular level and has been widely applied in the biomedical field. The number of samples in the microarray dataset is relatively small, and the number of genes is in the thousands. The error estimation is greatly affected by the samples. When the error is not properly estimated, inappropriate applications of classification methods will occur. To overcome this problem, a validation method called K-fold cross-validation is used to estimate the classification error. In the present invention, 10-fold cross-validation is used to verify the classification results when solving the accuracy in the classification process. The dataset is evenly divided into 10 parts, one of which is used as the test set, and the other nine parts are used as the training set. The final result is averaged by looping 10 times. The advantage of using 10-fold cross-validation is that the training set and test set can be fixed and emphasized in each round, and the error can be reduced.

[0062] Step S1: Based on the training sample set extracted from the above gene data microarray dataset, initialize the training sample population i = 1, 2,..., N, j = 1, 2,..., D, t = 0, where N is the number of training sample individuals, D is the dimension number of each training sample, and X t represents the population obtained at the t-th iteration, Denote the j-th eigenvalue of the i-th individual at the t-th iteration, where t represents the current iteration number and its value range is [0, 1000].

[0063] Step S2: Design a K-Nearest Neighbor (KNN) classifier based on the training sample set and perform classification;

[0064] Specifically, for example, design a KNN classifier based on the sample set and perform classification. Simulate each eigenvalue of each individual in the population X t into a binary coded value through Formula (1) and Formula (2);

[0065]

[0066]

[0067] wherein, denotes the j-th eigenvalue of the i-th individual generated in the t-th iteration, r is a random number in [0, 1], denotes the j-th binary coded value of the i-th individual generated in the t-th iteration, and Sig both denote the sigmoid function.

[0068] Step S3: Obtain the fitness values of the individuals in the current population through Formulas (4) and (5), sort them from smallest to largest in terms of fitness value, and update the optimal solution used in the salp swarm algorithm and the flame F t and the moth M t in the moth-flame algorithm, where the flame F t is specifically the population F t recombined from the fitness values of the obtained population X t in ascending order, and the moth M t is X t .

[0069] The KNN classification method determines which class the sample to be tested belongs to based on the distance between the test sample and the training samples. Generally, the K samples closest to the test sample are selected. When K = 1, if the test sample is the closest to a certain neighbor sample, its class is the same as that of this sample; when K ≥ 1, according to the fitness function defined based on the splitting accuracy in the KNN classifier, the test sample belongs to the same class as the majority of the nearest K samples. The steps of the KNN algorithm are as follows:

[0070] First, obtain the distance. When given the test data, calculate its distance from each object in the training data. The distance function determines which samples in the training set are the K neighbors of the test sample. The distance formula used in the present invention is the Euclidean distance, and the specific calculation method is as follows

[0071]

[0072] Among them, test i represents the i-th test vector, and train j represents the j-th training vector, and test i,k represents the k-th dimensional value of the i-th test vector, and train j,k represents the k-th dimensional value of the j-th training vector.

[0073] Secondly, find adjacent objects. According to the distance, find the K training samples with the closest distance as the neighbors of the test sample.

[0074] Finally, determine the category. According to the main categories to which these K neighbors belong, find the category with the largest proportion as the category to which the test sample belongs.

[0075] Gene selection can be regarded as a multi-objective optimization problem, in which two conflicting objectives need to be achieved, namely selecting the smallest number of genes and maximizing the classification accuracy. Therefore, we need to set an objective function to normalize these two objectives into one function. The specific fitness function is as follows:

[0076]

[0077]

[0078] Among them, Fitness i represents the fitness value of the i-th individual, and acc i represents the classification accuracy rate, and w A represents the classification accuracy weight, and w F represents the feature selection number weight. R refers to the number of '1's in each binary individual value, that is, the length of the feature subset of the gene data. D is the dimension of the individual, that is, the total number of attributes in the gene dataset, cc represents the number of correctly classified samples in the sample, and uc represents the number of misclassified samples.

[0079] Step S4: Set the relevant parameters of the sine-cosine optimization algorithm and obtain the population updated by the sine-cosine optimization algorithm;

[0080] Specifically, for example, set the relevant parameters r1, r2, r3, and r4 of the sine-cosine optimization algorithm, and use formula (6) to update to obtain a new population:

[0081]

[0082] Among them, r1 is a linearly decreasing function in [0, 2], r2 is a random number in [0, 2π], and r3 and r4 are random numbers in [0, 1], Denote the j-th eigenvalue of the i-th individual generated in the (t + 1)-th iteration, which is the j-th binary coding value of the i-th individual generated in the t-th iteration by formulas (1) and (2), Denote the j-th binary coding value of the individual corresponding to the minimum fitness value in the binary coding population obtained by using formulas (4) and (5) in the t-th iteration.

[0083] Step S5: Update the populations updated by the sine-cosine optimization algorithm through the salp swarm algorithm, moth-flame optimization algorithm, and opposition-based learning strategy respectively to obtain three populations;

[0084] Specifically, for example, first, the salp swarm algorithm update strategy transposes the population X obtained by formula (6) t+1 , denoted as (X t +1 ), T Specifically: when i < N / 2, update the first half of the transposed population by using formula (7); when i > N / 2 and i < N + 1, update the second half of the transposed population by using formula (8), and finally synthesize the transposed population and then perform transposition to obtain the new population S t+1 , where N is the same as above, and N is the number of training sample individuals;

[0085]

[0086]

[0087] Among them, t and t max are the current iteration number and the maximum iteration number respectively, c2 and c3 are random numbers in [0, 1], denotes the j-th binary coding value of the individual corresponding to the minimum fitness value in the binary coding population obtained by using formulas (4) and (5) in the t-th iteration, ub j is the upper bound value of the j-th dimension, lb j is the lower bound value of the j-th dimension, is the transposed value of the i-th individual of the current population X t+1 in the j-th dimension in the (t + 1)-th iteration, denotes the transposed value of the (i - 1)-th individual of the current population X t+1 in the j-th dimension in the (t + 1)-th iteration, is the transposed value of the i-th individual in the j-th dimension obtained by using the salp swarm algorithm update strategy in the (t + 1)-th iteration;

[0088] Secondly, the moth-flame optimization algorithm update strategy is to adopt the navigation method of moths, use the flame as the "wind vane" for moths to search in the search space, and update the current position in a spiral manner, and update the population M by using formulas (9) - (11)t+1 ;

[0089]

[0090]

[0091]

[0092] Among them, is the j - dimensional value of the i - th moth individual in the (t + 1)-th iteration, is the j - dimensional value of the i - th flame individual in the (t + 1)-th iteration, is the distance between the flame and the moth in the (t + 1)-th iteration, b is a constant coefficient, k is a random number from - 1 to 1, n represents the maximum number of flames, t represents the current iteration number, t max represents the maximum number of iterations, l represents the current number of flames, round represents rounding;

[0093] Finally, the reverse learning strategy is a reverse solution symmetric to the original solution; the reverse population O of the current population is obtained using formula (11) t+1 ;

[0094]

[0095] Among them, ub j is the upper bound value of the j - th dimension, lb j is the lower bound value of the j - th dimension, the j - dimensional value of the i - th individual in the (t + 1)-th iteration.

[0096] Step S6: Screen out the best population through greedy selection;

[0097] Specifically, for example, the three populations S t+1 , M t+1 and O t+1 obtained in step S5 are used to calculate the fitness values according to formulas (4) and (5), sorted from small to large, and the first N individuals with small fitness values are selected as the next - generation population X t+1 , where N is the same as above and is the number of training sample individuals.

[0098] Step S7: If the maximum number of iterations is reached, end the loop and output the optimal solution; otherwise, increment the iteration number by 1 and return to step S2.

[0099] According to a preferred embodiment of the present invention, compared with the gene selection method provided in the first embodiment of the present invention, the second embodiment of the present invention further provides a gene selection device, including a memory and a processor. The memory stores a computer program. Wherein, when the processor executes the computer program, the steps of the gene selection method in the first embodiment of the present invention are implemented. It should be noted that the process of the processor executing the computer program in the second embodiment of the present invention is consistent with the execution process of each step in the gene selection method provided in the first embodiment of the present invention. For specific details, reference can be made to the foregoing related content description.

[0100] According to a preferred embodiment of the present invention, the present invention further provides a method for denoising gene expression data, which is characterized in that the gene selection method described above is used to remove gene expression data irrelevant to sample classification.

[0101] Implementing the embodiments of the present invention has the following beneficial effects:

[0102] Aiming at the characteristics of gene microarray data, the salp swarm algorithm, the moth-flame optimization algorithm and the opposition-based learning strategy are combined into the SCA algorithm, which greatly reduces the computational burden and "noise" caused by irrelevant genes. It can even extract simple detection rules, while simplifying gene expression tests and significantly reducing the cost of detection.

[0103] Those of ordinary skill in the art can understand that all or part of the steps in implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, such as ROM / RAM, disk, optical disc, etc.

[0104] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.

Claims

1. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is adapted to be executed by a processor to implement the steps of a gene selection method, and the gene selection method includes the following steps: Step S1, obtaining a training set and a test set through a gene data microarray data set, and determining an initial population; Step S2, performing binary encoding on each feature value of each individual in the current population by using a transfer function; Step S3, calculating the fitness value of the current population, and updating relevant parameters in the salp swarm algorithm and the moth-flame optimization strategy; Step S4, setting relevant parameters of the sine-cosine optimization algorithm, and updating the population by using the iteration formula of the sine-cosine optimization algorithm; Step S5, sequentially updating the population obtained by the sine-cosine algorithm through the salp swarm algorithm, the moth-flame optimization strategy, and the opposition-based learning strategy to obtain three populations; Step S6, selecting the next-generation population through greedy selection; Step S7, if the maximum number of iterations is reached, ending the loop and outputting the optimal solution, otherwise continuing to iterate until the iterative calculation ends; Among them, in step S2, each eigenvalue of each individual in population X t is simulated into a binary coded value by formulas (1) and (2); Among them, represents the j-th eigenvalue of the i-th individual generated in the t-th iteration, r is a random number in [0, 1], represents the j-th binary coding value of the i-th individual generated in the t-th iteration, and Sig both represent the sigmoid function; In step S3, the population X is calculated using formulas (3) and (4). t The fitness values are calculated, and the optimal solution used in the salp swarm algorithm and the flames F involved in the moth-flame optimization algorithm t and moths M t are updated, where the flames F t are the population obtained by reordering the fitness values of the above population X t in ascending order, and the moths M t are X t . Among them, Fitness i represents the fitness value of the i-th individual, acc i represents the classification accuracy rate, w A represents the classification accuracy weight, w F represents the feature selection number weight. R refers to the number of '1's in each binary individual value, that is, the length of the feature subset of the gene data; D is the dimension of the individual, that is, the total number of attributes in the gene dataset, cc represents the number of correctly classified samples in the sample, and uc represents the number of misclassified samples; In step S4, setting relevant parameters r1, r2, r3, and r4 of the sine-cosine optimization algorithm, and updating to obtain a new population by using formula (5); Among them, r1 is a linearly decreasing function in [0, 2], r2 is a random number in [0, 2π], and r3 and r4 are random numbers in [0, 1]. represents the j-th eigenvalue of the i-th individual generated in the (t + 1)-th iteration is the j-th binary coding value of the i-th individual in the t-th iteration generated by formulas (1) and (2), P j t represents the j-th binary coding value of the individual corresponding to the minimum fitness value in the binary coding population obtained by using formulas (3) and (4) in the t-th iteration.

2. The computer-readable storage medium according to claim 1, wherein The training sample set obtained by feature extraction in step S1, let the initial training sample population where N is the number of training sample individuals, D is the dimensionality of each training sample, and X t represents the population obtained in the t-th iteration, represents the j-th eigenvalue of the i-th individual in the t-th iteration, and t represents the current iteration number, with a value range of [0, 1000].

3. The computer-readable storage medium according to claim 1, wherein In step S5, the steps of updating the population updated by the sine-cosine optimization algorithm through the salp swarm algorithm, the moth-flame optimization strategy, and the opposition-based learning strategy respectively to obtain three populations specifically include: First, the salp swarm update strategy transposes the population X obtained by formula (5), denoted as (X t+1 ), specifically: when i < N / 2, formula (6) is used for update to obtain the first half of the transposed population; when i > N / 2 and i < N + 1, formula (7) is used for update to obtain the second half of the transposed population. Finally, the transposed populations are combined and transposed again to obtain the new population S t+1 ), T , where N is the same as above, and N is the number of training sample individuals; t+1 ​ Among them, t and t max are the current iteration number and the maximum iteration number respectively. c2 and c3 are random numbers in [0, 1]. P j t represents the j-th binary coding value of the individual corresponding to the minimum fitness value in the binary coding population obtained by using formulas (3) and (4) at the t-th iteration. ub j is the upper bound value of the j-th dimension, and lb j is the lower bound value of the j-th dimension. is the transposed value of the i-th individual of the current population X t+1 at the (t + 1)-th iteration in the j-th dimension. represents the transposed value of the (i - 1)-th individual of the current population X t+1 at the (t + 1)-th iteration in the j-th dimension. is the transposed value of the i-th individual obtained by using the salp swarm optimization update strategy at the (t + 1)-th iteration in the j-th dimension. Secondly, the moth-to-flame update strategy adopts the navigation method of moths, takes the flame as the "wind vane" for moths to search in the search space, updates the current position in a spiral manner, and updates the population M using formulas (8) to (10). t+1 ; Among them, is the j - dimensional value of the i - th moth individual at the (t + 1)-th iteration, is the j - dimensional value of the i - th flame individual at the (t + 1)-th iteration, is the distance between the flame and the moth at the (t + 1)-th iteration, b is a constant coefficient, k is a random number from - 1 to 1, n represents the maximum number of flames, t represents the current iteration number, t max represents the maximum number of iterations, l represents the current number of flames, and round represents rounding; Finally, the reverse learning strategy is a reverse solution symmetric to the original solution; the reverse population O of the current population is obtained using Equation (11). t+1 ; where, ub j is the upper bound value of the j-th dimension, lb j is the lower bound value of the j-th dimension, and is the value of the j-th dimension of the i-th individual at the (t + 1)-th iteration.

4. The computer-readable storage medium according to claim 3, wherein Take the three populations S t+1 , M t+1 and O t+1 obtained in step S5, calculate the fitness values according to formulas (3) and (4), sort them from small to large, and select the first N individuals with small fitness values as the next-generation population X t+1 , where N is the same as above and is the number of training sample individuals.

5. The computer-readable storage medium according to claim 4, wherein If the maximum number of iterations is reached, ending the loop and outputting the optimal solution, otherwise adding 1 to the number of iterations, and returning to step S2.

6. A denoising device for gene expression data, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the gene selection method according to any one of claims 1 to 5 is used to remove gene expression data irrelevant to sample classification.

Citation Information

Patent Citations

  • Spark distribution-based feature selection method of a parallel binary system moth fire fighting algorithm

    CN109871934A

  • Method for constructing prediction model based on improved sine and cosine algorithm

    CN111079074A