Data feature selection method and system based on improved cross breeding algorithm
By segmenting the feature space of high-dimensional data into subspaces using an improved hybridization breeding algorithm, and optimizing rice populations using a ternary tournament strategy and simulated annealing mechanism, the problems of local optima and low computational efficiency in high-dimensional data are solved, achieving efficient feature selection and stable feature subset output.
Patent Information
- Application Number
- CN202510810551.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies are prone to getting stuck in local optima, have low computational efficiency, and lack population diversity in high-dimensional data environments, which affects the application effect of algorithms on large-scale, high-dimensional datasets.
An improved hybridization breeding algorithm is adopted to divide the data feature space into subspaces. The rice population is divided into maintainer lines, restorer lines and sterile lines using a ternary tournament strategy. Individuals are optimized through hybridization, self-pollination, mutation and simulated annealing mechanisms. Local search is performed in combination with simulated annealing mechanism to avoid local optima.
It improves search efficiency, avoids local optima, enhances the accuracy and stability of feature selection, is suitable for rapid dimensionality reduction operations on large-scale high-dimensional data, and enhances the algorithm's adaptability and search efficiency in complex problems.
Smart Images

Figure CN120873533A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of feature selection technology, and in particular to a data feature selection method and system based on an improved hybridization breeding algorithm. Background Technology
[0002] In today's era of industrial big data, high-dimensional features of data have become a common phenomenon, especially in machine learning and data analysis, where industrial big data often contains tens of thousands of feature variables. The feature selection problem for high-dimensional data refers to selecting features from a large pool of data that are significant for solving the problem. This process is a crucial step in data preprocessing, effectively improving model accuracy, reducing computational overhead, accelerating training, and enhancing the model's generalization ability. However, as the number of features increases, the difficulty of feature selection also rises significantly, and traditional feature selection methods are gradually revealing their limitations in high-dimensional data environments. Specifically, with the increase in feature dimensionality, existing algorithms are prone to getting trapped in local optima, experiencing low computational efficiency, and insufficient population diversity. These problems severely affect the application performance of algorithms on large-scale, high-dimensional datasets. Therefore, how to design a feature selection method that can effectively overcome these problems, improve search efficiency, and avoid local optima remains a critical issue that urgently needs to be addressed. Summary of the Invention
[0003] This invention provides a data feature selection method and system based on an improved hybridization breeding algorithm, which solves the technical problems of easy getting trapped in local optima, low computational efficiency, and insufficient population diversity in the prior art, and achieves the technical effect of improving search efficiency and avoiding local optima.
[0004] This invention provides a data feature selection method based on an improved hybridization breeding algorithm, comprising:
[0005] The feature space of the dataset to be selected is divided into several subspaces;
[0006] The rice population was divided into subspaces, and a three-way tournament strategy was used to classify the rice population individuals in each subspace into maintainer lines, restorer lines and sterile lines.
[0007] The expression for hybridization of maintainer lines and sterile lines is: ;in, For the new individuals obtained through hybridization, For the maintainer individuals in the current iteration, For sterile individuals in the current iteration, and A random number between 0 and 1;
[0008] The expression for self-crossing between restorer line individuals is: ;in, For new individuals obtained through self-fertilization, As the current globally optimal individual, Individuals randomly selected from the restorer line. For the restorer individuals in the current iteration, It is a random number between 0 and 1;
[0009] During the global search phase, sterile individuals with fitness values less than a preset value or maintainer individuals with fitness values greater than a threshold are subjected to mutation operations, the expression of which is: ;in, These are new individuals from the sterile line or new individuals from the maintainer line after mutation. This refers to sterile individuals with a current fitness value less than a preset value or maintainer individuals with a fitness value greater than a threshold. For variable asynchronous length, , For fitness guiding coefficients, The fitness value is the fitness value of sterile individuals whose current fitness value is less than the preset value or the fitness value of maintainer individuals whose fitness value is greater than the threshold. It is a random number between 0 and 1;
[0010] During the local search phase, the three lineages are optimized and updated; the expression is as follows: ;in, To optimize the updated individual, These are individuals from three lineages within the current population. For learning rate, The gradient of the current best individual. To simulate the random perturbation term of annealing, , For one in Random numbers within the interval For temperature parameters, This represents the current iteration number; expressed by the formula. The probability of an individual being accepted is calculated. ;in, and These are the fitness values of the new individual and the current individual, respectively. A random number between 0 and 1; if If the random number is greater than 0 and between 1, then the current individual will be... Update to a new individual ;if If the random number is less than or equal to 0 and between 1, then the current individual will be retained. ;
[0011] By combining the updated individuals from each subpopulation, the globally optimal individual is obtained. ;
[0012] Through formula The probability of a feature being selected is calculated. ;in, Standard deviation Index of features The mean; if A random number greater than 0 and 1 represents the globally optimal individual. medium-sized individuals Features If selected, the optimal feature subset is obtained when the iteration ends.
[0013] Specifically, dividing the feature space of the dataset to be selected into several subspaces includes:
[0014] Through formula Calculated features Relevance with class label C ;in, Features Mutual information with class label C, Features Information entropy The information entropy of class label C;
[0015] The features are sorted in descending order according to their correlation SU values to obtain the feature space. ;
[0016] Through formula The number of features in the eigenspace is calculated. ;in, The total number of features in the dataset. The number of subspaces;
[0017] Select the feature space Ranked 1st to Feature composition feature subspace Ranked arrive Feature composition feature subspace This process continues until all features have been assigned to their respective subspaces.
[0018] Specifically, the process of dividing the rice population into subspaces and using a ternary tournament strategy to classify rice population individuals in each subspace into maintainer lines, restorer lines, and sterile lines includes:
[0019] Through formula The characteristic subspace is calculated. Importance ;in, Features Relevance to class label C;
[0020] Through formula The initial number of individuals in the subpopulation was calculated. ;in, The total number of individuals in the population. For characteristic subspace The importance of;
[0021] For the division Each feature subspace is randomly initialized. Individual rice plants were used as the original population;
[0022] Three individuals are randomly selected from the original rice population to calculate their fitness values. The individual with the highest fitness value is assigned to the maintainer line, the individual with the lowest fitness value is assigned to the sterile line, and the last individual is assigned to the restorer line. This process is repeated until all individuals in the population have been assigned, resulting in a three-line rice population: maintainer line... Recovery system and sterile line .
[0023] Specifically, it also includes:
[0024] For restorer individuals that have reached the self-pollination limit, a reset update is performed, expressed as follows: ;in, To reset the updated individual, For the currently reset Restoration class individual, and These represent the upper and lower bounds of the search space, respectively. It is a random number between 0 and 1.
[0025] Specifically, it also includes:
[0026] Through formula The temperature parameters after iterative updates were calculated. ;in, Let be the attenuation factor, and take... , The temperature parameters are as shown before the update.
[0027] This invention also provides a data feature selection system based on an improved hybridization breeding algorithm, comprising:
[0028] The feature space segmentation module is used to segment the feature space of the dataset to be selected into several subspaces;
[0029] The rice population segmentation module is used to divide the rice population into various subspaces. The three-way tournament strategy is used to divide the rice population individuals in each subspace into maintainer lines, restorer lines and sterile lines.
[0030] The hybridization module is used to perform hybridization operations on maintainer lines and sterile lines; its expression is: ;in, For the new individuals obtained through hybridization, For the maintainer individuals in the current iteration, For sterile individuals in the current iteration, and A random number between 0 and 1;
[0031] The self-crossing module is used to perform self-crossing operations between restorer line individuals; its expression is: ;in, For new individuals obtained through self-fertilization, As the current globally optimal individual, Individuals randomly selected from the restorer line. For the restorer individuals in the current iteration, It is a random number between 0 and 1;
[0032] The mutation module is used during the global search phase to perform mutation operations on sterile individuals with fitness values less than a preset value or maintainer individuals with fitness values greater than a threshold. Its expression is: ;in, These are new individuals from the sterile line or new individuals from the maintainer line after mutation. This refers to sterile individuals with a current fitness value less than a preset value or maintainer individuals with a fitness value greater than a threshold. For variable asynchronous length, , For fitness guiding coefficients, The fitness value is the fitness value of sterile individuals whose current fitness value is less than the preset value or the fitness value of maintainer individuals whose fitness value is greater than the threshold. It is a random number between 0 and 1;
[0033] The optimization and update module is used to optimize and update individuals from the three lineages during the local search phase. Its expression is: ;in, To optimize the updated individual, These are individuals from three lineages within the current population. For learning rate, The gradient of the current best individual. To simulate the random perturbation term of annealing, , For one in Random numbers within the interval For temperature parameters, This represents the current iteration number; expressed by the formula. The probability of an individual being accepted is calculated. ;in, and These are the fitness values of the new individual and the current individual, respectively. A random number between 0 and 1; if If the random number is greater than 0 and between 1, then the current individual will be... Update to a new individual ;if If the random number is less than or equal to 0 and between 1, then the current individual will be retained. ;
[0034] The global optimal individual acquisition module combines the updated individuals from each subpopulation to obtain the globally optimal individual. ;
[0035] The optimal feature subset acquisition module is used to obtain the optimal feature subset through the formula. The probability of a feature being selected is calculated. ;in, Standard deviation Index of features The mean; if A random number greater than 0 and 1 represents the globally optimal individual. medium-sized individuals Features If selected, the optimal feature subset is obtained when the iteration ends.
[0036] Specifically, the feature space segmentation module includes:
[0037] Correlation calculation unit, used to calculate using formulas Calculated features Relevance with class label C ;in, Features Mutual information with class label C, Features Information entropy The information entropy of class label C;
[0038] The feature sorting unit is used to sort the features in descending order according to the correlation SU value to obtain the feature space. ;
[0039] The characteristic number calculation unit is used to calculate the characteristic number using the formula The number of features in the eigenspace is calculated. ;in, The total number of features in the dataset. The number of subspaces;
[0040] Feature partitioning unit, used to select the feature space Ranked 1st to Feature composition feature subspace Ranked arrive Feature composition feature subspace This process continues until all features have been assigned to their respective subspaces.
[0041] Specifically, the rice population segmentation module includes:
[0042] Importance calculation unit, used to calculate importance using formulas The characteristic subspace is calculated. Importance ;in, Features Relevance to class label C;
[0043] The initial population size calculation unit is used to calculate the initial population size using the formula. The initial number of individuals in the subpopulation was calculated. ;in, The total number of individuals in the population. For characteristic subspace The importance of;
[0044] The original population initialization unit is used for partitioning. Each feature subspace is randomly initialized. Individual rice plants were used as the original population;
[0045] The rice population partitioning unit is used to randomly select three individuals from the original rice population, calculate their fitness values, assign the individual with the highest fitness value to the maintainer line, the individual with the lowest fitness value to the sterile line, and the last individual to the restorer line. This operation is repeated until all individuals in the population have been partitioned, resulting in a three-line rice population: maintainer line... Recovery system and sterile line .
[0046] Specifically, it also includes:
[0047] The reset / update module is used to reset and update restorer individuals that have reached the self-crossing limit. Its expression is: ;in, To reset the updated individual, For the currently reset Restoration class individual, and These represent the upper and lower bounds of the search space, respectively. It is a random number between 0 and 1.
[0048] Specifically, it also includes:
[0049] Temperature parameter calculation module, used to calculate parameters using formulas The temperature parameters after iterative updates were calculated. ;in, Let be the attenuation factor, and take... , The temperature parameters are as shown before the update.
[0050] One or more technical solutions provided in this invention have at least the following technical effects or advantages:
[0051] First, high-dimensional big data is acquired from industrial IoT devices to obtain a dataset. Then, a feature space segmentation method based on feature importance is used to divide the feature space of the dataset into several subspaces. A subpopulation initialization method based on feature space importance is then used to divide the rice population into these subspaces. An improved hybridization breeding algorithm is then used for feature selection within each subspace, and the optimal feature subset is output. Specifically, by employing a subspace partitioning strategy based on symmetric uncertainty, the importance of features in a subspace can be quickly assessed in a relatively simple and efficient way, avoiding unnecessary computational overhead. Especially when dealing with large-scale high-dimensional data, dimensionality reduction can be completed in a short time, significantly improving processing efficiency and scalability. This makes it more suitable for applications with high real-time requirements or large data volumes. Furthermore, in the early stages of iteration, i.e., during the global search phase of the population, the mutation amplitude is adjusted based on the fitness value of individuals, enhancing population diversity and improving global search capabilities. This mechanism ensures that the population can both broadly explore the solution space and fully utilize existing excellent solutions, avoiding the risk of premature convergence and improving the algorithm's adaptability and search efficiency in complex problems. In the later stages of iteration, specifically during the local search phase of the population, the solutions for individual populations are further refined by strengthening local optimization. By incorporating simulated annealing, the over-reliance on gradient information in the local search is prevented, thereby enhancing the global nature of the local search and ensuring that it is less prone to getting trapped in local optima. This locally enhanced search strategy effectively improves the accuracy and stability of feature selection, ensuring that the final output feature subset can better solve practical problems. Attached Figure Description
[0052] Figure 1 A flowchart illustrating a data feature selection method based on an improved hybridization breeding algorithm provided in an embodiment of the present invention. Detailed Implementation
[0053] This invention provides a data feature selection method and system based on an improved hybridization breeding algorithm, which solves the technical problems of easy getting trapped in local optima, low computational efficiency, and insufficient population diversity in the prior art, and achieves the technical effect of improving search efficiency and avoiding local optima.
[0054] The technical solutions in the embodiments of the present invention are intended to solve the above-mentioned technical problems, and the overall approach is as follows:
[0055] The embodiments of this invention include the following steps: acquiring high-dimensional big data from the industrial production process, performing preprocessing operations on missing and outlier values to obtain a processed dataset; dividing the features into several subspaces based on feature importance; for each subspace, randomly initializing the population and adopting a binary encoding scheme based on Gaussian distribution; using a ternary tournament strategy to divide rice population individuals into maintainer lines, restorer lines, and sterile lines; in the early stage of population iteration, i.e., the global search stage, using a fitness-guided mutation strategy; in the later stage of population iteration, i.e., the local search stage, using a simulated annealing combined with gradient descent strategy to enhance the effect of local search; comprehensively evaluating the classification results and the number of features; determining whether the current iteration number has reached the set maximum iteration number; if so, outputting the optimal feature subset.
[0056] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0057] like Figure 1 As shown in the embodiment of the present invention, the data feature selection method based on the improved hybridization breeding algorithm includes:
[0058] The feature space of the dataset to be selected is divided into several subspaces;
[0059] This step involves dividing the feature space of the dataset to be selected into several subspaces, including:
[0060] Through formula Calculated features Relevance with class label C ;in, Features Mutual information with class label C, , Features Information entropy , The information entropy of class label C, ;
[0061] The features are sorted in descending order according to their correlation SU values to obtain the feature space. ;
[0062] Through formula The number of features in the eigenspace is calculated. ;in, The total number of features in the dataset. The number of subspaces;
[0063] Select feature space Ranked 1st to Feature composition feature subspace Ranked arrive Feature composition feature subspace This process continues until all features have been assigned to their respective subspaces.
[0064] The rice population was divided into subspaces, and a three-way tournament strategy was used to classify the rice population individuals in each subspace into maintainer lines, restorer lines and sterile lines.
[0065] This step is explained in detail: the rice population is divided into various subspaces, and a three-way tournament strategy is used to classify the rice population individuals in each subspace into maintainer lines, restorer lines, and sterile lines, including:
[0066] Through formula The characteristic subspace is calculated. Importance ;in, Features Relevance to class label C;
[0067] Through formula The initial number of individuals in the subpopulation was calculated. ;in, The total number of individuals in the population. For characteristic subspace The importance of;
[0068] For the division Each feature subspace is randomly initialized. Individual rice plants were used as the original population;
[0069] Three individuals are randomly selected from the original rice population to calculate their fitness values. The individual with the highest fitness value is assigned to the maintainer line, the individual with the lowest fitness value is assigned to the sterile line, and the last individual is assigned to the restorer line. This process is repeated until all individuals in the population have been assigned, resulting in a three-line rice population: maintainer line... Recovery system and sterile line .
[0070] Hybridization of maintainer lines and sterile lines enhances the traits of sterile lines, resulting in superior sterile lines. The expression for this is: ;in, For the new individuals obtained through hybridization, For the maintainer individuals in the current iteration, For sterile individuals in the current iteration, and A random number between 0 and 1; if The fitness value is greater than that of the current sterile line individuals. Then Updated to Otherwise, it will not be updated.
[0071] Self-crossing is performed between restorer line individuals, using the optimal individual to guide the evolution of other individuals, thereby improving the overall traits of the restorer line. The expression for this is: ;in, For new individuals obtained through self-fertilization, As the current globally optimal individual, Individuals randomly selected from the restorer line. For the restorer individuals in the current iteration, It is a random number between 0 and 1; if The fitness value is greater than that of the current restorer individuals. Then Updated to Otherwise, it will not be updated.
[0072] To avoid slow convergence or even getting trapped in local optima due to frequent self-crossing operations by some restorer individuals when they are close to local optima, the following measures are also included:
[0073] For restorer individuals that have reached the self-pollination limit, a reset update is performed, expressed as follows: ;in, To reset the updated individual, For the currently reset Restoration class individual, and These represent the upper and lower bounds of the search space, respectively. It is a random number between 0 and 1. If If the maximum number of self-crosses is reached, then... Updated to Otherwise, it will not be updated.
[0074] During the global search phase, sterile individuals with fitness values less than a preset value or maintainer individuals with fitness values greater than a threshold are subjected to mutation operations, the expression of which is: ;in, These are new individuals from the sterile line or new individuals from the maintainer line after mutation. This refers to sterile individuals with a current fitness value less than a preset value or maintainer individuals with a fitness value greater than a threshold. For variable asynchronous length, , For fitness guiding coefficients, The fitness value is the fitness value of sterile individuals whose current fitness value is less than the preset value or the fitness value of maintainer individuals whose fitness value is greater than the threshold. It is a random number between 0 and 1, used to generate random perturbations; if The fitness value is greater than that of the current sterile or maintainer line individuals. Then Updated to Conversely, no update is made if the fitness is high. In other words, during the global search phase, a larger variable time length is set for sterile individuals with low fitness to increase the breadth of exploration; while a smaller variable time length is used for maintainer individuals with high fitness to preserve the characteristics of their superior solutions.
[0075] During the local search phase, the three lineages are optimized and updated; the expression is as follows: ;in, To optimize the updated individual, These are individuals from three lineages within the current population. For learning rate, The gradient of the current best individual. To simulate the random perturbation term of annealing, , For one in Random numbers within the interval For temperature parameters, This represents the current iteration number; that is, in the local search, the update direction of the current individual is calculated based on the gradient of the objective function. A temperature decay mechanism is introduced, allowing for the acceptance of poorer individuals under certain conditions, thus helping the algorithm escape local optima and ultimately leading to a more accurate and efficient solution. Simulated annealing introduces a temperature parameter... This controls the probability of accepting weaker individuals. During the search process, if a new individual... fitness value Compared to the current individual fitness value Even worse, simulated annealing will accept new individuals based on probability; this acceptance mechanism helps avoid premature convergence to local optima. Specifically, through the formula... The probability of an individual being accepted is calculated. ;in, and These are the fitness values of the new individual and the current individual, respectively. A random number between 0 and 1; if If the random number is greater than 0 and between 1, then the current individual will be... Update to a new individual ;if If the random number is less than or equal to 0 and between 1, then the current individual will be retained. ;
[0076] In this embodiment, through the formula The temperature parameters after iterative updates were calculated. ;in, Let be the attenuation factor, and take... , The temperature parameters are as shown before the update. In other words, the temperature parameters change as the iteration progresses. Gradually decrease.
[0077] By combining the updated individuals from each subpopulation, the globally optimal individual is obtained. Specifically, for populations Subpopulation needs to be calculated Individuals in The fitness of the individual is then compared with the best individuals in other subpopulations. Combined to obtain ,Will The fitness value as an individual The fitness value. At the end of each iteration, record the globally optimal individual. .
[0078] Through formula The probability of a feature being selected is calculated. , Representation of features The probability of being selected; where, The standard deviation is used to control for search diversity, and it gradually decreases as the search progresses. Index of features The mean value represents the average level of feature selection. It is initially uniformly distributed and updated based on fitness as iterations progress. If... A random number greater than 0 and 1 represents the globally optimal individual. medium-sized individuals Features Selected; if A random number less than or equal to 0 and between 1 represents the globally optimal individual. medium-sized individuals Features If not selected, the optimal feature subset is obtained if the iteration ends. In other words, it is determined whether the termination condition is met; if so, the optimal feature subset is output; otherwise, the process of three-system partitioning, individual updating, and individual evaluation is returned until the termination condition is met.
[0079] Corresponding to the data feature selection method described above, this invention also provides a data feature selection system based on an improved hybridization breeding algorithm, comprising:
[0080] The feature space segmentation module is used to segment the feature space of the dataset to be selected into several subspaces;
[0081] Specifically, the feature space segmentation module includes:
[0082] Correlation calculation unit, used to calculate using formulas Calculated features Relevance with class label C ;in, Features Mutual information with class label C, , Features Information entropy , The information entropy of class label C, ;
[0083] The feature sorting unit is used to sort the features in descending order according to their relevance SU values to obtain the feature space. ;
[0084] The characteristic number calculation unit is used to calculate the characteristic number using the formula The number of features in the eigenspace is calculated. ;in, The total number of features in the dataset. The number of subspaces;
[0085] Feature partitioning unit, used to select feature space Ranked 1st to Feature composition feature subspace Ranked arrive Feature composition feature subspace This process continues until all features have been assigned to their respective subspaces.
[0086] The rice population segmentation module is used to divide the rice population into various subspaces. The three-way tournament strategy is used to divide the rice population individuals in each subspace into maintainer lines, restorer lines and sterile lines.
[0087] Specifically, the rice population segmentation module includes:
[0088] Importance calculation unit, used to calculate importance using formulas The characteristic subspace is calculated. Importance ;in, Features Relevance to class label C;
[0089] The initial population size calculation unit is used to calculate the initial population size using the formula. The initial number of individuals in the subpopulation was calculated. ;in, The total number of individuals in the population. For characteristic subspace The importance of;
[0090] The original population initialization unit is used for partitioning. Each feature subspace is randomly initialized. Individual rice plants were used as the original population;
[0091] The rice population partitioning unit is used to randomly select three individuals from the original rice population, calculate their fitness values, assign the individual with the highest fitness value to the maintainer line, the individual with the lowest fitness value to the sterile line, and the last individual to the restorer line. This operation is repeated until all individuals in the population have been partitioned, resulting in a three-line rice population: maintainer line... Recovery system and sterile line .
[0092] The hybridization module is used to perform hybridization operations on maintainer lines and sterile lines to improve the traits of sterile lines, thereby obtaining superior sterile lines. Its expression is: ;in, For the new individuals obtained through hybridization, For the maintainer individuals in the current iteration, For sterile individuals in the current iteration, and A random number between 0 and 1; if The fitness value is greater than that of the current sterile line individuals. Then Updated to Otherwise, it will not be updated.
[0093] The self-crossing module is used to perform self-crossing operations between restorer line individuals. It guides the evolution of other individuals through the optimal restorer line individual, thereby improving the overall traits of the restorer line individuals. Its expression is: ;in, For new individuals obtained through self-fertilization, As the current globally optimal individual, Individuals randomly selected from the restorer line. For the restorer individuals in the current iteration, It is a random number between 0 and 1; if The fitness value is greater than that of the current restorer individuals. Then Updated to Otherwise, it will not be updated.
[0094] To avoid slow convergence or even getting trapped in local optima due to frequent self-crossing operations by some restorer individuals when they are close to local optima, the following measures are also included:
[0095] The reset / update module is used to reset and update restorer individuals that have reached the self-crossing limit. Its expression is: ;in, To reset the updated individual, For the currently reset Restoration class individual, and These represent the upper and lower bounds of the search space, respectively. It is a random number between 0 and 1. If If the maximum number of self-crosses is reached, then... Updated to Otherwise, it will not be updated.
[0096] The mutation module is used during the global search phase to perform mutation operations on sterile individuals with fitness values less than a preset value or maintainer individuals with fitness values greater than a threshold. Its expression is: ;in, These are new individuals from the sterile line or new individuals from the maintainer line after mutation. This refers to sterile individuals with a current fitness value less than a preset value or maintainer individuals with a fitness value greater than a threshold. For variable asynchronous length, , For fitness guiding coefficients, The fitness value is the fitness value of sterile individuals whose current fitness value is less than the preset value or the fitness value of maintainer individuals whose fitness value is greater than the threshold. It is a random number between 0 and 1, used to generate random perturbations; if The fitness value is greater than that of the current sterile or maintainer line individuals. Then Updated to Conversely, no update is made if the fitness is high. In other words, during the global search phase, a larger variable time length is set for sterile individuals with low fitness to increase the breadth of exploration; while a smaller variable time length is used for maintainer individuals with high fitness to preserve the characteristics of their superior solutions.
[0097] The optimization and update module is used to optimize and update individuals from the three lineages during the local search phase. Its expression is: ;in, To optimize the updated individual, These are individuals from three lineages within the current population. For learning rate, The gradient of the current best individual. To simulate the random perturbation term of annealing, , For one in Random numbers within the interval For temperature parameters, This represents the current iteration number; that is, in the local search, the update direction of the current individual is calculated based on the gradient of the objective function. A temperature decay mechanism is introduced, allowing for the acceptance of poorer individuals under certain conditions, thus helping the algorithm escape local optima and ultimately leading to a more accurate and efficient solution. Simulated annealing introduces a temperature parameter... This controls the probability of accepting weaker individuals. During the search process, if a new individual... fitness value Compared to the current individual fitness value Even worse, simulated annealing will accept new individuals based on probability; this acceptance mechanism helps avoid premature convergence to local optima. Specifically, through the formula... The probability of an individual being accepted is calculated. ;in, and These are the fitness values of the new individual and the current individual, respectively. A random number between 0 and 1; if If the random number is greater than 0 and between 1, then the current individual will be... Update to a new individual ;if If the random number is less than or equal to 0 and between 1, then the current individual will be retained. ;
[0098] Specifically, embodiments of the present invention also include:
[0099] Temperature parameter calculation module, used to calculate parameters using formulas The temperature parameters after iterative updates were calculated. ;in, Let be the attenuation factor, and take... , The temperature parameters are as shown before the update. In other words, the temperature parameters change as the iteration progresses. Gradually decrease.
[0100] The global optimal individual acquisition module combines the updated individuals from each subpopulation to obtain the globally optimal individual. Specifically, for populations Subpopulation needs to be calculated Individuals in The fitness of the individual is then compared with the best individuals in other subpopulations. Combined to obtain ,Will The fitness value as an individual The fitness value. At the end of each iteration, record the globally optimal individual. .
[0101] The optimal feature subset acquisition module is used to obtain the optimal feature subset through the formula. The probability of a feature being selected is calculated. , Representation of features The probability of being selected; where, The standard deviation is used to control for search diversity, and it gradually decreases as the search progresses. Index of features The mean value represents the average level of feature selection. It is initially uniformly distributed and updated based on fitness as iterations progress. If... A random number greater than 0 and 1 represents the globally optimal individual. medium-sized individuals Features Selected; if A random number less than or equal to 0 and between 1 represents the globally optimal individual. medium-sized individuals Features If not selected, the optimal feature subset is obtained if the iteration ends. In other words, it is determined whether the termination condition is met; if so, the optimal feature subset is output; otherwise, the process of three-system partitioning, individual updating, and individual evaluation is returned until the termination condition is met.
[0102] In summary, the feature selection method and system provided by the embodiments of the present invention can complete the dimensionality reduction operation in a short time in a simple and efficient manner, improve the performance of data analysis and models, achieve accurate selection of high-dimensional features, retain key feature information, reduce computational costs, and improve the efficiency of classification and prediction models.
[0103] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0104] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0105] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0106] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0107] Any aspects of this invention not described in detail in the embodiments are well-known techniques to those skilled in the art. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this invention and not to limit it. Although this invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this invention without departing from the spirit and scope of this invention, and all such modifications and substitutions should be covered within the scope of the claims of this invention.
Claims
1. A data feature selection method based on an improved hybridization breeding algorithm, characterized in that, include: The feature space of the dataset to be selected is divided into several subspaces; The rice population was divided into subspaces, and a three-way tournament strategy was used to classify the rice population individuals in each subspace into maintainer lines, restorer lines and sterile lines. The expression for hybridization of maintainer lines and sterile lines is: ;in, For the new individuals obtained through hybridization, For the maintainer individuals in the current iteration, For sterile individuals in the current iteration, and A random number between 0 and 1; The expression for self-crossing between restorer line individuals is: ;in, For new individuals obtained through self-fertilization, As the current globally optimal individual, Individuals randomly selected from the restorer line. For the restorer individuals in the current iteration, It is a random number between 0 and 1; During the global search phase, sterile individuals with fitness values less than a preset value or maintainer individuals with fitness values greater than a threshold are subjected to mutation operations, the expression of which is: ;in, These are new individuals from the sterile line or new individuals from the maintainer line after mutation. This refers to sterile individuals with a current fitness value less than a preset value or maintainer individuals with a fitness value greater than a threshold. For variable asynchronous length, , For fitness guiding coefficients, The fitness value is the fitness value of sterile individuals whose current fitness value is less than the preset value or the fitness value of maintainer individuals whose fitness value is greater than the threshold. It is a random number between 0 and 1; During the local search phase, the three lineages are optimized and updated; the expression is as follows: ;in, To optimize the updated individual, These are individuals from three lineages within the current population. For learning rate, The gradient of the current best individual. To simulate the random perturbation term of annealing, , For one in Random numbers within the interval For temperature parameters, This represents the current iteration number; expressed by the formula. The probability of an individual being accepted is calculated. ;in, and These are the fitness values of the new individual and the current individual, respectively. A random number between 0 and 1; if If the random number is greater than 0 and between 1, then the current individual will be... Update to a new individual ;if If the random number is less than or equal to 0 and between 1, then the current individual will be retained. ; By combining the updated individuals from each subpopulation, the globally optimal individual is obtained. ; Through formula The probability of a feature being selected is calculated. ;in, Standard deviation Index of features The mean; if A random number greater than 0 and 1 represents the globally optimal individual. medium-sized individuals Features If selected, the optimal feature subset is obtained when the iteration ends.
2. The data feature selection method based on the improved hybridization breeding algorithm as described in claim 1, characterized in that, The step of dividing the feature space of the dataset to be selected into several subspaces includes: Through formula Calculated features Relevance with class label C ;in, Features Mutual information with class label C, Features Information entropy The information entropy of class label C; The features are sorted in descending order according to their correlation SU values to obtain the feature space. ; Through formula The number of features in the eigenspace is calculated. ;in, The total number of features in the dataset. The number of subspaces; Select the feature space Ranked 1st to Feature composition feature subspace Ranked arrive Feature composition feature subspace This process continues until all features have been assigned to their respective subspaces.
3. The data feature selection method based on the improved hybridization breeding algorithm as described in claim 2, characterized in that, The process of dividing the rice population into subspaces and using a ternary tournament strategy to classify rice population individuals in each subspace into maintainer lines, restorer lines, and sterile lines includes: Through formula The characteristic subspace is calculated. Importance ;in, Features Relevance to class label C; Through formula The initial number of individuals in the subpopulation was calculated. ;in, The total number of individuals in the population. For characteristic subspace The importance of; For the division Each feature subspace is randomly initialized. Individual rice plants were used as the original population; Three individuals are randomly selected from the original rice population to calculate their fitness values. The individual with the highest fitness value is assigned to the maintainer line, the individual with the lowest fitness value is assigned to the sterile line, and the last individual is assigned to the restorer line. This process is repeated until all individuals in the population have been assigned, resulting in a three-line rice population: maintainer line... Recovery system and sterile line .
4. The data feature selection method based on the improved hybridization breeding algorithm as described in claim 1, characterized in that, Also includes: For restorer individuals that have reached the self-pollination limit, a reset update is performed, expressed as follows: ;in, To reset the updated individual, For the currently reset Restoration class individual, and These represent the upper and lower bounds of the search space, respectively. It is a random number between 0 and 1.
5. The data feature selection method based on the improved hybridization breeding algorithm as described in any one of claims 1-4, characterized in that, Also includes: Through formula The temperature parameters after iterative updates were calculated. ;in, Let be the attenuation factor, and take... , The temperature parameters are as shown before the update.
6. A data feature selection system based on an improved hybridization breeding algorithm, characterized in that, include: The feature space segmentation module is used to segment the feature space of the dataset to be selected into several subspaces; The rice population segmentation module is used to divide the rice population into various subspaces. The three-way tournament strategy is used to divide the rice population individuals in each subspace into maintainer lines, restorer lines and sterile lines. The hybridization module is used to perform hybridization operations on maintainer lines and sterile lines; its expression is: ;in, For the new individuals obtained through hybridization, For the maintainer individuals in the current iteration, For sterile individuals in the current iteration, and A random number between 0 and 1; The self-crossing module is used to perform self-crossing operations between restorer line individuals; its expression is: ;in, For new individuals obtained through self-fertilization, As the current globally optimal individual, Individuals randomly selected from the restorer line. For the restorer individuals in the current iteration, It is a random number between 0 and 1; The mutation module is used during the global search phase to perform mutation operations on sterile individuals with fitness values less than a preset value or maintainer individuals with fitness values greater than a threshold. Its expression is: ;in, These are new individuals from the sterile line or new individuals from the maintainer line after mutation. This refers to sterile individuals with a current fitness value less than a preset value or maintainer individuals with a fitness value greater than a threshold. For variable asynchronous length, , For fitness guiding coefficients, The fitness value is the fitness value of sterile individuals whose current fitness value is less than the preset value or the fitness value of maintainer individuals whose fitness value is greater than the threshold. It is a random number between 0 and 1; The optimization and update module is used to optimize and update individuals from the three lineages during the local search phase. Its expression is: ;in, To optimize the updated individual, These are individuals from three lineages within the current population. For learning rate, The gradient of the current best individual. To simulate the random perturbation term of annealing, , For one in Random numbers within the interval For temperature parameters, This represents the current iteration number; expressed by the formula. The probability of an individual being accepted is calculated. ;in, and These are the fitness values of the new individual and the current individual, respectively. A random number between 0 and 1; if If the random number is greater than 0 and between 1, then the current individual will be... Update to a new individual ;if If the random number is less than or equal to 0 and between 1, then the current individual will be retained. ; The global optimal individual acquisition module combines the updated individuals from each subpopulation to obtain the globally optimal individual. ; The optimal feature subset acquisition module is used to obtain the optimal feature subset through the formula. The probability of a feature being selected is calculated. ;in, Standard deviation Index of features The mean; if A random number greater than 0 and 1 represents the globally optimal individual. medium-sized individuals Features If selected, the optimal feature subset is obtained when the iteration ends.
7. The data feature selection system based on the improved hybridization breeding algorithm as described in claim 6, characterized in that, The feature space segmentation module includes: Correlation calculation unit, used to calculate using formulas Calculated features Relevance with class label C ;in, Features Mutual information with class label C, Features Information entropy The information entropy of class label C; The feature sorting unit is used to sort the features in descending order according to the correlation SU value to obtain the feature space. ; The characteristic number calculation unit is used to calculate the characteristic number using the formula The number of features in the eigenspace is calculated. ;in, The total number of features in the dataset. The number of subspaces; Feature partitioning unit, used to select the feature space Ranked 1st to Feature composition feature subspace Ranked arrive Feature composition feature subspace This process continues until all features have been assigned to their respective subspaces.
8. The data feature selection system based on the improved hybridization breeding algorithm as described in claim 7, characterized in that, The rice population segmentation module includes: Importance calculation unit, used to calculate importance using formulas The characteristic subspace is calculated. Importance ;in, Features Relevance to class label C; The initial population size calculation unit is used to calculate the initial population size using the formula. The initial number of individuals in the subpopulation was calculated. ;in, The total number of individuals in the population. For characteristic subspace The importance of; The original population initialization unit is used for partitioning. Each feature subspace is randomly initialized. Individual rice plants were used as the original population; The rice population partitioning unit is used to randomly select three individuals from the original rice population, calculate their fitness values, assign the individual with the highest fitness value to the maintainer line, the individual with the lowest fitness value to the sterile line, and the last individual to the restorer line. This operation is repeated until all individuals in the population have been partitioned, resulting in a three-line rice population: maintainer line... Recovery system and sterile line .
9. The data feature selection system based on the improved hybridization breeding algorithm as described in claim 6, characterized in that, Also includes: The reset / update module is used to reset and update restorer individuals that have reached the self-crossing limit. Its expression is: ;in, To reset the updated individual, For the currently reset Restoration class individual, and These represent the upper and lower bounds of the search space, respectively. It is a random number between 0 and 1.
10. The data feature selection system based on the improved hybridization breeding algorithm as described in any one of claims 6-9, characterized in that, Also includes: Temperature parameter calculation module, used to calculate parameters using formulas The temperature parameters after iterative updates were calculated. ;in, Let be the attenuation factor, and take... , The temperature parameters are as shown before the update.