Multi-group firefly optimization method combined with simulated annealing strategy

By combining simulated annealing strategies and multiple sets of mechanisms to improve the firefly algorithm, the problems of premature convergence and redundant search in high-dimensional feature space are solved, efficient and accurate feature selection is achieved, and the performance and computing efficiency of the model are improved.

CN120493738APending Publication Date: 2025-08-15JIUJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510598502.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional firefly algorithms are prone to generate redundant search paths in high-dimensional and large-scale feature spaces, resulting in early maturity convergence and local optimum, making it difficult to meet the model accuracy and computing efficiency requirements of application scenarios such as intelligent diagnosis and financial risk control.

Method used

Combining simulated annealing strategies and multiple sets of mechanisms to improve the firefly optimization method, enhancing the ability to jump out of local extreme values ​​through simulated annealing, using multiple sets of local collaboration mechanisms to reduce redundant attraction, and using adaptive control parameters and Cauchy jump to update the optimal firefly position.

Benefits of technology

It significantly improves the globality and convergence stability of feature subset search, reduces feature dimensions, improves the accuracy and robustness of classifiers, shortens model training and inference time, and optimizes resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493738A_ABST
    Figure CN120493738A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data models, and discloses a multi-group firefly optimization method combined with a simulated annealing strategy, which comprises a simulated annealing method; multiple groups of strategies; and analyzing the complexity of the MSAFA. The invention provides an improved firefly algorithm variant, the variant adopts a simulated annealing (SA) algorithm to escape local traps, and a plurality of groups of strategies are utilized to enhance the global search capability of a firefly algorithm. When a database has the problems of class imbalance and high dimension, the performance of a spatial data mining (SDP) model constructed by a basic classification algorithm is weak. The proposed firefly algorithm variant significantly improves the performance of the firefly algorithm, and applies the firefly algorithm variant to the SDP model to reduce redundant features and improve the efficiency of the SDP model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to but is not limited to the field of data model technology, and in particular relates to a multi-group firefly optimization method combined with a simulated annealing strategy. Background Art

[0002] Generally speaking, traditional feature selection (FS) methods can be categorized into filter-based methods (FFS), wrapper-based methods (WFS), and embedded methods (EFS) based on their evaluation mechanisms. FFS is the most computationally efficient, while WFS and EFS offer superior classification performance due to their consideration of the interactions between selected features and the classifier. However, finding the optimal solution is a challenging task for FS due to the vast search space for feature subsets. Evolutionary computation (EC) is one of the most effective search techniques for solving NP-hard problems and has been widely used in feature selection optimization problems, such as particle swarm optimization (PSO), artificial bee colony optimization (ABC), ant colony optimization (ACO), firefly algorithm (FA) (Yang, 2009), and cuckoo search (CS). FA, leveraging the self-learning and self-organizing capabilities of the population, can effectively search for the optimal solution, thereby solving feature selection optimization problems. However, according to the FA principle, interactions between any two fireflies can lead to premature convergence and trapping in local optima.

[0003] In existing technologies, the traditional firefly algorithm (FA), proposed by Yang (2009), has been applied to feature selection tasks. This algorithm simulates the attraction mechanism of fireflies' luminescence behavior to optimize search within a feature subset space. This method dynamically adjusts search direction based on the attraction between individuals, possesses self-learning and self-organizing capabilities, and can balance global exploration with local exploitation to a certain extent. Consequently, it has achieved promising results in areas such as feature selection and classifier optimization.

[0004] However, in industrial applications, traditional FA has obvious technical bottlenecks: since there is mutual attraction between any two fireflies, the algorithm is prone to generate redundant search paths in high-dimensional, large-scale feature spaces, which in turn causes premature convergence and falls into local optimality, reducing the effect of feature selection results on improving classifier performance; at the same time, in actual application scenarios such as intelligent diagnosis, financial risk control, and medical prediction, feature selection algorithms have higher requirements on model accuracy and computational overhead. The existing FA is still difficult to meet the needs of complex engineering applications in terms of search efficiency and the globality of the solution. Summary of the Invention

[0005] In response to the problems existing in the prior art, the present invention provides a multi-group firefly optimization method combined with a simulated annealing strategy and its feature selection application.

[0006] The present invention is implemented as follows: a multi-group firefly optimization method combined with a simulated annealing strategy, the method comprising:

[0007] S1: simulated annealing method;

[0008] S2: multiple group strategies;

[0009] S3: Complexity analysis of MSAFA.

[0010] Furthermore, the S1 specifically includes:

[0011] Simulated annealing is a variant of the Metropolis algorithm. It simulates the physical annealing process by controlling temperature parameters to find the optimal state of the material. The basic simulated annealing method is shown in the following formula:

[0012] x i (t+1)=x i (t)+λ

[0013]

[0014] T(t+1)=T(t)ρ

[0015] When λ follows a random distribution, the algorithm generates new solutions to escape the local minimum. When a poor solution is generated, it is accepted with probability p, thereby maintaining the diversity of the population, where is a random number in the interval [0,1]. In addition, as the number of iterations increases, the temperature parameter decays according to the control parameter ρ. Therefore, the simulated annealing algorithm SA contains two random processes: one is responsible for generating new solutions, and the other determines the acceptance mechanism of new solutions;

[0016] The SA algorithm is used to optimize the overall position of the population in each iteration. To accelerate convergence, an adaptive control parameter α is introduced to generate new solutions. The improved attractive force movement formula is shown below. In terms of parameter settings, α(0) = 0.2, ρ = 0.6, initial temperature T(0) = 1, and termination temperature Tmin = 0.1.

[0017]

[0018]

[0019] Furthermore, the S2 specifically includes:

[0020] In the firefly algorithm (FA), there is attraction between any two fireflies, which may lead to some wrong directions and redundant attraction, causing the algorithm to converge prematurely;

[0021] In order to reduce the redundant attraction of the firefly algorithm, a multi-group strategy based on the position information of fireflies is proposed. In the multi-group mechanism, for a firefly, the fireflies with better fitness and located in a circle with the firefly as the center and the distance from the firefly to the best firefly as the radius form a group. Taking fireflies i and j as an example, fireflies a, b, c and d form one group, and fireflies e, f and d form another group. Each group effectively performs local search in its circular space, while all groups perform global search in the entire space. In particular, firefly d is a common member of the two groups and plays the role of an information exchange center, thereby enhancing the global search capability. When performing global search in the multi-group mechanism, the number of migrations plays a key role, but this number cannot be determined in advance. Therefore, a control parameter times is proposed to determine the number of migrations. The maximum value of times is set to 8, that is, each firefly can be a member of up to 8 groups at the same time.

[0022] In order to increase the diversity of the population, Cauchy jump is used to update the position of the best firefly, as shown in the following formula;

[0023] in, represents the position of the current best firefly in the dth dimension in each iteration, Cauchy is a random value of the standard Cauchy distribution, α represents the control parameter, β represents the attraction rate, γ represents the Euler distance between fireflies, Distance represents the Euler distance between the current firefly and the best firefly, and ε represents a random number value [0,1];

[0024] distance i =||x i -x best ||

[0025] x i (t+1)=x i (t)+β(x j (t)-x i (t))+α(t)ε

[0026]

[0027]

[0028] Furthermore, the S3 specifically includes:

[0029] In order to clearly study and evaluate the complexity of the proposed algorithm, it is compared with the complexity of some improved Firefly algorithm variants, including MFA, RaFA, and LiFA;

[0030] First, we analyze the complexity of these advanced firefly algorithm variants. We assume that each competitor uses the same settings, the termination condition is Max, the population size is N, and the dimension of the problem is D. In RaFA, the execution time to evaluate a solution is t1. Therefore, the time complexity of RaFA is O(MaxND*(t1)), which can be simplified to O(N).

[0031] In MFA, the execution time to evaluate a solution is t2; therefore, the time complexity of MFA is O(MaxNND(t2)), which, like FA, can be expressed as O(N 2 ); In LiFA, the execution time of evaluating a solution is t3, and the time of selecting a group G is t4; therefore, the time complexity of LiFA is O(MaxNGD(t3+t4)), which can be expressed as O(N*G); the time complexity of LiFA depends on the size of the adaptive group G; then, the time complexity of the proposed algorithm is analyzed. The time of updating the new position of the entire population is t5, the time of calculating the distance between the current firefly and other fireflies is t6, the time of evaluating the generated group M consisting of k individuals is t7, and the time of performing a Cauchy jump on the current best firefly is t8;

[0032] Therefore, DSFA can be expressed as O(N*K 2 ); Theoretically, the value of k is between [0,19]; through preliminary experiments, the distribution of parameter k is between [0,11]; in other words, the time complexity of MSAFA is O(N 2 ) and O(N 3 ); From the above analysis, it can be seen that RaFA performs the best in time complexity, followed by MFA, LiFA and MSAFA; although the time complexity of MSAFA is worse than these firefly algorithm variants, the increased computational cost is worthwhile for the proposed method; one of the reasons is that the proposed simulated annealing SA is used to help the firefly algorithm increase the probability of escaping from the local trap; another reason is that the multi-group mechanism can enhance the global search ability of the firefly algorithm.

[0033] Another object of the present invention is to provide a feature selection application, which utilizes feature selection problems for evaluation, specifically comprising:

[0034] The well-known NASA database was used to evaluate the quality of the SDP model. The database contains 13 projects, which are characterized by class imbalance and high dimensionality. To verify the performance of the proposed feature selection method, six datasets were selected from the NASA repository, including MW1, KC3, PC1, KC1, PC2 and MC1. The scale of these datasets varies, with the number of modules ranging from 403 to 9466. In addition, three classification algorithms were used to build the SDP model, including support vector machine (SVM), K-nearest neighbor (KNN) and random forest (RF). In terms of feature selection, MSAFA was compared with particle swarm optimization (PSO), artificial bee colony algorithm (ABC) and the five firefly algorithm (FA) variants mentioned above to improve the performance of the SDP model. The area under the curve was used to evaluate the SDP model for each selected feature subset. This metric provides a comprehensive measure of performance across all possible classification thresholds. The maximum number of iterations was set to 1000 and the population size was set to 20. All compared algorithms were run 10 times to search for the optimal feature subset to improve the performance of the SDP model.

[0035] Each algorithm significantly improved the performance of the SDP model. In comparison with competitors, the proposed method obtained the best values on 17 of the 18 SDP models evaluated, followed by the ABC algorithm, which achieved the best results on 10 of the 18 SDP models. PSO and LiFA algorithms both provided the best AUC values for 8 SDP models. On the KC3 dataset using the SVM classification algorithm, MFA achieved the best performance of the SDP model. On the PC1 dataset using RF for the SDP model, DFA outperformed the other algorithms. The Friedman test tool was still used to analyze the performance of all feature selection methods. MSAFA achieved the best ranking score. 1.8611, followed by ABC, LiFA, PSO, DFA, FA, MFA and NaFA; the feature selection process plays a key role in improving the performance of the SDP model; the original continuous representation method proposed by ABC can solve the feature selection problem, where each element of the continuous vector corresponds to an original feature; each element of the feature vector is first represented in the [0,1] space; then, if the element is greater than a given constant θ, the corresponding feature is selected; otherwise, the feature is not selected; therefore, each firefly can represent a feature subset, where each dimension of the firefly corresponds to a feature, and the optimal feature subset is searched as a continuous optimization problem with θ set to 0.5.

[0036] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0037] First, in industrial application scenarios such as smart manufacturing, financial risk control, and medical diagnosis, feature selection algorithms need to efficiently select subsets that contribute most to model performance in high-dimensional feature spaces to reduce model complexity and improve prediction accuracy and inference speed. The application of traditional firefly algorithms in feature selection is limited by redundant attraction, premature convergence, and local optimal traps, making it difficult to effectively balance feature relevance and redundancy on large-scale datasets, which in turn affects the model's generalization ability and actual deployment effect. Therefore, the industry urgently needs an optimization method that can enhance population diversity, improve global search capabilities, and maintain reasonable computational overhead to adapt to the needs of complex tasks for efficient and accurate feature selection.

[0038] The improved firefly optimization method (MSAFA) combining simulated annealing strategy and multi-group mechanism is introduced. The ability to jump out of local extreme values is enhanced by simulated annealing, and redundant attraction is reduced by multi-group local cooperation mechanism, which significantly improves the globality and convergence stability of feature subset search. In industrial applications, the use of this method for feature selection not only effectively reduces the feature dimension and improves the accuracy and robustness of the classifier on data sets in different fields, but also shortens the model training and inference time, optimizes resource consumption and deployment efficiency, and meets the actual needs of efficient and intelligent feature selection technology in a big data environment. The present invention proposes a multi-group firefly optimization method combined with a simulated annealing strategy and its feature selection application. Simulated annealing (SA) is used to help fireflies escape from local optimality, and multi-group strategies are designed to protect the diversity of the population that quickly traverses the search space. To verify the performance of the proposed method, the present invention uses 28 optimization benchmark functions and 6 software projects selected from the National Aeronautics and Space Administration (NASA) repository for evaluation. In addition, to fairly demonstrate the competitiveness and superiority of the proposed algorithm, this paper compares it with seven state-of-the-art swarm intelligence algorithms, which are tested on an SDP model containing three classification algorithms: k-nearest neighbor (KNN), random forest (RF), and support vector machine (SVM).

[0039] This paper proposes an improved Firefly algorithm variant that employs simulated annealing (SA) to escape local traps and leverages multiple group strategies to enhance the global search capabilities of the Firefly algorithm. Spatial data mining (SDP) models constructed using basic classification algorithms perform poorly when the database suffers from class imbalance and high dimensionality. The proposed Firefly algorithm variant significantly improves the performance of the Firefly algorithm and is applied to the SDP model to reduce redundant features and improve its efficiency. Two sets of experiments were conducted to test the performance of the proposed method. The first set of experiments evaluated the efficiency of the proposed Firefly algorithm variant using 28 optimization problems; the second set of experiments used six datasets with class imbalance and high dimensionality to test the proposed method's feature selection capabilities in SDP. In comparisons with seven state-of-the-art swarm intelligence algorithms, the Multi-Strategy Adaptive Firefly Algorithm (MSAFA) significantly outperformed its competitors in solving continuous optimization problems. In future work, we will utilize the proposed algorithm to optimize multi-objective solutions to SDP problems.

[0040] Second, this technical solution directly promotes the cost-effectiveness of enterprises in extracting value from data by balancing accuracy and efficiency, while building business moats in multiple dimensions such as compliance, agility, and innovation. Based on the experimental results of other technical solutions, it is inferred that the expected benefits of applications in various fields will generally increase by 5%-15%. Its value is not only reflected in the improvement of technical indicators, but also in the full empowerment of the business decision-making chain, specifically: improving model efficiency and accuracy; avoiding the collection of irrelevant data and reducing data collection costs by screening key features; visualizing key risk factors and enhancing business explainability; achieving real-time edge analysis of industrial equipment failure prediction to support real-time decision-making scenarios; discovering non-intuitive feature associations through optimization strategies to drive business innovation; avoiding model reliance on accidental features, improving business stability, and achieving risk control and compliance; and the autonomously optimized feature selection framework can serve as the core technology asset of the enterprise to build its competitive advantage. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flow chart of a method for a multi-group firefly optimization method combined with a simulated annealing strategy provided by an embodiment of the present invention.

[0042] Figure 2 is a framework diagram of a distance guidance mechanism provided by an embodiment of the present invention;

[0043] Figure 3 These are the first 10 selected features provided by the embodiment of the present invention. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0045] like Figure 1 As shown, an embodiment of the present invention provides a multi-group firefly optimization method combined with a simulated annealing strategy and a feature selection application thereof, the method comprising:

[0046] S1: simulated annealing method;

[0047] S2: multiple group strategies;

[0048] S3: Complexity analysis of MSAFA.

[0049] Said S1 specifically includes:

[0050] Simulated annealing is a variant of the Metropolis algorithm. It simulates the physical annealing process by controlling temperature parameters to find the optimal state of the material. The basic simulated annealing method is shown in the following formula:

[0051] x i (t+1)=x i (t)+λ

[0052]

[0053] T(t+1)=T(t)ρ

[0054] When λ follows a random distribution, the algorithm generates new solutions to escape the local minimum. When a poor solution is generated, it is accepted with probability p, thereby maintaining the diversity of the population, where is a random number in the interval [0,1]. In addition, as the number of iterations increases, the temperature parameter decays according to the control parameter ρ. Therefore, the simulated annealing algorithm SA contains two random processes: one is responsible for generating new solutions, and the other determines the acceptance mechanism of new solutions;

[0055] The SA algorithm is used to optimize the overall position of the population in each iteration. To accelerate convergence, an adaptive control parameter α is introduced to generate new solutions. The improved attractive force movement formula is shown below. In terms of parameter settings, α(0) = 0.2, ρ = 0.6, initial temperature T(0) = 1, and termination temperature Tmin = 0.1.

[0056]

[0057]

[0058] The S2 specifically includes:

[0059] The proposed multiple groups of mechanisms such as Figure 2 As shown in the figure, in the firefly algorithm FA, there is attraction between any two fireflies, which may lead to some wrong directions and redundant attraction, causing the algorithm to converge prematurely;

[0060] In order to reduce the redundant attraction of the firefly algorithm, a multi-group strategy based on the position information of fireflies is proposed. In the multi-group mechanism, for a firefly, the fireflies with better fitness and located in a circle with the firefly as the center and the distance from the firefly to the best firefly as the radius form a group. Taking fireflies i and j as an example, fireflies a, b, c and d form one group, and fireflies e, f and d form another group. Each group effectively performs local search in its circular space, while all groups perform global search in the entire space. In particular, firefly d is a common member of the two groups and plays the role of an information exchange center, thereby enhancing the global search capability. When performing global search in the multi-group mechanism, the number of migrations plays a key role, but this number cannot be determined in advance. Therefore, a control parameter times is proposed to determine the number of migrations. The maximum value of times is set to 8, that is, each firefly can be a member of up to 8 groups at the same time.

[0061] In order to increase the diversity of the population, Cauchy jump is used to update the position of the best firefly, as shown in the following formula;

[0062] in, represents the position of the current best firefly in the dth dimension in each iteration, Cauchy is a random value of the standard Cauchy distribution, α represents the control parameter, β represents the attraction rate, γ represents the Euler distance between fireflies, Distance represents the Euler distance between the current firefly and the best firefly, and ε represents a random number value [0,1];

[0063] distance i =||x i -x best ||

[0064] x i (t+1)=x i (t)+β(x j (t)-x i (t))+α(t)ε

[0065]

[0066]

[0067] The S3 specifically includes:

[0068] In order to clearly study and evaluate the complexity of the proposed algorithm, it is compared with the complexity of some improved Firefly algorithm variants, including MFA, RaFA, and LiFA;

[0069] First, we analyze the complexity of these advanced firefly algorithm variants. We assume that each competitor uses the same settings, the termination condition is Max, the population size is N, and the dimension of the problem is D. In RaFA, the execution time to evaluate a solution is t1. Therefore, the time complexity of RaFA is O(MaxND*(t1)), which can be simplified to O(N).

[0070] In MFA, the execution time to evaluate a solution is t2; therefore, the time complexity of MFA is O(MaxNND(t2)), which, like FA, can be expressed as O(N 2 ); In LiFA, the execution time of evaluating a solution is t3, and the time of selecting a group G is t4; therefore, the time complexity of LiFA is O(MaxNGD(t3+t4)), which can be expressed as O(N*G); the time complexity of LiFA depends on the size of the adaptive group G; then, the time complexity of the proposed algorithm is analyzed. The time of updating the new position of the entire population is t5, the time of calculating the distance between the current firefly and other fireflies is t6, the time of evaluating the generated group M consisting of k individuals is t7, and the time of performing a Cauchy jump on the current best firefly is t8;

[0071] Therefore, DSFA can be expressed as O(N*K 2 ); Theoretically, the value of k is between [0,19]; through preliminary experiments, the distribution of parameter k is between [0,11]; in other words, the time complexity of MSAFA is O(N 2 ) and O(N 3 ); From the above analysis, it can be seen that RaFA performs the best in time complexity, followed by MFA, LiFA and MSAFA; although the time complexity of MSAFA is worse than these firefly algorithm variants, the increased computational cost is worthwhile for the proposed method; one of the reasons is that the proposed simulated annealing SA is used to help the firefly algorithm increase the probability of escaping from the local trap; another reason is that the multi-group mechanism can enhance the global search ability of the firefly algorithm.

[0072] An embodiment of the present invention provides a feature selection application, which utilizes feature selection problems for evaluation, specifically including:

[0073] The well-known NASA database was used to evaluate the quality of the SDP model. The database contains 13 projects, which are characterized by class imbalance and high dimensionality. To verify the performance of the proposed feature selection method, six datasets were selected from the NASA repository, including MW1, KC3, PC1, KC1, PC2 and MC1. The scale of these datasets varies, with the number of modules ranging from 403 to 9466. In addition, three classification algorithms were used to build the SDP model, including support vector machine (SVM), K-nearest neighbor (KNN) and random forest (RF). In terms of feature selection, MSAFA was compared with particle swarm optimization (PSO), artificial bee colony algorithm (ABC) and the five firefly algorithm (FA) variants mentioned above to improve the performance of the SDP model. The area under the curve was used to evaluate the SDP model for each selected feature subset. This metric provides a comprehensive measure of performance across all possible classification thresholds. The maximum number of iterations was set to 1000 and the population size was set to 20. All compared algorithms were run 10 times to search for the optimal feature subset to improve the performance of the SDP model.

[0074] Each algorithm significantly improved the performance of the SDP model. In comparison with competitors, the proposed method obtained the best values on 17 of the 18 SDP models evaluated, followed by the ABC algorithm, which achieved the best results on 10 of the 18 SDP models. PSO and LiFA algorithms both provided the best AUC values for 8 SDP models. On the KC3 dataset using the SVM classification algorithm, MFA achieved the best performance of the SDP model. On the PC1 dataset using RF for the SDP model, DFA outperformed the other algorithms. The Friedman test tool was still used to analyze the performance of all feature selection methods. MSAFA achieved the best ranking score. 1.8611, followed by ABC, LiFA, PSO, DFA, FA, MFA and NaFA; the feature selection process plays a key role in improving the performance of the SDP model; the original continuous representation method proposed by ABC can solve the feature selection problem, where each element of the continuous vector corresponds to an original feature; each element of the feature vector is first represented in the [0,1] space; then, if the element is greater than a given constant θ, the corresponding feature is selected; otherwise, the feature is not selected; therefore, each firefly can represent a feature subset, where each dimension of the firefly corresponds to a feature, and the optimal feature subset is searched as a continuous optimization problem with θ set to 0.5.

[0075] This technical solution was applied to software defect prediction on the MW1, KC3, PC1, KC1, PC2, and MC1 datasets from the NASA database. First, a software defect prediction model was constructed using advanced machine learning algorithms (SVM, KNN, RF). Next, the Firefly variant of this solution was applied for feature selection. Finally, the effectiveness of this solution was scientifically evaluated by comparing it with multiple sets of advanced algorithms. Compared to seven advanced solutions, this solution achieved significant advantages over 13 of the 18 prediction models, with prediction accuracy improvements ranging from 1% to 20%.

[0076] Relevant evidence of the technical effects achieved by the embodiments of the present invention.

[0077] 1. Evaluation using CEC2013 benchmark functions

[0078] The CEC2013 benchmark consists of 28 optimization problems proposed by Liang et al. (2013). These problems are widely used to evaluate the performance of optimization algorithms, including unimodal functions (f1-f6), basic multimodal functions (f7-f20), and composite functions (f21-f28). In comparisons with the original Firefly Algorithm (FA) and four state-of-the-art Firefly Algorithm variants, all algorithms were run 30 times on these optimization problems in 30 and 50 dimensions, respectively. The parameters of each algorithm were set according to the original settings. The maximum number of iterations and population size of each algorithm were set to 5E5 and 20, respectively. The mean (mean) and standard deviation (SD) of the best solution for each problem were recorded, and their efficiency was analyzed using the Wilcoxon rank sum test and Friedman test. Detailed experimental results are shown in Tables 1 and 2. The best solution for each problem is marked in bold. The results of the Wilcoxon rank sum test are indicated by "+ / ≈ / -" with a significance level of 0.05, and the results of the Friedman test are presented as ranking scores. “+ / ≈ / -” means better than, equivalent to, and worse than, respectively.

[0079] Table 1 Comparative experimental results with FA, MFA, NaFA, DFA and LiFA on the 30-dimensional CEC2013 problem

[0080]

[0081] In Table 1, MSAFA achieved optimal solutions on 16 problems, while other algorithms achieved optimal solutions on the remaining 12. The Wilcoxon rank sum test results at the bottom of the table show that, out of 28 functions, MSAFA outperformed FA on 24, MFA on 22, NaFA on 19, DFA on 21, and LiFA on 14. Among the 28 problems, MSAFA was inferior to these algorithms on 4, 3, 6, 4, and 4 problems, respectively. The Friedman test results show that MSAFA achieved the best ranking score of 1.875. The remaining ranking order is LiFA, MFA, NaFA, DFA, and FA. It can be concluded that the proposed algorithm outperforms other algorithms on most 30-dimensional CEC2013 problems.

[0082] Table 2 Comparative experimental results with FA, MFA, NaFA, DFA and LiFA on the 50-dimensional CEC2013 problem

[0083]

[0084] In Table 2, MSAFA achieved optimal solutions on 16 of the 28 problems. In the "+ / ≈ / -" notation, MSAFA outperformed FA on 24 of the 28 functions, MFA on 22, NaFA on 19, DFA on 21, and LiFA on 15. Among the 28 problems, MSAFA was inferior to them in 3, 3, 5, 4, and 3 problems, respectively. According to the Friedman test results, MSAFA achieved the best ranking score of 1.9821. The remaining ranking order was LiFA, MFA, NaFA, DFA, and FA. It can be concluded that the proposed algorithm still significantly outperforms other algorithms on most 50-dimensional CEC2013 problems.

[0085] Table 3. Running time of the algorithm on 30-dimensional and 50-dimensional CEC2013 problems (seconds)

[0086]

[0087] Table 4 As the problem dimension increases from 30 to 50, the ratio of computational cost also increases

[0088]

[0089] The computational cost of each algorithm and the ratio of the increase in computational cost as the problem dimension increases from 30 to 50 are shown in Tables 3 and 4. Although MSAFA performs slightly worse than other algorithms on the 30- and 50-dimensional CEC2013 problems, its computational cost ratio outperforms most other algorithms, achieving a second-place ranking score of 2.18 using the Friedman test tool. In other words, taking computational cost into account, MSAFA is likely to outperform most algorithms as the dimension of the optimization problem increases.

[0090] 2. Evaluation using feature selection problems

[0091] NASA, a well-known database for evaluating the quality of SDP models, is available on Shepperd et al. (2013). It contains 13 projects characterized by class imbalance and high dimensionality. To validate the performance of the proposed feature selection method, six datasets were selected from the NASA repository, including MW1, KC3, PC1, KC1, PC2, and MC1. The sizes of these datasets range from 403 to 9466 modules, as shown in Table 5. Furthermore, three classification algorithms were employed to construct the SDP model: support vector machine (SVM), K-nearest neighbor (KNN), and random forest (RF). Regarding feature selection, MSAFA was compared with particle swarm optimization (PSO), artificial bee colony algorithm (ABC), and the five aforementioned firefly algorithm (FA) variants to improve the performance of the SDP model. The area under the curve (AUC, Fawcett, 2006) was used to evaluate the SDP model for each selected feature subset. This metric provides a comprehensive measure of performance across all possible classification thresholds. The maximum number of iterations was set to 1000, and the population size was set to 20. All algorithms were run 10 times to search for the optimal feature subset to improve the performance of the SDP model. The experimental results are shown in Table 6. Figure 3 The top 10 occurrences of the selected features are shown.

[0092] Table 5 Dataset description

[0093]

[0094] Table 6 AUC evaluation results of each feature selection scheme in SDP model application

[0095]

[0096] Table 6 shows that this solution is more advanced than other solutions, and through professional data analysis and evaluation, its advantages are significant, with a score gap nearly double that of ABC, which ranks second.

[0097] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0098] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.

Claims

1. A multi-group firefly optimization method combined with simulated annealing strategy, characterized in that: The method includes: S1, using the simulated annealing method to update the position of the individuals in the population, the simulated annealing method is iterated based on the temperature control parameter and uses the adaptive control parameter α to generate a new solution; S2, grouping the firefly population based on a multi-group strategy, determining group members according to the distance between each individual and the current optimal individual, performing local search within the group and global search between groups; S3, uses a jump mechanism based on the standard Cauchy distribution to update the position of the current best individual; S4, analyze the time complexity of the multi-group simulated annealing firefly optimization method (MSAFA), assuming that the population size is N, the number of individuals in a single group is k, and the complexity is O(N 2 ) to O(N 3 )between; S5 is applied to the feature selection task. It encodes feature subsets based on continuous vectors, where each element corresponds to a feature. Features are selected by threshold θ in the continuous space [0, 1], and θ is set to 0.

5.

2. The method according to claim 1, characterized in that In the simulated annealing method, the initial temperature T(0) is set to 1, the end temperature Tmin is set to 0.1, and the temperature decay is performed according to the control parameter ρ=0.

6.

3. The method according to claim 1, characterized in that The adaptive control parameter α is initially set to 0.2 and is dynamically adjusted during the iteration process to accelerate convergence.

4. The method according to claim 1, wherein In the multi-group strategy, any individual can belong to up to 8 different groups. The group members are selected based on their fitness and are located within a range with the individual as the center and the distance to the optimal individual as the radius.

5. The method according to claim 1, wherein The jump mechanism based on the standard Cauchy distribution acts on each dimension coordinate of the current best individual in each iteration, and realizes position update by introducing Cauchy distribution random perturbation.

6. The method according to claim 1, characterized in that During the feature selection process, each continuous vector element is selected when it is greater than the threshold θ, and is eliminated when it is less than or equal to θ. Support vector machine (SVM), K-nearest neighbor (KNN), and random forest (RF) are used as classifiers for evaluation.

7. The method according to claim 1, characterized in that During the optimization process, the maximum number of iterations was set to 1000, the population size was set to 20, and all compared algorithms were run 10 times to determine the optimal feature subset.

8. The method according to claim 1, characterized in that The MSAFA method was evaluated for feature selection on six datasets including MW1, KC3, PC1, KC1, PC2, and MC1 in the NASA repository, and the area under the curve (AUC) was used as the performance evaluation indicator.