Service Value Chain Business Data Feature Selection Method Based on Golden Jackal Optimization Algorithm
The key features in the service value chain business data are screened out through the multi-objective binary Jinzhai optimization algorithm, solving the problems of many redundant features and heavy computing load in the existing technology, and realizing the reduction of data dimensions and improvement of classification performance.
Patent Information
- Application Number
- CN202211637549.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-16
AI Technical Summary
The existing Jinzhai optimization algorithm is only suitable for single-objective optimization problems with continuous decision space, and cannot effectively solve the problem of multi-objective discrete optimization, especially in the selection of business data features in multiple service value chains, resulting in the problem of many redundant features and heavy computing load.
A multi-objective binary Jinzhai optimization algorithm is proposed. Through the construction of fitness function and non-dominant solution sets, combining clustering algorithms and KNN classifiers, key features are selected, redundant features are eliminated, and feature subsets are generated to optimize computing resources.
It achieves the maintenance of good classification performance while reducing the data dimensions, optimizes the use of computing resources, and improves the accuracy of the classifier.
Smart Images

Figure CN115858534B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method for selecting service value chain business data features based on the golden jackal optimization algorithm. Background Art
[0002] With the accelerating development of the digital economy, data has become an important production factor, and data-driven business operations have become a hot topic in the academic and industrial circles. Since the third-party cloud platform provides information services for multiple service value chains at the same time and connects multiple service value chains, the third-party cloud platform aggregates data from multiple value chains. By integrating multi-chain data, it is possible to effectively break the information islands existing between value chains, which has important economic significance for revitalizing the business resources within each chain and enhancing the stability of the entire system. Specifically, the cloud platform integrates business data such as maintenance, claims, and used parts accumulated in each service value chain, analyzes and mines the business data to discover the deficiencies in business operations, and provides decision-making support for managers.
[0003] Compared with processing the business data of a single value chain, the business data of multiple integrated service value chains has characteristics such as a large amount of data, redundant features, and a high dimension. The existence of redundant features not only contaminates the data set with irrelevant features, but sometimes also reduces the performance of the classifier, while increasing the computational load of the platform for processing data. Selecting data features and reducing the number of data features can save more computing resources, which is an important method for reducing the data computing load of the platform. Feature selection is an optimization problem. Compared with traditional feature selection methods, feature selection methods based on swarm intelligence algorithms have better performance. Through the feature selection method based on the swarm intelligence algorithm, the key features in the data are screened out, and redundant data features are removed. Through feature selection, while reducing the dimension of the data, good classification performance is still ensured, realizing the optimization of computing resources.
[0004] Currently, the golden jackal optimization algorithm in the swarm intelligence algorithm is an algorithm inspired by the hunting strategy of golden jackals, proposed by Nitish Chopra et al. in 2022. The individuals in the golden jackal population are monogamous, and adult golden jackals usually hunt in pairs and cooperate to hunt. The golden jackal optimization algorithm has better performance than other swarm intelligence algorithms, but the golden jackal algorithm is only applicable to single-objective optimization algorithms with continuous decision spaces and does not have the ability to solve multi-objective discrete optimization problems. Therefore, there is an urgent need for a feature selection method based on a multi-objective binary golden jackal optimization algorithm to solve multi-objective data feature selection problems. Summary of the Invention
[0005] The object of the present invention is to: in view of the fact that the current golden jackal algorithm is only applicable to single-objective optimization algorithms with continuous decision spaces and does not have the ability to solve multi-objective discrete optimization problems, a feature selection method for a multi-objective binary golden jackal optimization algorithm is proposed to solve the multi-objective data feature selection problem.
[0006] In order to achieve the above object of the invention, the present invention provides the following technical solutions:
[0007] A service value chain business data feature selection method based on the golden jackal optimization algorithm includes the following steps:
[0008] Step S1: Extract the required business data from a third-party cloud platform, and use the business data as the input set for fitness function calculation. The fitness function is two objective functions F = {f1, f2}, where f1 is the retention rate of the number of features in the original data set, and f2 is the classification performance of the currently selected data features;
[0009] Step S2: Initialize the parameters, including the number of prey populations N, the dimension M of individuals in the population, the total number of iterations mg, and a feasible solution of the objective function is X i = [x1, x2,..., x n , where n is the maximum dimension of the data set features, and where x i = 0 or 1, 0 indicates that the feature is not selected, and 1 indicates that the feature is selected. Solve the objective function based on different solutions;
[0010] Step S3: Generate an initial prey population, where each individual in the prey population is a solution of the objective function. The initial prey population is an N*M matrix, and each element in the matrix is randomly generated 0 or 1, 0 indicates that the feature is not selected, and 1 indicates that the feature is selected. The generation formula is expressed as Prey = randint([0, 1], N, M);
[0011] Step S4: Calculate the maleset and femaleset sets. The maleset is the set of male golden jackals, and the femaleset is the set of female golden jackals. The maleset set and the femaleset set guide the position update of the prey population. The maleset set and the femaleset set are constructed based on the non-dominated solution set. The maleset set is calculated based on a uniformly distributed vector, and the individual in the non-dominated solution set with the smallest angle with the uniformly distributed vector is set as the individual with the optimal fitness. The femaleset set obtains individuals with uniform distribution through a clustering algorithm, and the clustering center point obtained by clustering is set as the sub-optimal fitness individual;
[0012] Step S5: Update the prey position. Randomly select individuals from the maleset set and the femaleset set as male and female to guide the update of the prey position. When the prey's escape ability is greater than the hunting ability of the golden jackal, it is the exploration stage of the golden jackal for the prey. When the prey's escape ability is less than the hunting ability of the golden jackal, it is the encirclement stage of the golden jackal for the prey, Y M (g) is the position of the male golden jackal, Y FM (g) is the position of the female golden jackal, Y(g + 1) is the updated prey position, the The Y(g + 1) is a continuous value, and the population individuals are mapped into binary vectors through a transfer function
[0013] Step S6: Select the prey position. Prey(g) is the current prey position. Let PreyCandidate(g + 1) = Y(g + 1) ∪ Prey(g). Based on non - dominated sorting and crowding distance sorting, N individuals are selected from the PreyCandidate(g + 1) set as Prey(g + 1). Update g = g + 1 to determine whether the iteration ends. When the iteration g < mg, enter the next iteration and jump to step 4. When g = mg, the iteration ends. The non - dominated solutions of the objective function are the results of various data feature selections.
[0014] Further, the specific steps of step S4 are as follows
[0015] Step S41: The uniformly distributed vector is W = {w1, w2, w3}, and the target space is divided into two regions
[0016] Step S42: Perform Pareto non - dominated sorting on the target solution set to generate a non - dominated solution set. The non - dominated solution set is PF. When the number of individuals in the PF is more than 5, that is, |PF| > 5, enter step S43. When the number of individuals in the PF is 4 or 5, that is, 4 ≤ |PF| ≤ 5, enter step S44. When the number of individuals in the PF is less than 3 or equal to 3, that is, |PF| ≤ 3, enter step S45
[0017] Step S43: Since |PF| > 5, two individuals are selected. Among them, W is used as the basis for screening the maleset solution set, and the solution with a smaller angle between the PF and W is selected as the maleset solution set, maleset = {c1, c3, c5}. Based on the k - medoids algorithm, clustering is performed on the remaining elements, and the central elements of the two clusters are selected as the femaleset solution set
[0018] Step S44: When 4 ≤ |PF| ≤ 5, select male and female respectively based on the angles between the solutions in the PF and W = {w1, w2, w3}, and take the non-dominated solution set with the smallest angle with {w1, w2, w3} as male, and the rest as female;
[0019] Step S45: When |PF| ≤ 3, randomly select two solutions from the PF as male and female.
[0020] Further, {w1, w2, w3} are uniformly distributed vectors, w1 is (0, 1), w2 is (1, 0), w3 is (0.5, 0.5), the boundary solutions in the non-dominated solution set are retained, the boundary solutions are obtained through w1 and w2, and w3 is the solution located in the middle position of the solution distribution.
[0021] Further, f1 is the retention rate of the number of features in the original dataset, and the In the formula, FN is the number of data features in the original dataset, and NN is the number of data features after screening;
[0022] f2 is the classification performance of the currently selected data features, the classification performance is measured by the classification error rate, Acc is the classification accuracy of the current KNN classifier, the classification error rate obtained by using the KNN classifier to calculate the current feature subset is based on the feature subset x i Extract data from the dataset as the input of the KNN classifier, train and verify, and then obtain the classification error rate of the feature subset x i of, the where D is the number of data records in the validation dataset, A is the number of correctly classified ones in the validation dataset, the
[0023] The values of the objective function {f1, f2} are the basis for population selection, and non-dominated solutions are obtained based on this value.
[0024] Further, the prey escape ability is E, and the where r is any random number in the interval [0, 1], c1 is the constant 1.5, g is the current iteration number, mg is the total number of iterations, rl = 0.05 * LF(y) is used to prevent the solution process from falling into local optimum, in the formula LF is the flight function, the In the formula μ and v are random numbers uniformly distributed in [0, 1], and β takes 1.5;
[0025] When |E| ≥ 1, it is the exploration stage of the golden jackal for the prey, Y M (g) is the position of the male golden jackal, Prey(g) is the current prey position, Y(g + 1) is the updated prey position, YM (g)-E|Y M (g)-rl.Prey(g)| = Y1(g), Y FM (g) is the position of the female golden jackal, Y FM (g)-E|Y FM (g)-rl.Prey(g)| = Y2(g), the
[0026] When |E| < 1, it is the surrounding stage of the golden jackal to the prey, Y M (g) is the position of the male golden jackal, Prey(g) is the current prey position, Y(g + 1) is the updated prey position, Y M (g)-E|rl.Y M (g)-Prey(g)| = Y1(g), Y FM (g) is the position of the female golden jackal, Y FM (g)-E|rl.Y FM (g)-Prey(g)| = Y2(g), the
[0027] Furthermore, the service data is divided into a training data set and a validation data set. The training data set is used to train the KNN classifier, and the validation data set is used to verify the performance of the KNN classifier. The training data set and the validation data set are used as the input set for fitness function calculation.
[0028] Furthermore, the third-party cloud platform includes n service value chains The n service value chains form unified maintenance, claim, and used parts business data by integrating the maintenance, claim, and used parts business data included in each service value chain, extract the data that needs to be feature selected for normalization, and form the service business data after missing value supplementation data preprocessing.
[0029] Compared with the prior art, the beneficial effects of the present invention:
[0030] In the solution of this application, based on the feature selection algorithm of this patent, feature selection is performed on the service data A integrated by the platform, and a set of [x1, x2,..., x n vector set feature subsets are generated, where n is the maximum dimension of the data set features, and where x i= 0 or 1, indicating whether the current feature is selected to filter out the key features in the data and eliminate redundant data features. This algorithm is a two-objective optimization algorithm that can generate a set of feature subsets. The decision maker can select an optimization scheme for the feature subset according to the decision-making requirements, and then generate new business data B based on the selected feature subset scheme combined with business data A. At this time, business data B has a lower dimension than business data A. When using this business data for classification, due to the lower dimension of business data B and maintaining good classification performance, the optimization of computing resources is achieved. Description of the Drawings
[0031] Figure 1 It is a schematic diagram of the multi-chain data feature selection framework for the service value chain of a third-party cloud platform;
[0032] Figure 2 It is a schematic diagram of feature selection for the business data of a third-party cloud platform
[0033] Figure 3 It is a schematic diagram of the process of a feature selection method based on a multi-objective binary golden jackal optimization algorithm;
[0034] Figure 4 It is a schematic diagram of the solution methods for male individuals and female individuals;
[0035] Figure 5 It is to solve the male set solution set and female set solution set when |PF|≥5;
[0036] Figure 6 It is the error rate of the MOBGJO algorithm and the MOBPSO-CD algorithm when selecting the same number of features. Detailed Implementation Modes
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments.
[0038] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely represents some embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0039] It should be noted that, without conflict, the embodiments in the present invention and the features and technical solutions in the embodiments can be combined with each other.
[0040] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0041] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present invention is usually placed during use, or the orientation or positional relationship commonly understood by those skilled in the art. Such terms are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0042] Embodiment 1: Refer to Figures 1 - 6 as shown in
[0043] A method for selecting service value chain business data features based on the golden jackal optimization algorithm provided in this embodiment, refer to Figure 2 and Figure 3 as shown in, includes the following steps
[0044] Step S1: Extract the required business data from a third-party cloud platform, and use the business data as the input set for calculating the fitness function. The fitness function is two objective functions F = {f1, f2}, where f1 is the retention rate of the number of features in the original data set, and f2 is the classification performance of the currently selected data features.
[0045] Step S2: Initialize the parameters, including the number of prey populations N, the dimension M of individuals in the population, the total number of iterations mg, and a feasible solution of the objective function is X i =[x1, x2,..., x n , where n is the maximum dimension of the data set features, and where x i = 0 or 1, 0 indicates not selecting the feature, and 1 indicates selecting the feature. Solve the objective function based on different solutions.
[0046] Step S3: Generate an initial prey population, where each individual in the prey population is a solution of the objective function. The initial prey population is an N*M matrix, and each element in the matrix is randomly generated 0 or 1. 0 indicates not selecting the feature, and 1 indicates selecting the feature. The generation formula is expressed as Prey = randint([0, 1], N, M).
[0047] Step S4: Calculate the maleset and femaleset sets. The maleset is the set of male golden jackals, and the femaleset is the set of female golden jackals. The maleset and femaleset sets guide the position update of the prey population. The maleset and femaleset sets are constructed based on the non-dominated solution set. The maleset set is calculated based on a uniform distribution vector, and the individual in the non-dominated solution set with the smallest angle with the uniform distribution vector is set as the individual with the optimal fitness. The femaleset set obtains individuals with uniform distribution through a clustering algorithm, and the clustering center point obtained by clustering is set as the sub-optimal fitness individual;
[0048] Step S5: Update the prey position. Randomly select individuals from the maleset and femaleset sets as male and female to guide the update of the prey position. When the prey's escape ability is greater than the golden jackal's hunting ability, it is the exploration stage of the golden jackal for the prey. When the prey's escape ability is less than the golden jackal's hunting ability, it is the encirclement stage of the golden jackal for the prey. Y M (g) is the position of the male golden jackal, Y FM (g) is the position of the female golden jackal, Y(g + 1) is the updated prey position, and the The Y(g + 1) is a continuous value, and the population individuals are mapped to binary vectors through a transfer function
[0049] Step S6: Select the prey position. Prey(g) is the current prey position. Let PreyCandidate(g + 1) = Y(g + 1) ∪ Prey(g). Based on non-dominated sorting and crowding distance sorting, N individuals are selected from the PreyCandidate(g + 1) set as Prey(g + 1). Update g = g + 1 to determine whether the iteration ends. When the iteration g < mg, enter the next iteration and jump to step 4. When g = mg, the iteration ends. The non-dominated solution of the objective function is the result of various data feature selections.
[0050] In the solution of this application, based on the feature selection algorithm of this patent, feature selection is performed on the service data A integrated by the platform to generate a set of [x1, x2,..., x n vector set feature subsets, where n is the maximum dimension of the dataset features, and where x i = 0 or 1, indicating whether the current feature is selected to filter out the key features in the data and eliminate redundant data features. This algorithm is a multi-objective optimization algorithm that can generate a set of feature subsets. The decision maker can select an optimization scheme for the feature subset according to the decision-making requirements, and then generate new business data B based on the selected feature subset scheme combined with business data A. At this time, business data B has a lower dimension than business data A. When classifying using this business data, since business data B has a lower dimension and maintains good classification performance, the optimization of computing resources is achieved.
[0051] Furthermore, selecting male individuals and female individuals is a crucial step. In the original golden jackal algorithm, since the objective function is single-objective, the optimal and sub-optimal golden jackal positions can be directly obtained. The data feature subset in this patent is an optimization problem with two objectives. Due to the existence of two objectives, the optimal solution and sub-optimal solution cannot be directly calculated. In a multi-objective optimization algorithm, the diversity and convergence of the objective values obtained by the algorithm are important indicators for measuring the performance of the algorithm. To ensure that this algorithm has good diversity and convergence, the method of Pareto non-dominated sorting is used to ensure convergence, and the diversity of the solutions obtained by the algorithm is maintained by presetting 3 two-dimensional uniformly distributed vectors and clustering the non-dominated solutions. The specific steps are as Figure 4 shown, and the specific steps in step S4 are as follows:
[0052] Step S41: The uniformly distributed vector is W = {w1, w2, w3}, and the objective space is divided into two regions. See Figure 5 part a of.
[0053] Step S42: Perform Pareto non-dominated sorting on the objective solution set to generate a non-dominated solution set, which is PF and is used to ensure the convergence of the population. When the number of individuals in PF is more than 5, that is, |PFλ| > 5, enter step S43. When the number of individuals in PF is 4 or 5, that is, 4 ≤ |PF| ≤ 5, enter step S44. When the number of individuals in PF is less than 3 or equal to 3, that is, |PF| ≤ 3, enter step S45;
[0054] Step S43: Since |PF| > 5, two individuals are selected. Among them, W is used as the basis for screening the maleset solution set. See Figure 5 part b of. The solution in the non-dominated solution set with a smaller angle with W is selected as the maleset solution set, maleset = {c1, c3, c5}. Based on the k-medoids algorithm, the remaining non-dominated solutions are clustered, and two cluster center elements are selected as the femaleset solution set;
[0055] Step S44: When 4 ≤ |PF| ≤ 5, select male and female respectively based on the angles between the non-dominated solutions and W = {w1, w2, w3}, and take the non-dominated solution set with the smallest angle with {w1, w2, w3} as male, and the rest as female;
[0056] Step S45: When |PF| ≤ 3, randomly select two solutions from the non-dominated solution set as male and female.
[0057] Further, {w1, w2, w3} are uniformly distributed vectors, w1 is (0, 1), w2 is (1, 0), w3 is (0.5, 0.5), the boundary solutions in the non-dominated solution set are retained, the boundary solutions are obtained through w1 and w2, and w3 is used to obtain the solutions in the middle position of the solution distribution.
[0058] Further, f1 is the retention rate of the number of features in the original data set, the where FN is the number of data features in the original data set, and NN is the number of data features after screening;
[0059] f2 is the classification performance of the currently selected data features, the classification performance is measured by the classification error rate, Acc is the classification accuracy rate of the current KNN classifier, the classification error rate obtained by calculating the current feature subset using the KNN classifier, based on the feature subset x i Extract data from the data set as the input of the KNN classifier, train and verify to obtain the classification error rate of the feature subset x i of, the where D is the number of data records in the validation data set, A is the number of correctly classified ones in the validation data set, the
[0060] The values of the objective function {f1, f2} are the basis for population selection, and non-dominated solutions are obtained based on this value.
[0061] Further, the prey escape ability is E, the where r is any random number in the interval [0, 1], c1 is the constant 1.5, g is the current iteration number, mg is the total number of iterations, rl = 0.05 * LF(y) is used to prevent the solution process from falling into local optimum, where LF is the flight function, the where μ and v are random numbers uniformly distributed in [0, 1], and β takes 1.5;
[0062] When |E| ≥ 1, it is the exploration stage of the golden jackal for the prey, Y M (g) is the position of the male golden jackal, Prey(g) is the current prey position, Y(g + 1) is the updated prey position, YM (g)-E|Y M (g)-rl.Prey(g)| = Y1(g), YF M (g) is the position of the female golden jackal, Y FM (g)-E|Y FM (g)-rl.Prey(g)| = Y2(g), the
[0063] When |E| < 1, it is the encirclement stage of the golden jackal towards the prey, Y M (g) is the position of the male golden jackal, Prey(g) is the current prey position, Y(g + 1) is the updated prey position, Y M (g)-E|rl.Y M (g)-Prey(g)| = Y1(g), Y FM (g) is the position of the female golden jackal, Y FM (g)-E|rl.Y FM (g)-Prey(g)| = Y2(g), the
[0064] Furthermore, the service data is divided into a training data set and a validation data set. The training data set is used to train the KNN classifier, and the performance of the validation KNN classifier is evaluated. The training data set and the validation data set are used as the input set for fitness function calculation.
[0065] Furthermore, for the multi-chain data feature selection framework based on the service value chain of the third-party cloud platform, refer to Figure 1 As shown, the third-party cloud platform aggregates n service value chains The platform aggregates the data. By integrating the maintenance, claim, and used parts business data in each service value chain, the maintenance, claim, and used parts business data in a globally unified mode is formed. By configuring the globally integrated data, the source data "business data A" that needs to perform feature selection is extracted. After data preprocessing such as data normalization and missing value supplementation, an available data set is formed. The new feature set is screened using the feature extraction method, and a data set containing only new features is generated using the new feature set. This data set will be used to construct a higher-performance classifier to feed back to the service business and support decision-making problems in the operation of the service value chain.
[0066] The effects of the present invention can be illustrated by comparative experiments. This experiment uses the business operation data of multi-chain service providers, which includes a total of 16 features, including: claim review passing rate, claim working hours, timeliness of claim reports, maintenance working hours, timeliness rate of maintenance reports, passing rate of request reports review, passing rate of on-site reviews, passing rate of on-site report reviews, etc. The population size is set to 50, the maximum number of iterations is 100, and knn is selected as the classifier, with the value of k taken as 5. In the experiment, a method for feature selection of multi-chain data in the service value chain using a multi-objective binary discrete golden jackal algorithm is abbreviated as (MOBGJO).
[0067] An experimental comparison is made between MOBGJO and the multi-objective binary particle swarm optimization algorithm based on crowding distance (MOBPSO-CD) to verify that the MOBGJO algorithm can effectively select key features, eliminate redundant features, and further improve the classification accuracy rate. Figure 6 It shows the Pareto front results obtained by each algorithm. Figure 6 It shows that MOBGJO has a lower error rate than MOBPSO-CD when selecting the same number of features.
[0068] The above embodiments are only used to illustrate the present invention and do not limit the technical solutions described in the present invention. Although this specification has described the present invention in detail with reference to the above respective embodiments, the present invention is not limited to the above specific implementation manners. Therefore, any modification or equivalent replacement made to the present invention; and all technical solutions and their improvements that do not depart from the spirit and scope of the invention are covered by the scope of the claims of the present invention.
Claims
1. A method for selecting service value chain business data features based on the golden jackal optimization algorithm, characterized in that: including the following steps, Step S1: Extract the required business data from a third-party cloud platform, and use the business data as the input set for fitness function calculation. The fitness function is two objective functions , the is the retention rate of the number of features in the original data set, and the is the classification performance of the currently selected data features; Step S2: Initialize parameters, including the number of prey populations N, the dimension M of individuals in the population, and the total number of iterations mg. A feasible solution of the objective function is , where n is the maximum dimension of the dataset features, and where = 0 or 1, 0 indicates not selecting this feature, and 1 indicates selecting this feature; Step S3: Generate an initial prey population, where each individual in the prey population is a solution to the objective function, and the initial prey population is N*M a matrix of, where each element in the matrix is randomly generated as 0 or 1, 0 indicates not selecting the feature, and 1 indicates selecting the feature. The generation formula is expressed as ; Step S4: Calculate the maleset and femaleset sets. The maleset is the set of male golden jackals, and the femaleset is the set of female golden jackals. The maleset set and the femaleset set guide the position update of the prey population. The maleset set and the femaleset set are constructed based on the non-dominated individual set. The maleset set is calculated based on the uniform distribution vector, and the non-dominated individual with the smallest angle with the uniform distribution vector is set as the individual with the optimal fitness. The femaleset set screens the remaining non-dominated individuals through a clustering algorithm to obtain evenly distributed individuals, and the clustering center point obtained by clustering is set as the sub-optimal fitness individual; Step S5: Update the prey position. Randomly select individuals from the maleset set and the femaleset set as male and female to guide the update of the prey position. When the prey's escape ability is greater than the hunting ability of the golden jackal, it is the exploration stage of the golden jackal for the prey. When the prey's escape ability is less than the hunting ability of the golden jackal, it is the encirclement stage of the golden jackal for the prey. is the position of the male golden jackal, is the position of the female golden jackal, is the updated prey position, the ; the is a continuous value, and the population individuals are mapped into binary vectors through the transfer function Step S6: Select the prey location, which is the current prey location. Set , and screen out N individuals from as based on non-dominated sorting and crowding distance sorting. Update g = g + 1 to determine whether the iteration ends. When the iteration g < mg, jump to step 4 to enter the next iteration. When g = mg, the iteration ends. The non-dominated solutions of the objective function are the results of various data feature selections; The third-party cloud platform includes n service value chains , n By integrating the data of maintenance, claims, and used parts business included in each service value chain, a unified set of data for maintenance, claims, and used parts business is formed. After preprocessing the data that requires feature selection, including normalization and missing value supplementation, service business data is formed.
2. The method for selecting service value chain business data features based on the golden jackal optimization algorithm according to claim 1, wherein: The specific steps of step S4 include the following steps, Step S41: The uniform distribution vector is , which divides the target space into two regions; Step S42: Perform Pareto non-dominated sorting on the target solution set to generate a non-dominated solution set, where the non-dominated solution set is PF. When the number of individuals in the PF is more than 5, i.e., then proceed to step S43. When the number of individuals in the PF is 4 or 5, i.e., then proceed to step S44. When the number of individuals in the PF is less than 3 or equal to 3, i.e., then proceed to step S45; Step S43: Since , two individuals are screened out, where is used as the basis for screening the maleset solution set. Among the solutions in the PF with a smaller angle with , the solutions are selected as the maleset solution set. , based on the k-medoids algorithm, the remaining elements are clustered, and the central elements in the two clusters are screened out as the femaleset solution set; Step S44: When occurs, select male and female respectively based on the angles between the solutions in the PF and . Take the non-dominated solution set with the smallest angle with as male, and the rest as female; Step S45: When is true, randomly select two solutions from the PF as male and female.
3. The method for selecting service value chain business data features based on the golden jackal optimization algorithm according to claim 2, wherein: The said is a uniformly distributed vector, is (0, 1), is (1, 0), is (0.5, 0.5), and the boundary solutions in the non-dominated solution set are retained. The boundary solutions are obtained through and to obtain the boundary solutions, is to obtain the solutions located in the middle position of the solution distribution.
4. The method for selecting service value chain business data features based on the golden jackal optimization algorithm according to claim 1, wherein: The is the retention rate of the number of features in the original dataset, and the , where FN is the number of data features in the original dataset and NN is the number of data features after screening.
5. The method for selecting service value chain business data features based on the golden jackal optimization algorithm according to claim 4, wherein: The classification performance of the currently selected data feature, where the classification performance is measured by the classification error rate, Acc is the classification accuracy of the current KNN classifier, and the classification error rate obtained by calculating the current feature subset using the KNN classifier is based on the feature subset Data is extracted from the dataset as the input of the KNN classifier, trained and verified, and then the feature subset is obtained, and the classification error rate of the is calculated, where D is the number of data records in the validation dataset, A is the number of correctly classified data in the validation dataset, and the ; The objective function value is the basis for population selection, and non-dominated solutions are solved based on this value.
6. The method for selecting service value chain business data features based on the golden jackal optimization algorithm according to claim 1, wherein: The prey escape ability is , the , where r is any random number in the interval [0, 1], is the constant 1.5, is the current iteration number, is the total number of iterations, is used to prevent the solution process from falling into a local optimum. In the formula, is the flight function, the , in the formula , and are random numbers uniformly distributed in [0, 1], takes 1.
5.
7. The method for selecting service value chain business data features based on the golden jackal optimization algorithm according to claim 6, wherein: It is the exploration stage of the golden jackal for its prey, is the position of the male golden jackal, is the current prey position, is the updated prey position, , is the position of the female golden jackal, , the said .
8. The method for selecting service value chain business data features based on the golden jackal optimization algorithm according to claim 7, wherein: The is the stage of the golden jackal surrounding the prey, is the position of the male golden jackal, is the current prey position, is the updated prey position, , is the position of the female golden jackal, , the .
9. The method for selecting service value chain business data features based on the golden jackal optimization algorithm according to claim 1, wherein: The service data is divided into a training data set and a validation data set. The training data set is used to train the KNN classifier, and the validation data set is used to verify the performance of the KNN classifier. The training data set and the validation data set are used as the input set for fitness function calculation.
Citation Information
Patent Citations
Analysis method and device for network static business
CN106888465A
Method and system for automatically sharing procedural knowledge
US20210256869A1