A Symbolic Regression Method Based on Federated Genetic Programming
Through the federal genetic programming method that processes data locally on the client and adopts the mean drift aggregation mechanism, data privacy and security issues and insufficient search performance in distributed genetic programming are solved, and efficient symbol regression and data protection are achieved.
Patent Information
- Application Number
- CN202210366425.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-08
AI Technical Summary
Existing genetic programming algorithms fail to effectively consider data privacy and security issues when processing distributed data, and prediction errors cannot effectively guide model optimization when the data samples do not sufficiently cover the input space.
A symbol regression method based on federal genetic programming is proposed. By processing data in parallel locally on the client, the original data transmission is avoided, combined with the mean drift aggregation mechanism, the local fitness is aggregated, and the fitness function is adjusted through weights, improving the effect of symbol regression.
It realizes training of global models without centralizing data, protecting data privacy and security, reducing data acquisition time, and improving the search performance of genetic programming algorithms, and is suitable for symbol regression problems in real environments.
Smart Images

Figure CN114840873B_ABST
Abstract
Description
Technical Field
[0001] The specific technical field involved in the present invention is intelligent computing and high-performance computing, and particularly relates to a symbolic regression method based on federated genetic programming. Background Art
[0002] With the popularization of intelligence, various edge devices have become an essential part of life, such as smart phones, smart computers, smart appliances, etc. Various data are stored dispersedly in each device. If the data is centrally stored in a server, it will bring great security risks during the transmission process, and the communication overhead is huge. At present, the security of the cyberspace has a great impact on individuals and even the whole country. How to design a machine learning framework using the data of these edge devices on the premise of protecting data privacy and security is the focus of current research.
[0003] In recent years, due to the poor interpretability of deep learning models and high requirements for hardware, more and more researchers have started to focus on interpretable machine learning, making symbolic regression a hot topic. The genetic programming (GP) algorithm is the mainstream method for studying symbolic regression problems currently. The principle of genetic programming is to optimize the non-linear tree-structured program, that is, the chromosome in the genetic algorithm, and at the same time, the chromosome needs to be parsed. Currently, genetic programming is widely used in fields such as pattern recognition, image analysis, and symbolic regression.
[0004] However, the existing genetic programming algorithms have the following deficiencies: on the one hand, current technologies do not consider data privacy and data security issues from the dimension of data. On the other hand, the current genetic programming search is purely driven by the prediction errors observed on the training data samples. When the data samples cannot fully cover the input space, the prediction errors cannot provide sufficient guidance for the desired model (A large-scale symbolic regression method and system based on an adaptive parallel genetic algorithm). Summary of the Invention
[0005] Starting from protecting the privacy and security of data, to solve the technical problems not considered in distributed genetic programming. The present invention proposes a symbolic regression method based on federated genetic programming, which can train a global model without centralizing data. Each client can process local data in parallel locally without sending the original data to the server. This method not only protects the privacy and security of data, but also reduces the time for data collection. In addition, a mean shift aggregation mechanism is also proposed to aggregate local fitness. Considering the relative importance of samples, this mechanism studies the possibility of improving the symbolic regression of real data by incorporating weights into the fitness function.
[0006] The object of the present invention is achieved by at least one of the following technical solutions.
[0007] A symbolic regression method based on federated genetic programming, comprising the following steps:
[0008] S1: Initialization: Create multiple threads, determine the number of clients accessing the server, and ensure successful client access to the server; randomly initialize the population, and the population size is NP;
[0009] S2: Client fitness calculation: Multiple clients perform fitness calculation of the population in parallel, determine whether the fitness value reaches the termination condition, if so, exit, otherwise execute step S3;
[0010] S3: Server fitness aggregation: Perform fitness aggregation according to the mean shift aggregation mechanism to obtain the aggregated population fitness F;
[0011] S4: Gene selection: The process of selecting individuals according to the aggregated population fitness F, and the selected individuals will be used as the paternal line in the future to breed the next-generation program individuals through genetic operators;
[0012] S5: Gene mutation: Genes are randomly changed to new values with a certain probability;
[0013] S6: Gene crossover: Each gene crosses with the mutation vector to generate a population, and return to step S2.
[0014] Further, in step S1, a symbolic regression system for symbolic regression is constructed. The symbolic regression system includes multiple clients and a central server, i.e., the server. The server sends the population to the clients, and the clients calculate the fitness based on their own data and return it to the server. Neither of the data transmitted between the two parties is the original data, which solves the problem of data non-sharing in the privacy environment.
[0015] Further, in step S1, start the server and multiple clients; the server monitors in real time whether there is a server end applying for access or a connected client that needs to be disconnected. When a new client requests access, the server immediately responds to the client's access; after all clients are successfully connected, perform population initialization on the server; in the server, the server confirms the IP and port of the connected clients, and then uniformly sends the initial population to the clients;
[0016] The population initialization in the server refers to generating NP random chromosomes to form the initial population, which is specifically expressed as follows:
[0017] X = {X i |X i = [x i,1 ,x i,2 ,...,xi,L , i = 1, 2, ..., NP} (1)
[0018] Among them, X i is the vector representing the i-th chromosome, i is the index of the chromosome in the population, and x i,j is the j-th element of the i-th chromosome X i . L is the length of the chromosome, and NP represents the population size; each chromosome includes a main program and multiple sub-functions, and both the main program and the sub-functions are composed of gene expressions at the head and the tail;
[0019] In the client, before startup, confirm the IP address and port number of the server to be connected. After successfully connecting to the server, wait for the server to send the population for fitness calculation.
[0020] Furthermore, in step S2, after the client obtains the population, each chromosome in the population is encoded into an expression equal in length to the chromosome; assume that the data sets of all clients are represented as follows:
[0021] D = {D1, ..., D k , ..., D K} (2)
[0022] Among them, D k represents the data of the k-th client connected to the server, k = 1 ~ K, and K is the number of clients connected to the server; after chromosome encoding and calculation, the fitness f of the entire population is obtained, which is expressed as follows:
[0023]
[0024] Among them, NP represents the population size, and f k (X i ) represents the fitness value calculated by the i-th chromosome in the population on the k-th client, i = 1 ~ NP.
[0025] Furthermore, in step S3, using the mean shift aggregation mechanism, each chromosome aggregates multiple fitness values according to the importance of each client. The specific algorithm of the mean shift aggregation mechanism is as follows:
[0026] S3.1: Initialize the aggregated population fitness F = 0, and obtain a random center point x;
[0027] S3.2: Input the kernel bandwidth h, the aggregation termination distance s d , the entire population fitness f, and the client weights W = [w1, w2, ..., w k ;
[0028] S3.3: Calculate all the distances from the fitness f of the entire population to the random center point x, and then find all the points within the kernel bandwidth h, which is called the set M;
[0029] S3.4: Calculate the vectors from the random center point x to each point in the set M, and add up all the vectors to get M h (x);
[0030] S3.5: The random center point x moves along the direction of M h (x), and the center point becomes x' = x + ||M h (x)||;
[0031] S3.6: Loop through steps S3.3 - S3.5 until |M h (x)|| < s d , and execute step S3.7;
[0032] S3.7: Output the aggregated population fitness F;
[0033] The kernel bandwidth h in the mean - shift aggregation mechanism algorithm is an important parameter of the Gaussian kernel function, and different values result in different aggregation effects; the weight W of the client is calculated according to the percentage of the client's data volume in the total data volume of all clients.
[0034] Furthermore, the specific calculation formula of M h (x) is as follows:
[0035]
[0036] Among them, x i represents the i - th chromosome in the population, w k represents the weight of the k - th client, represents the Gaussian kernel function.
[0037] Furthermore, in step S4, based on the aggregated population fitness F = {f c (X1),..., f c (X i ),..., f c (X NP )} obtained in step S3, select offspring to replace the chromosomes of the parent generation to form a new population, specifically as follows:
[0038]
[0039] Among them, f(U i ) represents the fitness of the parent - generation chromosome U i , and the parent - generation chromosome represents the chromosome of the previous round of training, f c (X i ) represents the fitness of the i - th chromosome Xi Aggregated fitness value.
[0040] Further, in step S5, based on the traditional DE mutation scheme "DE / current-to-best / 1", the genes in the chromosome are randomly changed to new values with a certain probability, specifically as follows:
[0041] Y i = X i + β(X best - X i ) + β{X r1 - X r2}(5)
[0042] Where Y i represents the mutation vector of the i-th chromosome X i in the population, X best is the best individual in the population, X r1 , X r2 and X i are three different individuals respectively, X r1 and X r2 are randomly selected from the population; β is a scaling factor with a value of rand(0,1).
[0043] Further, in step S6, each element in the i-th chromosome X i in the population crosses with each element of the mutation vector Y i to create a new trial vector, enabling the population to search for better solutions in the solution space; each gene in the i-th chromosome X i in the population creates a new trial vector Z i through the mutation vector Y i , specifically as follows:
[0044]
[0045] Where z i,j , y i,j and x i,j represent the j-th elements of the trial vector Z i , the mutation vector Y i and the chromosome X i respectively; CR represents the crossover probability with a value of rand(0,1); l is a random integer between 1 and L, where L is the length of the chromosome;
[0046] After the crossover operation, a new population is generated; the new population is sent to the client, and step S2 is returned.
[0047] Further, in step S2, the root mean square error (RMSE) in symbolic regression is used for fitness calculation. Given a set value of the root mean square error (RMSE), when the population fitness f is less than the set value, the termination condition is reached and the symbolic regression is completed.
[0048] Compared with the prior art, the advantages of the present invention are as follows:
[0049] (1) For the existing distributed GP technology, the present invention can protect the privacy and security of data by training a global model through federated learning. At the same time, the clients in the present invention have absolute freedom and can enter or exit the entire system at any time, which is more in line with the application scenarios in the real environment.
[0050] (2) The present invention further improves the search performance of the genetic programming algorithm by adopting a mean shift aggregation method, and also considers the different weights given to different degrees of importance of data samples, thereby effectively solving the symbolic regression problem in the real environment.
[0051] (3) The symbolic regression method of the present invention can make full use of data information and has better effects compared with the traditional gene programming algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 The algorithm framework diagram of a symbolic regression method based on federated genetic programming in an embodiment of the present invention;
[0053] Figure 2 The schematic diagram of chromosome coding in an embodiment of the present invention;
[0054] Figure 3 The schematic diagram of symbolic regression solved in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following examples are given in conjunction with the accompanying drawings to describe the specific implementation of the present invention in detail.
[0056] Example 1:
[0057] The main purpose of this work is to solve the symbolic regression problem when data is scattered on different local machines and is not allowed to be transmitted to the central server. At the same time, the data distribution of each client does not cover the entire sample space, and the amount of data of each client is different. When each client trains the model alone, each client can train multiple different function expressions, which are far from the approximate function. As Figure 3 shown, the present invention proposes a federated training method to jointly train the data of multiple clients, and finally the desired target function can be obtained.
[0058] A symbolic regression method based on federated genetic programming, as Figure 1 shown, includes the following steps:
[0059] S1: Initialization: Create multiple threads, determine the number of clients accessing the server, and ensure successful client access to the server; randomly initialize the population, with the population size being NP;
[0060] Construct a symbolic regression system for symbolic regression. The symbolic regression system includes multiple clients and a central server, i.e., the server. The server sends the population to the clients, and the clients calculate the fitness based on their own data and return it to the server. Neither the data transmitted by both parties is the original data, solving the problem of data non - sharing in a privacy environment.
[0061] Start the server and multiple clients; the server monitors in real - time whether there is a server - side application for access or a connected client that needs to be disconnected. When a new client requests access, the server immediately responds to the client's access; after all clients are successfully connected, perform population initialization on the server; in the server, the server confirms the IP and port of the connected clients, and then uniformly sends the initial population to the clients;
[0062] The population initialization in the server means generating NP random chromosomes to form the initial population, which is specifically expressed as follows:
[0063] X = {X i |X i = [x i,1 , x i,2 ,..., x i,L , i = 1, 2,..., NP} (1)
[0064] Among them, X i is the vector representing the i - th chromosome, i is the index of the chromosome in the population, x i,j is the j - th element of the i - th chromosome X i , L is the length of the chromosome, and NP represents the population size; each chromosome includes a main program and multiple sub - functions, and both the main program and the sub - functions are composed of gene expressions at the head and the tail, as Figure 2 shown;
[0065] In the client, confirm the IP address and port number of the server to be connected before starting. After successfully connecting to the server, wait for the server to send the population for fitness calculation.
[0066] S2: Client fitness calculation: Multiple clients perform parallel fitness calculations on the population, determine whether the fitness value reaches the termination condition. If so, exit; otherwise, execute step S3;
[0067] After the client obtains the population, each chromosome in the population is encoded as an expression equal in length to the chromosome; assume that the data sets of all clients are represented as follows:
[0068] D = {D1,..., D k ,..., D K} (2)
[0069] where D k represents the data of the k-th client connected to the server, k = 1 ~ K, and K is the number of clients connected to the server; after chromosome encoding and calculation, the fitness f of the entire population is obtained, which is expressed as follows:
[0070]
[0071] where NP represents the population size, and f k (X i ) represents the fitness value calculated by the i-th chromosome in the population on the k-th client, i = 1 ~ NP.
[0072] The root mean square error (RMSE) in symbolic regression is used for fitness calculation. Given the set value of the root mean square error (RMSE), when the population fitness f is less than the set value, the termination condition is reached and the symbolic regression is completed.
[0073] S3: Server fitness aggregation: Aggregate the fitness according to the Mean shift aggregation mechanism to obtain the aggregated population fitness F;
[0074] Using the Mean shift aggregation mechanism, each chromosome aggregates multiple fitness values according to the importance of each client. The specific algorithm of the Mean shift aggregation mechanism is as follows:
[0075] S3.1: Initialize the aggregated population fitness F = 0 and obtain a random center point x;
[0076] S3.2: Input the kernel bandwidth h, the aggregation termination distance s d , the entire population fitness f, and the client weights W = [w1, w2,..., w k ;
[0077] S3.3: Calculate all the distances from the entire population fitness f to the random center point x, and then find all the points within the kernel bandwidth h, which is called the set M;
[0078] S3.4: Calculate the vectors from the random center point x to each point in the set M, and add all the vectors to get M h (x);
[0079] S3.5: The random center point x moves along M hMove in the direction of (x), and the center point becomes x' = x + ||M h (x)||;
[0080] S3.6: Loop through steps S3.3 - S3.5 until |M h (x)|| < s d , and execute step S3.7;
[0081] S3.7: Output the aggregated population fitness F;
[0082] The kernel bandwidth h in the mean - shift aggregation mechanism algorithm is an important parameter of the Gaussian kernel function, and different values result in different aggregation effects; the weight W of the client is calculated according to the percentage of the client's data volume in the total data volume of all clients.
[0083] M h (x) The specific calculation formula is as follows:
[0084]
[0085] Among them, x i represents the i - th chromosome in the population, w k represents the weight of the k - th client, represents the Gaussian kernel function.
[0086] S4: Gene selection:
[0087] Based on the aggregated population fitness F = {f c (X1),..., f c (X i ),..., f c (X NP )} obtained in step S3, select offspring to replace the chromosomes of the parent generation to form a new population, specifically as follows:
[0088]
[0089] Among them, f(U i ) represents the fitness of the parent chromosome U i , and the parent chromosome represents the chromosome of the previous round of training, f c (X i ) represents the aggregated fitness value of the i - th chromosome X i .
[0090] S5: Gene mutation:
[0091] Based on the traditional DE mutation scheme "DE / current - to - best / 1", the genes in the chromosome are randomly changed to new values with a certain probability, specifically as follows:
[0092] Y i = X i + β(X best - X i ) + β{X r1 - X r2} (5)
[0093] Wherein, Y i represents the mutation vector of the i-th chromosome X i in the population, X best is the best individual in the population, X r1 , X r2 and X i are three different individuals respectively, X r1 and X r2 are randomly selected from the population; β is a scaling factor with a value of rand(0,1).
[0094] S6: Gene crossover:
[0095] Each element in the i-th chromosome X i in the population is crossed with each element of the mutation vector Y i to create a new trial vector, enabling the population to search for better solutions in the solution space; each gene in the i-th chromosome X i in the population creates a new trial vector Z i through the mutation vector Y i , which is specifically expressed as follows:
[0096]
[0097] Wherein, z i,j , y i,j and x i,j represent the j-th elements of the trial vector Z i , the mutation vector Y i and the chromosome X i respectively; CR represents the crossover probability with a value of rand(0,1); l is a random integer between 1 and L, where L is the length of the chromosome;
[0098] After the crossover operation is completed, a new population is generated; the new population is sent to the client, and step S2 is returned.
[0099] In this embodiment, in order to verify the performance of the algorithm framework of the present invention, it is first verified on 5 artificially customized standard data sets. The parameter settings of the algorithm of the present invention are: the population size is NP = 30, the maximum number of iterations is R = 20000, s d = 0.5, the kernel bandwidth h = 3, and the fitness value termination value is RMSE < 10 -4 .
[0100] Example 2:
[0101] In this embodiment, in order to further verify the effectiveness of the present invention, it was verified on 5 noise data sets. The parameter settings of the algorithm of the present invention are: the population size is NP = 50, the maximum number of iterations is R = 20000, s d = 0.5, the kernel bandwidth h = 3, and the termination value of the fitness value is RMSE < 10 -4 .
[0102] Example 3:
[0103] In this embodiment, finally, the present invention was verified on 2 real - scenario data sets. The parameter settings of the algorithm of the present invention are: the population size is NP = 50, the maximum number of iterations is R = 20000, s d = 0.5, the kernel bandwidth h = 3, and the termination value of the fitness value is RMSE < 10 -4 .
[0104] The final results of the three implementation cases all show that the present invention is significantly superior to the existing genetic programming algorithms in terms of RMSE and convergence speed of data sets in different environments. This shows that adopting the present invention can not only protect data information, but also improve the search ability of the genetic programming algorithm.
Claims
1. A symbolic regression method based on federated genetic programming, characterized in that, It includes the following steps: S1: Initialization: Create multiple threads, determine the number of clients accessing the server, and ensure successful client access to the server; randomly initialize the population with a population size of NP; S2: Client fitness calculation: Multiple clients perform fitness calculation of the population in parallel, judge whether the fitness value reaches the termination condition. If so, exit. Otherwise, execute step S3; S3: Server fitness aggregation: Perform fitness aggregation according to the Mean shift aggregation mechanism to obtain the aggregated population fitness F; adopt the Mean shift aggregation mechanism, and each chromosome aggregates multiple fitness values according to the importance of each client. The specific algorithm of the Mean shift aggregation mechanism is as follows: S3.1: Initialize the aggregated population fitness F = 0, and obtain a random center point x; S3.2: Input the kernel bandwidth h, the aggregation termination distance s d , the fitness f of the entire population, and the client weights W = [w1, w2,..., w k ; S3.3: Calculate all distances from the fitness f of the entire population to the random center point x, and then find all points within the kernel bandwidth h, which is called the set M; S3.4: Calculate the vectors from the random center point x to each point in the set M, and add all the vectors to obtain M h (x); M h (x) The specific calculation formula is as follows: where x i represents the i-th chromosome in the population, w k represents the weight of the k-th client, represents the Gaussian kernel function; S3.5: The random center point x moves along the direction of M h (x), and the center point becomes x' = x + ||M h (x)||; S3.6: Loop through steps S3.3 - S3.5 until |M h (x)||<s d , and execute step S3.7; S3.7: Output the aggregated population fitness F; The kernel bandwidth h in the Mean shift aggregation mechanism algorithm is an important parameter of the Gaussian kernel function, and different values result in different aggregation effects; the weight W of the client is calculated according to the percentage of the client data volume in the total data volume of all clients; S4: Gene selection: The process of selecting individuals according to the aggregated population fitness F. The selected individuals will be used as the paternal line and breed the next-generation program individuals through genetic operators; S5: Gene mutation: Genes are randomly changed to new values with a certain probability; S6: Gene crossover: Each gene crosses with the mutation vector to generate a population, and return to step S2.
2. The symbolic regression method based on federated genetic programming according to claim 1, wherein In step S1, a symbolic regression system for symbolic regression is constructed. The symbolic regression system includes multiple clients and a central server, i.e., the server side. The server side sends the population to the clients, and the clients calculate the fitness based on their own data and return it to the server side. Neither of the two parties transmits the original data, which solves the problem of data non-sharing in the privacy environment.
3. The symbolic regression method based on federated genetic programming according to claim 2, wherein In step S1, start the server and multiple clients; the server monitors in real time whether there is a server side applying for access or a connected client that needs to be disconnected. When a new client requests access, the server immediately responds to the client's access; after all clients are successfully connected, perform population initialization on the server; in the server, the server confirms the IP and port of the connected clients, and then uniformly sends the initial population to the clients; The population initialization in the server refers to generating NP random chromosomes to form the initial population, which is specifically expressed as follows: X = {X i | X i = [x i,1 , x i,2 ,..., x i,L , i = 1, 2,..., NP} (1) Among them, X i is a vector representing the i-th chromosome, where i is the index of the chromosome in the population, and x i,j is the j-th element of the i-th chromosome X i ; L is the length of the chromosome, and NP represents the population size; each chromosome includes a main program and multiple sub-functions, and both the main program and the sub-functions are composed of gene expressions at the head and the tail. In the client, confirm the IP address and port number of the server to be connected before starting. After successfully connecting to the server, wait for the server to send the population for fitness calculation.
4. A symbolic regression method based on federated genetic programming according to claim 1, characterized in that In step S2, after the client obtains the population, each chromosome in the population is encoded as an expression equal to the chromosome length; assume that the data sets of all clients are expressed as follows: D = {D1,..., D k ,..., D K}(2) Among them, D k represents the data of the k-th client connected to the server, where k = 1 to K, and K is the number of clients connected to the server; the fitness f of the entire population is obtained through chromosome coding and calculation, which is expressed as follows: Among them, NP represents the population size, and f k (X i ) represents the fitness value calculated by the i-th chromosome in the population on the k-th client, where i = 1 to NP.
5. A symbolic regression method based on federated genetic programming according to claim 1, characterized in that In step S4, based on the aggregated population fitness F = {f c (X1),..., f c (X i ),..., f c (X NP )} obtained in step S3, offspring are selected to replace the chromosomes of the parent generation to form a new population, specifically as follows: Among them, f(U i ) represents the fitness of the parental chromosome U i . The parental chromosome refers to the chromosome of the previous round of training. f c (X i ) represents the fitness value aggregated by the i-th chromosome X i .
6. The symbolic regression method based on federated genetic programming according to claim 1, wherein In step S5, based on the traditional DE mutation scheme "DE / current-to-best / 1", the genes in the chromosome are randomly changed to new values with a certain probability, which is specifically as follows: Y i = X i + β(X best - X i ) + β{X r1 - X r2}(5) Among them, Y i represents the mutation vector of the i-th chromosome X i in the population, X best is the best individual in the population, X r1 , X r2 and X i are three different individuals respectively, X r1 and X r2 are randomly selected from the population; β is a scaling factor, and its value is rand(0,1).
7. A symbolic regression method based on federated genetic programming according to claim 1, characterized in that In step S6, each element in the i-th chromosome X in the population i is crossed with each element of the mutation carrier Y i to create a new trial carrier, enabling the population to search for better solutions in the solution space; each gene in the i-th chromosome X in the population i creates a new trial vector Z i through the mutation carrier Y i , which is specifically expressed as follows: Among them, z i,j , y i,j and x i,j respectively represent the j-th element of the test vector Z i , the mutant vector Y i and the chromosome X i ; CR represents the crossover probability, taking the value of rand(0,1); l is a random integer between 1 and L, where L is the length of the chromosome; After the crossover operation, a new population is generated; the new population is sent to the client, and step S2 is returned.
8. A symbolic regression method based on federated genetic programming according to any one of claims 1 to 7, characterized in that In step S2, the root mean square error (RMSE) in symbolic regression is used for fitness calculation. Given the set value of the root mean square error (RMSE), when the population fitness f is less than the set value, the termination condition is reached and the symbolic regression is completed.
Citation Information
Patent Citations
Multi-phase orthogonal code generating method based on improved immune genetic algorithm
CN104376363A
Large-scale symbol regression method and system based on adaptive parallel genetic algorithm
CN110135584A
Waste mobile phone identification method based on differential evolution algorithm-deep forest algorithm
CN113298107A