An adaptive symbolic regression method based on multi-task genetic programming algorithm
By adopting multi-task genetic programming algorithm and adaptive computing resource allocation mechanism in symbol regression, the problem of inefficient multi-task symbol regression in the existing technology is solved, and more efficient search and resource utilization are achieved.
Patent Information
- Application Number
- CN202211061344.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-01
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-09-01
AI Technical Summary
When solving multiple symbol regression problems at the same time, the prior art is inefficient and lacks effective cross-reorganization operations and computing resource allocation mechanisms, resulting in low information migration and resource utilization among tasks.
Adaptive symbol regression method based on multitasking genetic programming algorithm is adopted to form a unified search space by constructing a scalable gene expression encoding method, and multi-factor evaluation and cross-recombination are realized. At the same time, an adaptive computing resource redistribution mechanism and an adaptive control strategy for cross-parameter IR are designed to optimize evolutionary behavior and resource allocation.
The search efficiency and result quality of multi-task symbol regression are improved, the correlation information between different tasks is effectively utilized, and the utilization rate of computing resources and the overall search efficiency of tasks are improved.
Smart Images

Figure CN115543556B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent computing and multi-task optimization, and in particular to an adaptive symbolic regression method based on a multi-task genetic programming algorithm. Background Art
[0002] Symbolic regression simultaneously searches for the optimal form of both the function and parameter set for a given problem. It is a powerful regression technique when prior knowledge of the model structure or data distribution is limited. As a supervised learning problem, symbolic regression aims to find a symbolic mathematical formula that establishes the relationship between given input and output variables. Its functional model structure and associated parameters are automatically discovered during the learning and evolutionary process, offering broad scalability and facilitating the discovery of hidden patterns in datasets that are often difficult to detect.
[0003] Genetic programming is currently the mainstream approach for solving symbolic regression problems. However, traditional genetic programming algorithms can only solve a single task in a single, independent run, which is inefficient when multiple tasks need to be solved simultaneously. In this context, multi-factor optimization has been proposed as a new evolutionary multitasking paradigm. It attempts to search for multiple tasks simultaneously in a single, independent run. In real-world applications, problems rarely exist in isolation, and they often contain useful information that, if properly leveraged, can enhance the problem-solving process when encountering another related problem. For example, if two problems happen to share a common global optimum, solving one problem simultaneously leads to solving the other. A practical example in industrial applications is cloud services. With the rapid development of cloud computing, current cloud services face the challenge of simultaneously accepting multiple optimization tasks from multiple users. These tasks can have similar properties or belong to completely different domains. Therefore, the need for cloud computing to concurrently handle the combined requirements of multiple clients is a major driving force behind the rapid rise of multitasking optimization.
[0004] Multifactorial optimization (MFO) is an optimization paradigm that applies a multitasking optimization framework to evolutionary computation, targeting a single population for evolutionary multitasking. Its primary goal is to simultaneously perform evolutionary searches across multiple concurrent search spaces corresponding to different optimization tasks. Because MFO maintains a single population for multiple tasks, useful information from different tasks can be implicitly transferred between individuals through a process called crossover recombination.
[0005] Currently, practical research on solving multiple symbolic regression problems simultaneously is still in its early stages, and the use of a multi-task paradigm to search across multiple problems remains in its early application stages. The specific operational steps are still understudied, and effective guidance for cross-combination operations between individual tasks is lacking. When using genetic programming algorithms based on a multi-task optimization paradigm to solve symbolic regression problems, there is a lack of proper management of the allocation of computational resources between different tasks during the concurrent search process. Given a relatively fixed resource allocation, balancing the computational resources allocated to each task to achieve the best possible results for all tasks has always been a key consideration. Summary of the Invention
[0006] The purpose of the present invention is to solve the above-mentioned defects in the prior art and to provide an adaptive symbolic regression method based on a multi-task genetic programming algorithm.
[0007] The purpose of the present invention can be achieved by taking the following technical solutions:
[0008] An adaptive symbolic regression method based on a multi-task genetic programming algorithm comprises the following steps:
[0009] S1. Construct a scalable gene expression encoding to form a unified search space for uniformly encoding multiple different symbolic regression problems; initialize all individuals in the population according to this encoding; and perform multi-factor evaluation on each individual across all tasks by converting the unified scalable gene expression encoding representation into a task-specific chromosome encoding decoding operation;
[0010] S2. According to the crossover parameter IR, the population is subjected to a selective crossover operation or a guided mutation operation, and the above operation is performed on each individual in the population one by one to produce a new generation of offspring population;
[0011] S3. Perform a hybrid join operation on the newly generated offspring population and the parent population to form a temporary population. Multi-factor evaluation and scalar fitness calculation are then re-performed on the temporary population. The hybrid join operation includes a join operation and a selection operation. A one-to-one selection is performed between the parent and offspring individuals based on scalar fitness to select the individuals that will enter the next generation population. An adaptive computing resource reallocation mechanism is constructed to ensure that individuals from populations with relatively poor performance enter the next generation population.
[0012] S4. Count the operation sources of the best individuals for each task in the next generation population, and adopt an adaptive control strategy for the cross parameter IR to adjust the evolutionary behavior;
[0013] S5. Determine whether the precision conditions or the number of fitness evaluations specified for all tasks have reached the specified maximum number. If so, terminate the evolution and output the solution of the best individual corresponding to each task in the final population; otherwise, continue to return to the evolution process from step S2 to step S5 to perform a cyclic evolutionary search of the population.
[0014] Furthermore, the scalable gene expression encoding method in step S1 is as follows: (1) a fixed-length string is used as an individual chromosome; (2) a main function and multiple automatically defined functions (ADFs) are included; and (3) integers are uniformly used to represent functions and terminals, and a single integer is used to represent various symbols of different tasks. In this way, the unique function and terminal symbols that may have different tasks can be encoded in the same encoding method, forming a unified public search space, which facilitates information transfer operations between different tasks.
[0015] Furthermore, the decoding process of converting the unified chromosome encoding representation into the chromosome of a specific task in step S1 is specifically as follows:
[0016] S11. Find the corresponding symbol type represented by each dimension by checking the range of the integer x on each dimension of the chromosome of the unified encoding method;
[0017] S12. Scaling the integer x according to the integer range of the type defined by the specific task to obtain an integer x' within the range defined by the specific task, that is:
[0018]
[0019] Among them, L x and U x are the minimum and maximum values within the numerical range of the corresponding type of the element found in step S11, N x is the maximum integer used in the corresponding type element, It is an operator that rounds a down to an integer and returns the largest integer not greater than a. The above steps map the integer x to 0 to N. x An integer x' between -1 is used as the index, and the scaled value x' is used to map x to a meaningful symbol for a specific task. The chromosome encoding structure in the public search space is decoded and converted into the chromosome encoding structure in the search space of a specific task, forming a solution specific to the task, which facilitates fitness evaluation on the task.
[0020] Furthermore, the selection crossover operation in step S2 is specifically as follows:
[0021] Use single-point crossover by adding randomly selected parent individuals X r1 and X iCrossover operation to generate offspring individuals U i , where X r1 represents a randomly selected individual, X i is the i-th individual in the population that performs single-point crossover; for the offspring individual U i Perform a unified mutation operation, and the generated offspring individuals U i The skill factor is inherited randomly from one of the parents with equal probability through vertical cultural transmission. Since the two parents may have their own independent skill factors, this crossover operation enhances the mutual transfer of useful genetic information found in different tasks.
[0022] Furthermore, the guided mutation operation in step S2 is specifically as follows:
[0023] The “DE / current-to-best / 1” mutation strategy in the differential evolution algorithm is used to i Perform mutation operation to generate offspring individuals U i , where X i is the i-th individual in the population that performs the guided mutation operation, and the offspring individual U i The chromosome has n dimensions, and the value u on each dimension i,j Calculated by the following formula:
[0024]
[0025] Where j represents U i The jth dimension in the chromosome, t i,j is a randomly generated value, CR is a user-defined mutation factor between 0 and 1, k is a random integer between 1 and n, and u i,j and x i,j U i and X i The j-th dimension variable in, rand(0,1) returns a random number between 0 and 1, the mutation probability parameter The value of is calculated by the following formula:
[0026]
[0027] where F is a user-defined scaling factor, x best,j is the X in the current population i The best individual for the task represented by the skill factor X best The value of the j-th dimension in the chromosome, x r2,j and x r3,j are two different individuals X randomly selected from the parent population r2 and X r3The value of the j-th dimension in the chromosome of [a]′ represents the Iverson bracket. If the condition a in the bracket is true, it returns 1, otherwise it returns 0. By controlling the mutation parameter u i,j is assigned a new value instead of x i,j In addition, u i,j It is allocated according to the frequency of occurrence of each function and terminal in the current population, that is, elements that appear more frequently in individuals in the population are more likely to be selected and assigned to u i,j .
[0028] Furthermore, the adaptive control strategy adopted for the cross parameter IR in step S4 is as follows:
[0029] When generating a new generation of population in each generation, the adaptive update operation of the parameter IR is determined by statistically analyzing the source of the best individuals of all current parallel tasks. The update formula is as follows:
[0030]
[0031] Where ρ is the decay factor. In the adaptive control strategy, the corresponding update rule is determined based on the generation path of the best individuals for each task. If the majority of the best individuals in the new generation come from the selection crossover operation, the value of the parameter IR tends to 1, increasing the probability of crossover between individuals and improving the frequency of useful information transfer between different tasks. Similarly, if the majority of the best individuals come from random mutations in the guided mutation operation, the parameter IR tends to 0, reducing the crossover frequency while increasing the frequency of mutation search within each individual's local range. If the majority of the best individuals are not updated, the parameter IR cautiously tends to 0.3, maintaining a relatively balanced state between the selection crossover and guided mutation operations.
[0032] Furthermore, the connection operation and selection operation process in step S3 are as follows:
[0033] The newly generated offspring population newpop is mixed and connected with the parent population pop to form a temporary population temppop. The individuals in temppop are sorted on each task according to the corresponding fitness evaluation value of each individual, and the skill factor and scalar fitness of each individual are updated. A one-to-one selection strategy is used in temppop for selection. Each offspring individual U i The corresponding parent individual X i For comparison,
[0034]
[0035] Where φ(X) returns the scalar fitness value of individual X.i Identical individual Xs k and when k < i, X i is considered redundant.
[0036] Furthermore, the adaptive computing resource reallocation mechanism in step S3 is as follows:
[0037] S31. Determine whether the condition for resource reallocation is met. If K - 1 out of all K concurrently executed tasks in the evolutionary process have reached the preset task accuracy threshold, then continue to execute the following steps S32 to S34; otherwise, perform the normal selection operation process;
[0038] S32. Find the worst-performing task by comparing the fitness values of the best-performing individuals of each task;
[0039] S33. Select the top M individuals in terms of factor ranking from the individuals to which the worst-performing task belongs and retain them directly in the next-generation population. If the number m of individuals to which this task belongs is less than M, then retain all of them, and for the rest, circularly copy the top M - m individuals;
[0040] S34. The remaining individuals in the temporary population temppop continue to perform the selection operation.
[0041] Through the above steps, the method proposed by the present invention will neither delay the search for other tasks nor, to a certain extent, tilt the computing resources towards tasks with poor performance, enabling all tasks to obtain better results as much as possible. In addition, after the search accuracy of tasks with better performance reaches the preset standard, the redundant computing resources are allocated to other tasks, which also improves the utilization rate of computing resources and the overall search efficiency of tasks.
[0042] The present invention has the following advantages and effects compared with the prior art:
[0043] 1. By adopting an adaptive control strategy for the crossover parameter IR in the crossover recombination operation, the present invention takes into account the knowledge transfer between tasks and improves the search quality of solutions in the local area during the evolutionary process, reducing the probability of ineffective crossover recombination occurring between uncorrelated tasks. By this operation, the evolutionary behavior is adjusted to improve the search efficiency.
[0044] 2. The present invention designs an adaptive computing resource reallocation mechanism to balance different tasks, enabling tasks with relatively poor performance to obtain more computing resource allocation, so that all concurrently executed tasks can obtain the optimal solution as much as possible. In addition, after the search accuracy of tasks with better performance reaches the preset standard, the redundant computing resources can be allocated to other tasks, improving the utilization rate of resources and the overall search efficiency of tasks.
[0045] 3. The multi-task symbolic regression method provided by the present invention can not only search multiple tasks simultaneously in one execution, effectively saving computational overhead, but also make full use of the correlation and complementary information between different tasks, promote the search of different tasks, and speed up the problem solving speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0047] Figure 1 This is a diagram illustrating an example of the use of the scalable gene expression encoding method of the present invention;
[0048] Figure 2 This is a flow chart of the adaptive symbolic regression method based on the multi-task genetic programming algorithm in the present invention. DETAILED DESCRIPTION
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0050] Example 1
[0051] Assume that there are K symbolic regression tasks to be solved simultaneously, where the function set and terminal set of the i-th task are F i and T i These terminal elements and the function elements together constitute the element set of the solution. The genetic programming algorithm finds the optimal formula that satisfies the training data and the objective function in the given construction element set.
[0052] like Figure 2 As shown, this embodiment is an adaptive symbolic regression method based on a multi-task genetic programming algorithm, comprising the following steps:
[0053] T1. Construct a scalable gene expression encoding method and initialize the population.
[0054] A unified encoding scheme for multiple symbolic regression problems. In this scalable gene encoding scheme, fixed-length strings are used to encode individual chromosomes, which contain a main function and multiple automatically defined functions (ADFs). Furthermore, since different tasks across domains may have unique functions and terminals, integers are used to represent the symbols for each function and terminal in the scalable gene expression encoding scheme to address this issue.
[0055] To ensure that each chromosome in the population can be correctly converted into a valid expression tree for the corresponding task, the head length h and tail length l are subject to the following restrictions:
[0056]
[0057] Among them, §(f) returns the number of parameters of function f, F k Represents the kth function, max{a} returns the maximum value in the set a.
[0058] After determining the length of the chromosome, we use integers to represent the elements in each chromosome. First, the number of ADFs in each chromosome and the number of input parameters in each ADF are N. a and N g After that, four integer ranges are defined to represent these elements, namely function, terminal, ADF and input parameters, which are [0, A-1], [A, B-1], [B, C-1] and [C, D-1]. The values of A, B, C and D are determined by the following formula:
[0059]
[0060] B=A+N a
[0061]
[0062] D=C+N g
[0063] Among them, |F i |Return to set F i The number of elements in , max{a} returns the maximum value in set a. With this common chromosome representation, the unified common chromosome can be converted into task-specific chromosomes and then decoded and evaluated.
[0064] After determining the scalable gene expression encoding method, the search population is initialized. This step randomly generates the initial population pop, with a population size of N. Each individual is represented by an integer vector using this encoding representation method, that is,
[0065] X=[x1,x2,…,x n ]
[0066] Where n is the length of each individual chromosome vector, x n Is an integer in the feasible value range. During initialization, randomly select a feasible type for the corresponding position among {function, ADF, terminal, input parameter} to initialize the corresponding X dimension type. Next, x n Set to a random integer within the range of possible values for the chosen type.
[0067] The population is initialized according to the encoding method, and each individual is evaluated on all problems using multiple factors. In this embodiment, the decoding process of converting the unified chromosome encoding representation into the chromosome of a specific task is as follows: (1) by checking the range of the integer x on each dimension of the chromosome of the unified encoding method, the corresponding symbol type represented by each dimension is found; (2) the integer x is scaled according to the integer range of the type defined by the specific task, and the integer x' is obtained in the range defined by the specific task, that is:
[0068]
[0069] Among them, L x and U x are the minimum and maximum values of the corresponding type of the element found in step (1), N x is the maximum integer used in the corresponding type element, It is a floor operator that returns the largest integer not greater than a; the above steps map the integer x to 0 to N x An integer x' between -1 and x' is used as an index to map x to a meaningful symbol for the specific task.
[0070] Figure 1 is a demonstration of decoding using a scalable gene expression encoding method. Figure 1In this example, two symbolic regression tasks are solved simultaneously, with their function sets, ADFs, terminal sets, and input parameters defined as {+, -, *, / }, {G1}, {x1, x2, x3}, and {T1, T2}, and {+, -, *, / , sin, cos}, {G1}, {x1, x2}, and {T1, T2}, respectively. Therefore, the maximum number of functions, terminals, ADFs, and input parameters for the two problems is 6, 1, 3, and 2, respectively. Based on the proposed S-ADF representation, we have A = 6, B = 7, C = 10, and D = 12. In this example, the chromosome is structured to have a single ADF, with the lengths of its main function and ADF head and tail set to (4, 5) and (2, 3), respectively. Let [1, 5, 6, 2, 7, 7, 9, 8, 9, 0, 2, 10, 11, 10] be an example of an S-ADF-encoded chromosome. In this example, the expression of the master function can be obtained based on the first nine dimensions (i.e., [1, 5, 6, 2, 7, 7, 9, 8, 9]), while the ADF is obtained based on the remaining dimensions (i.e., [0, 2, 10, 11, 10]). Then, in order to transform the chromosome into a solution to the corresponding symbolic regression problem, each dimension of the chromosome must be mapped to a symbol that makes sense for the problem based on the number of defined functions, terminals, and ADFs. Next, the first step is to identify x i The type of , and then assign the correct symbol to x according to the corresponding type obtained in the first stage i For example, for the first symbolic regression problem, the value of the first dimension is 1, which belongs to [0, A-1], so the element type of the first dimension is a function. Then, we assign the correct function symbol to x by calculating the scaled value of the element. i Since (1-0) / (6-0)*3=0, we set xi to the first function (+). In this way, we can translate the chromosome into a solution corresponding to the first symbolic regression problem [+,*,G1,-,x1,x1,x2,x1,x2,+,*,T1,T2,T1], which can be further decoded into the final expression.
[0071] T2. Choose to perform a selection crossover operation or a guided mutation operation on the population based on the crossover parameter IR.
[0072] By counting the operation sources of the best individual in each task, an adaptive control strategy is adopted for the control parameter IR, thereby adjusting the evolutionary behavior. In this process, the operation is performed on each individual one by one to produce a new generation of offspring population.
[0073] The selection crossover operation uses a single-point crossover by randomly selecting the parent individual X r1 and X i Crossover operation to generate offspring individuals U i, where X r1 represents a randomly selected individual, X i is the i-th individual in the population that performs single-point crossover; for the offspring individual U i Perform a unified mutation operation, and the generated offspring individuals U i The skill factor is transmitted vertically through culture and is inherited randomly from one of the parents with the same probability.
[0074] On the other hand, the guided mutation operation adopts the “DE / current-to-best / 1” mutation strategy in the differential evolution algorithm to i Perform mutation operation to generate offspring individuals U i , where X i is the i-th individual in the population that performs the guided mutation operation, and the offspring individual U i The chromosome has n dimensions, and the value u on each dimension i,j Calculated by the following formula:
[0075]
[0076] Where j represents U i The jth dimension in the chromosome, t i,j is a randomly generated value, CR is a user-defined mutation factor between 0 and 1, k is a random integer between 1 and n, and u i,j and x i,j U i and X i The j-th dimension variable in, rand(0,1) returns a random number between 0 and 1, the mutation probability parameter The value of is calculated by the following formula:
[0077]
[0078] where F is a user-defined scaling factor, x best,j is the X in the current population i The best individual for the task represented by the skill factor X bset The value of the j-th dimension in the chromosome, x r2,j and x r3,j are two different individuals X randomly selected from the parent population r2 and X r3 The value of the jth dimension in the chromosome of [a]′ represents the Iverson bracket. If the condition a in the bracket is true, it returns 1, otherwise it returns 0. By controlling the mutation condition, u i,j is assigned a new value instead of x i,j probability.
[0079] Finally, the role of the adaptive control strategy for the cross parameter IR is to control whether the population selects the assortative mating recombination method or the mutation method to generate offspring individuals.
[0080] T3, connection, and selection operations are used to select the individuals of the next generation population. <00,00275>The newly generated offspring population is mixed and connected with the parent population to form a temporary population, and the factor rankings and scalar fitness values corresponding to each task are recalculated in the temporary population; then, based on the scalar fitness, one-to-one selection is performed between the parent individuals and the offspring individuals to select the individuals entering the next generation population; in this process, by constructing an adaptive computing resource reallocation mechanism, the population individuals belonging to the tasks with relatively poor performance can enter the next generation population more, and obtain more computing resource allocations. <000027,7>In the connection operation, the newly generated offspring population newpop and the parent population pop are mixed and connected together to form a temporary population temppop. Subsequently, the individuals in temppop are sorted independently on each task according to their corresponding factor costs. Then, the skill factors and scalar fitness values of each individual are updated.
[0083] After that, a one-to-one selection strategy is adopted in temppop for selection, and each offspring individual U i is compared with its corresponding parent individual X i That is
[0084]
[0085] where φ(X) returns the scalar fitness value of individual X. When there is another separate X i identical to X j completely, and j < i, X i is considered redundant.
[0086] Finally, the adaptive computing resource reallocation mechanism balances the computing resources allocated to each task, ensuring that all tasks achieve the best possible search results, thereby improving the algorithm's overall resource utilization. The first step is to determine whether the conditions for resource reallocation have been met. During the evolutionary process, if K-1 of the K concurrently executing tasks have reached the preset task accuracy threshold, the next step proceeds; otherwise, the search proceeds normally. The subsequent operation consists of three steps. The first step is to identify the worst-performing task by comparing the fitness values of the best-performing individuals for each task. The second step is to select the top M individuals in the factor ranking from the individuals belonging to the worst-performing task and retain them directly to the next generation of the population. If the number m of individuals belonging to the task is less than M, all are retained, and the remaining individuals are cyclically replicated to the top Mm individuals. The third step is to continue the selection process among the remaining individuals.
[0087] T4. Adaptively update the cross parameter IR.
[0088] When generating a new generation of population in each generation, the adaptive update operation of the parameter IR is determined by statistically analyzing the source of the best individuals of all current parallel tasks. The update formula is as follows:
[0089]
[0090] Where ρ is the decay factor. In the adaptive control strategy, the corresponding update rule is determined based on the generation path of the best individuals for each task. If the majority of the best individuals in the current generation come from assortative mating and recombination via the selective crossover operation, the crossover parameter IR approaches 1, increasing the probability of crossover between individuals and improving the frequency of useful information transfer between different tasks. Similarly, if the majority of the best individuals come from random mutations via the guided mutation operation, the crossover parameter IR approaches 0, reducing the crossover frequency while increasing each individual's mutation search within its local range. If the majority of the best individuals are not updated, the crossover parameter IR cautiously approaches 0.3 to maintain a relatively balanced state between the selective crossover and mutation operations.
[0091] T5. Determine the termination condition.
[0092] If the accuracy conditions specified for all tasks are met or the number of fitness evaluations reaches the specified maximum number, the evolution is terminated and the solution of the best individual corresponding to each task in the final population is output; otherwise, the evolution process continues to return to steps T2-T6 for a cyclic search of the population.
[0093] In order to verify and evaluate the performance of the algorithm framework of this embodiment, this embodiment is verified by simulation experiments using 5 symbolic regression benchmark test problem data sets. This embodiment will perform multi-task concurrent execution of these 5 symbolic regression tasks at the same time, as shown in steps T1--T5. The specific parameters of this embodiment are set as follows: the population size is N=100, the head length of the main function in the chromosome structure is h=10; the number of ADFs is set to 1, and the head length in the ADF is h′=3, the crossover control parameter IR=0.3, the attenuation factor ρ=0.8, and the number of retained individuals M=45. Table 1 lists the expressions of these 5 symbolic regression benchmark test problems, and each problem sniper contains 500 data.
[0094] In this example, the root mean square error (RMSE) is used as the fitness evaluation method for each individual on all tasks. The search accuracy of the multi-task method proposed in this invention and the existing single-task search SLGEP algorithm are evaluated. The experimental results are shown in Table 2.
[0095] From the results in Table 2, it can be concluded that the search accuracy of the method proposed in the present invention is better than that of the general single-task search algorithm in the simulation test, and the result accuracy is significantly higher. This shows that the method proposed in the present invention is effective in improving the search ability of the genetic programming algorithm for multiple symbolic regression problems.
[0096] Table 1. Detailed list of expressions for the five symbolic regression benchmark tests
[0097]
[0098] Table 2. RMSE comparison of experimental results in Example 1
[0099] <![CDATA[F1]]> <![CDATA[F2]]> <![CDATA[F3]]> <![CDATA[F4]]> <![CDATA[F5]]> Example 1 0.0427 0.0101 0.312 0.00780 0.01 SLGEP 0.0633 0.0528 0.356 0.0142 0.0107
[0100] Example 2
[0101] To further verify the effectiveness of the method proposed in the present invention, the adaptive control strategy for the cross parameter IR is not adopted in Example 2, and the remaining steps are consistent with those in Example 1. This embodiment mainly includes the following steps:
[0102] P1-P4: Refer to steps T1-T4 in Example 1.
[0103] P5. Determine the termination condition.
[0104] If the accuracy conditions specified for all tasks are met or the number of fitness evaluations reaches the specified maximum number, the evolution is terminated and the solution of the best individual corresponding to each task in the final population is output; otherwise, the evolution process continues to return to steps P2-P5 for a cyclic search of the population.
[0105] This example also uses the five symbolic regression benchmark datasets in Table 1 for simulation verification. The specific experimental parameter settings remain consistent with those in Example 1. The experimental results of this example are compared with those in Example 1 and the experimental results of the single-task search SLGEP algorithm. The experimental results are shown in Table 3.
[0106] Table 3. RMSE comparison of experimental results in Example 2
[0107] <![CDATA[F1]]> <![CDATA[F2]]> <![CDATA[F3]]> <![CDATA[F4]]> <![CDATA[F5]]> Example 2 0.0478 0.0277 0.326 0.0105 0.0112 Example 1 0.0427 0.0101 0.312 0.00780 0.01 SLGEP 0.0633 0.0528 0.356 0.0142 0.0107
[0108] As can be seen from Table 3, after not adopting the adaptive control strategy for the cross parameter IR, the experimental accuracy of Example 2 decreases, and the performance is not as good as that of Example 1. This demonstrates the effectiveness of the cross parameter adaptive control strategy proposed in the present invention.
[0109] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. An adaptive symbolic regression method based on a multi-task genetic programming algorithm, characterized in that: The adaptive symbolic regression method comprises the following steps: S1. Construct a scalable gene expression encoding method to form a unified search space for uniformly encoding multiple different symbolic regression tasks, where the symbolic regression task includes a function set and a terminal set corresponding to the task; initialize all individuals in the population according to the encoding method; perform multi-factor evaluation on each individual on all tasks by converting the unified scalable gene expression encoding representation into a decoding operation of the chromosome encoding method of a specific task; S2, select the population to perform a selection crossover operation or a guided mutation operation according to the crossover parameter IR, and perform the above operations on each individual in the population one by one to produce a new generation of offspring population; S3, perform a mixed connection operation on the newly generated offspring population and the parent population to form a temporary population, and re-perform multi-factor evaluation and scalar fitness calculation in the temporary population, wherein the mixed connection operation includes a connection operation and a selection operation; perform a one-to-one selection between the parent individual and the offspring individual according to the scalar fitness to select the individuals that enter the next generation population; construct an adaptive computing resource reallocation mechanism so that more individuals of the population to which the tasks with poor performance belong enter the next generation population; S4, count the operation sources of the best individuals for each task in the next generation population, and adopt an adaptive control strategy for the crossover parameter IR to adjust the evolutionary behavior; S5. Determine whether the precision conditions or the number of fitness evaluations specified for all tasks have reached the specified maximum number. If so, terminate the evolution and output the best individual solution corresponding to each task in the final population; otherwise, continue to return to the evolutionary process from step S2 to step S5 to perform a population cyclic evolutionary search.
2. The adaptive symbolic regression method based on a multi-task genetic programming algorithm according to claim 1, characterized in that: The scalable gene expression encoding method in step S1 is as follows: (1) a fixed-length string is used as an individual chromosome; (2) a main function and multiple automatically defined functions (ADFs) are included; and (3) integers are uniformly used to represent functions and terminals, and a single integer is used to represent various symbols for different tasks.
3. The adaptive symbolic regression method based on a multi-task genetic programming algorithm according to claim 2, characterized in that: The specific decoding process of converting the unified chromosome encoding representation into the chromosome of a specific task in step S1 is as follows: S11, by checking the range to which the integer x on each dimension of the chromosome of the unified encoding method belongs, find out the corresponding symbol type represented by each dimension; S12, scaling the integer x according to the integer range of the type defined by the specific task, to obtain an integer x' in the range defined by the specific task, that is: Among them, L x and U x are the minimum and maximum values of the corresponding type of the element found in step S11, respectively, N x is the maximum integer used in the elements of the corresponding type, It is the floor operator for a, returning the largest integer not greater than a; the above steps map the integer x to 0 to N x -1, and use the scaled value x' as an index to map x to a meaningful symbol for the specific task.
4. The adaptive symbolic regression method based on a multi-task genetic programming algorithm according to claim 1, characterized in that: The selection crossover operation in step S2 is specifically as follows: Using single-point crossover, randomly select the parent individual X r1 and X i Crossover operation to generate offspring individuals U i , where X r1 represents a randomly selected individual, X i is the i-th individual in the population that performs single-point crossover; for the offspring individual U i Perform a unified mutation operation, and generate the offspring individual U i The skill factor is transmitted vertically through culture and is randomly inherited from one of the parents with the same probability.
5. The adaptive symbolic regression method based on a multi-task genetic programming algorithm according to claim 4, characterized in that: The guided mutation operation in step S2 is specifically as follows: The "DE / current-to-best / 1" mutation strategy in the differential evolution algorithm is used to i Perform mutation operation to generate offspring individuals U i , where X i is the i-th individual in the population that performs the guided mutation operation, and the offspring individual U i The chromosome has n dimensions, and the value u in each dimension i,j Calculated by the following formula: Where j represents U i The jth dimension in the chromosome, t i,j is a randomly generated value, CR is a user-defined mutation factor between 0 and 1, k is a random integer between 1 and n, and u i,j and x i,j U i and X i The j-th dimension variable in, rand(0,1) returns a random number between 0 and 1, and the mutation probability parameter The value of is calculated by the following formula: where F is a user-defined scaling factor, x best,j is the current population X i The best individual X for the task represented by the skill factor best The value of the jth dimension in the chromosome, x r2,j and x r3,j are two different individuals X randomly selected from the parent population. r2 and X r3 The value of the j-th dimension in the chromosome of , [a]′ represents an Iverson bracket. If the condition a in the bracket is true, it returns 1, otherwise it returns 0.
6. The adaptive symbolic regression method based on a multi-task genetic programming algorithm according to claim 1, characterized in that: The adaptive control strategy adopted for the crossover parameter IR in step S4 is specifically as follows: When a new generation of population is generated in each generation, the adaptive update operation of the parameter IR is determined by statistically analyzing the source of the best individuals of all current parallel tasks. The update formula is as follows: Where ρ is the decay factor. In the adaptive control strategy, the corresponding update rule is determined according to the generation path of the best individual for each task.
7. The adaptive symbolic regression method based on a multi-task genetic programming algorithm according to claim 1, characterized in that: The connection operation and selection operation process in step S3 are as follows: The newly generated offspring population newpop is mixed and connected with the parent population pop to form a temporary population temppop. The individuals in temppop are sorted on each task according to the corresponding fitness evaluation value of each individual, and the skill factor and scalar fitness of each individual are updated; a one-to-one selection strategy is adopted in temppop for selection, and each offspring individual U i The corresponding parent individual X i For comparison, where φ(X) returns the scalar fitness value of individual X, and when there exists another separate X i that is exactly the same as X k and k < i, then X i is considered redundant.
8. The adaptive symbolic regression method based on a multi-task genetic programming algorithm according to claim 1, characterized in that: The adaptive computing resource reallocation mechanism in step S3 is as follows: S31, judging whether the conditions for resource reallocation are met, if K-1 tasks among all K tasks executed concurrently in the evolution process have reached the preset task accuracy threshold, then continue to execute the next steps S32 to S34, otherwise, the selection operation process is carried out normally; S32, by comparing the fitness values of the tasks of the best performing individuals belonging to each task, thereby finding the worst performing task; S33, among the individuals belonging to the worst performing task, select the top M individuals in the factor ranking and retain them, and directly enter the next generation population. If the number m of individuals belonging to the task is less than M, all of them are retained, and the rest are cyclically replicated to the top Mm individuals; S34. The remaining individuals in the temporary population temppop continue to perform the selection operation.
Citation Information
Patent Citations
Large-scale symbol regression method and system based on adaptive parallel genetic algorithm
CN110135584A
Modeling of systems using canonical form functions and symbolic regression
US20070208548A1