Adaptive knowledge migration method for multi-task Bayesian optimization

By introducing an explicit knowledge transfer mechanism and a kernel-based autoencoding mechanism, the problem of handling nonlinear relationships between tasks in multi-task Bayesian optimization is solved, which improves the efficiency of knowledge sharing and optimization accuracy, reduces the negative impact of interfering tasks, and realizes selective collaboration between tasks.

CN120996138AActive Publication Date: 2025-11-21SOUTH CHINA UNIV OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511092384.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-21
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing multi-task Bayesian optimization methods rely on implicit knowledge transfer mechanisms when dealing with multiple related optimization tasks, which leads to the algorithm performance being affected by interfering tasks and fails to effectively achieve joint optimization between tasks.

Method used

An explicit knowledge transfer mechanism is introduced, which captures the nonlinear relationships between tasks through a kernel-based automatic encoding mechanism. Using a multivariate Gaussian distribution model and a pheromone concentration heuristic rule, auxiliary tasks are dynamically selected and the optimal solution is transferred. A mapping matrix between tasks is established to improve the efficiency of knowledge sharing.

Benefits of technology

It significantly improves the efficiency of knowledge sharing in multi-task optimization, reduces the negative impact of interfering tasks on optimization performance, and achieves selective collaboration between tasks and higher optimization accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996138A_ABST
    Figure CN120996138A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task Bayesian optimization self-adaptive knowledge migration method, which comprises the following steps of: capturing a non-linear relationship among data sets by utilizing a kernel automatic coding mechanism so as to measure the similarity among tasks; the priority of auxiliary task selection is dynamically adjusted through a heuristic rule, selective cooperation between tasks is ensured, and interference is reduced. Bayesian optimization modeling and black box function optimization are used to reduce the number of times of task real evaluation. According to the method, an explicit knowledge migration mechanism is introduced, the nonlinear mapping matrix is used for converting the optimal solution of the auxiliary task, efficient knowledge migration and collaborative optimization among the multiple related tasks are achieved, and the optimization efficiency and the resource utilization rate are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-task Bayesian optimization, and mainly to an adaptive knowledge transfer method for multi-task Bayesian optimization. Background Technology

[0002] Bayesian optimization, as an efficient global optimization method, is widely used in solving complex and expensive black-box function optimization problems, and mainly consists of two core components: a surrogate model and a sampling function. The surrogate model is constructed to approximate the objective function; this model predicts the function values ​​and their uncertainties in the unsampled region using a probability distribution. Common surrogate models include Gaussian processes, neural networks, and random forests. The sampling function helps select new data points; the selected data points are added to the training dataset after being evaluated by the objective function, and the surrogate model is updated. Due to its high sampling efficiency, Bayesian optimization can find the global optimum with fewer evaluations of the objective function.

[0003] While traditional Bayesian optimization methods have achieved success, real-world scenarios often involve multiple related optimization tasks that need to be solved simultaneously. To address this, multi-task Bayesian optimization, as an extension method, has received increasing attention in recent years. Its core idea is to leverage the correlation between multiple tasks, improving optimization efficiency through information sharing. This is particularly effective when the tasks are highly similar, significantly reducing the number of evaluations of the objective function.

[0004] Existing multi-task Bayesian optimization methods mainly focus on improving surrogate model construction and optimizing sampling strategies. Regarding surrogate models, the most commonly used multi-task Gaussian process model can simultaneously model the objective functions of multiple tasks (Swersky K, Snoek J, Adams R P. Multi-task bayesian optimization[J]. Advances in neural information processing systems, 2013, 26.), and captures the correlation between tasks by sharing the covariance matrix, thereby achieving collaborative optimization between tasks. Other studies have obtained the collected signals by training two Gaussian process models (Hakhamaneshi K, Abbeel P, Stojanovic V, et al. Jumbo: Scalable multi-task bayesian optimization using offline data[J]. arXiv preprint arXiv:2106.00942, 2021.), namely a cold Gaussian process running directly in the input domain and a warm Gaussian process running in the feature space of a deep neural network pre-trained using offline data. Regarding sampling strategies, cost-sensitive multi-task sampling functions based on entropy search and multi-task GP-UCB sampling functions are all improvements on single-task sampling functions. Some studies also suggest that sharing evaluation for highly similar tasks and parallel evaluation for different tasks can maximize resource efficiency and accelerate optimization. However, most of these multi-task Bayesian optimization methods rely on implicit knowledge transfer mechanisms, making them susceptible to poor algorithm performance due to interfering tasks. Furthermore, these methods primarily utilize existing knowledge to optimize the objective function, rather than achieving joint optimization of multiple related tasks. To address these challenges, this invention proposes an adaptive knowledge transfer method for multi-task Bayesian optimization. This method introduces an explicit knowledge transfer mechanism, models the nonlinear mapping between tasks, and directly transfers the optimal solution from related tasks, thereby improving optimization efficiency. Summary of the Invention

[0005] The technical problem solved by this invention is to provide an adaptive knowledge transfer method with multi-task Bayesian optimization.

[0006] The present invention is achieved by at least one of the following technical solutions.

[0007] An adaptive knowledge transfer method using multi-task Bayesian optimization includes the following steps:

[0008] S1. Initialize the pheromone concentration between tasks, randomly generate n initial solutions for each task, and establish a separate Gaussian process model.

[0009] S2. Begin evolution with a period of G. In the g-th generation, perform Bayesian optimization for each task and update the optimal solution.

[0010] S3. If the current algebra g is divisible by the set threshold, update the mapping matrix between any two tasks, calculate the similarity between each pair of tasks, select a suitable auxiliary task for each main task and migrate its optimal solution after mapping, and update the pheromone concentration.

[0011] S4. If the termination condition is met, terminate and return the current optimal solution for each task; otherwise, proceed to step S2.

[0012] Further, in step S1, the pheromone concentration is initialized as follows:

[0013] τ(T m ,T a )=(Phe max +Phe min ) / twenty one);

[0014] Among them Phe max and Phe min τ(T) represents the upper and lower limits of pheromone concentration, respectively. m ,T a ) for auxiliary task T a Main Task Support T m The concentration of pheromones.

[0015] Furthermore, in step S1, for the i-th task T i The expression for establishing the Gaussian process model is:

[0016] f(x)~GP(m(x),k(x,x′))(2); where m(·) is the prior mean function, k(·) is the kernel function, and let the i-th task T be denoted as f(x)~GP(m(x),k(x,x′))(2); i Given a finite dataset (X, y), then x, x′∈X, and f(x) represents the i-th task T. i The function expression, GP(·) represents Gaussian process modeling, x represents the input variable, X represents the set of sampled input variables, and T i Let represent the i-th task in multi-task optimization, y represent the set of observations, and the mean and variance of the candidate point x′ are calculated using the following expressions:

[0017] μ(x′)=k(x′;X)K(X;X) -1 y (3);

[0018] σ 2 (x′)=k(x′;x′)-k(x′;X)K(X;X)-1 k(x′;X) T (4); where μ(x′) represents the mean of candidate point x′, σ 2 (x′) represents the variance of candidate point x′, k(x′;X) calculates the correlation between x′ and each x in X using a kernel function, and K(X;X) is the covariance matrix containing information about the internal correlations of X. The kernel function chosen for the Gaussian process is the Matérn kernel.

[0019]

[0020] Where v is the smoothing coefficient, l is the length scale parameter, and K... v It is the modified Bessel function, and k(x,x′) is the kernel function.

[0021] Furthermore, in step S2, Bayesian optimization uses LogEI as the acquisition function, with the specific expression as follows:

[0022]

[0023] Where x represents the input variable, The function representing the logarithm of the desired improvement at candidate point x is log_h(.), which is a function that performs a logarithmic transformation on a portion of the calculations in the desired improvement. μ(x) represents the mean of candidate point x, σ(x) represents the standard deviation of candidate point x, and y * The current optimal value of the function is log_h, which can be calculated using the following expression:

[0024]

[0025] Where c1 = log(2π) / 2, log1me(z) = log(1-exp(z)), and ∈ represents the numerical precision. c2=log(π / 2) / 2, erfcx(z)=exp(z 2 )erfc(z), φ(z) is the probability density function of the standard normal distribution, z is the standardized variable in the standard normal distribution, Φ(z) is the cumulative distribution function of the standard normal distribution, exp(·) is the exponential function, erfcx(·) represents the scaled form of the complementary error function, and erfc(·) is the complementary error function.

[0026] Furthermore, in step S3, when g is divisible by interval, the specific process of adaptive knowledge transfer is as follows:

[0027] S31. Calculate the mapping matrix between two tasks through a kernel-based self-encoding mechanism. There are two types of mapping matrices: one records the mapping between decision spaces and the other records the mapping between joint spaces.

[0028] S32. For each task, randomly select kN test points on the domain of the function and retain N sampling points that are close to the known points and the current best point. Use Gaussian process regression to evaluate these sampling points to obtain the similarity measurement dataset of the task.

[0029] S33. For each pair of auxiliary tasks and main tasks, the similarity measurement dataset of the auxiliary tasks is mapped to the problem domain of the main task. The distribution of the similarity measurement dataset of the auxiliary tasks and the main tasks is approximated by a multivariate Gaussian distribution model, and the similarity between the two is calculated by Kullback-Leibler divergence.

[0030] S34. Combining kernel-based task similarity measurement and pheromone concentration-based heuristic rules, select auxiliary tasks for each task and transfer the current optimal solution of the auxiliary tasks to the problem domain of the main task through mapping.

[0031] S35, Update pheromone concentration.

[0032] Furthermore, in step S31, the mapping M obtained by the kernel-based automatic encoding mechanism k The closed-form solution is:

[0033] M k =S2K(S1,S1) T (K(S1,S1)K(S1,S1) T ) -1 (8);

[0034] Where S1 and S2 are the corresponding initial solution sets of the two optimization tasks T1 and T2, and the (m,n)th element of the kernel matrix K(S1,S1) is κ(s 1m ,s 1n ), κ(·,·) are kernel functions; s 1m and s 2m Let S1 and S2 represent the m-th solutions, respectively. If we want to map the optimal solution set RS of optimization task T1 to the problem domain of optimization task T2, denoted as MPS after mapping, the calculation method is as follows:

[0035] MPS = M k K(S1,PS) (9).

[0036] Further, in step S32, obtaining the similarity measurement dataset for the task includes the following steps:

[0037] S321, Note Dcandidates ={x s |x s ~f i (x),s=1,2,…,kN},D candidates Let x be the set of candidate points obtained by randomly sampling on the domain of the function for task i, where x s Let f represent the s-th candidate point. i (x) represents the objective function of task i, and x represents the input variable of task i;

[0038] For each candidate point x∈D candidates The posterior standard deviation was calculated using a Gaussian process model. Let Var[·] be the function value predicted by the Gaussian process model at candidate point x for task i, and Var[·] represent the variance estimate of the Gaussian process model at candidate point x. Mapping the posterior standard deviation to the deterministic score, we get:

[0039]

[0040] Where ∈ is a positive number, used to avoid the denominator being zero; C certainty (x) is the deterministic score of candidate point x, C certainty(x′) P is the deterministic score of candidate point x′. certainty (x) represents the importance of the prediction confidence of the Gaussian process model at the candidate point x.

[0041] S322, For each candidate point x∈D candidates Calculate the current optimal point x from x to task i. cb The Euclidean distance d(x,x) cb )=||xx cb || 2 Similarly, the distance is mapped to a distance weight W. distance (x), and calculate the normalized probability:

[0042]

[0043] Where P distance (x) reflects the importance of x among all candidate points based on distance, W distance (x′) represents the distance weight of candidate point x′;

[0044] S323. Introduce a weighting coefficient ω>0 to calculate the final joint probability of candidate point x:

[0045]

[0046] P combined(x) represents the final joint probability of candidate point x; P tmp (x) and P tmp (x′) represent the provisional joint probability at candidate points x and x′, combining prediction confidence and distance-based importance, respectively;

[0047] S324, According to the joint probability P combined (x), from the candidate point set D candidates We will perform sampling without replacement to select N valid sampling points for task i to be used for similarity measurement:

[0048] D selected ={x s |x s ~P combined (x),i=1,2,…,N} (16);

[0049] D selected Let x represent the final set of selected sampling points. s This represents the s-th selected sampling point.

[0050] Furthermore, the mapping in step S33 is a mapping between joint spaces, which are formed by concatenating the decision space and the target space. That is, for two datasets D... A =(X A ,y A ) and D B =(X B ,y B ), X A and X B Let y be the set of input variables. A and y B The mapping is established for the set of target values ​​corresponding to input variables by [X] A →X B Upgraded to [X] A ,y A →X B ,y B Furthermore, the formula for calculating the Kullback–Leibler divergence is as follows:

[0051]

[0052] Where Σ represents the covariance matrix, μ P′ μ Q Both represent the average vector, tr(·) and det(·) represent the rank and determinant of the matrix, respectively. P′ and Q represent two probability distributions, d represents the dimension of the data, and KLD(P′||Q) measures the difference between the two probability distributions.

[0053] The main task T is obtained by taking the reciprocal and normalizing.m And auxiliary tasks T a The similarity ξ(T) between them m ,T a ):

[0054]

[0055] Where T represents the main task T m Auxiliary task set. KLD(T) a ||T m ) represents the main task T m And auxiliary tasks T a Kullback–Leibler divergence between them, KLD(t||T) m ) represents the main task T m Kullback–Leibler divergence between the auxiliary task t and the auxiliary task t.

[0056] Furthermore, the method for selecting the auxiliary task in step S34 is as follows:

[0057]

[0058] Where τ(T) m ,T a ) for auxiliary task T a Main Task Support T m pheromone concentration, P kt (T m ,T a ) is T a Support T m The probability of knowledge transfer, ξ(T) m ,T a ) as the main task T m And auxiliary tasks T a The similarity between them; T represents the main task T m Auxiliary task set; pre-set probability P select Used in tk P kt A secondary task is randomly selected from the main task; rand(0,1) returns a random floating-point number between 0 and 1. Thus, for the main task T... m The auxiliary task selection for knowledge transfer is represented as:

[0059]

[0060] Where KT(T) m ) represents the main task T m In the selection of auxiliary tasks during knowledge transfer, P select This represents the probability of selecting one task from the current tk tasks as an auxiliary task.

[0061] Furthermore, the method for updating the pheromone concentration in step S35 is as follows:

[0062]

[0063] Where τ(T) m ,T a ) for auxiliary task T a Main Task Support T m pheromone concentration, α d and α r The reward and decay rates of pheromones.

[0064] A computer device according to the present invention includes: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements the aforementioned multi-task Bayesian optimization adaptive knowledge transfer method.

[0065] The present invention has the following advantages and effects compared with the prior art:

[0066] 1) The adaptive knowledge transfer method for multi-task Bayesian optimization disclosed in this invention effectively captures the nonlinear relationships between tasks through a kernel-based autoencoding mechanism, which can handle more complex inter-task relationships and significantly improves the knowledge sharing efficiency in multi-task optimization.

[0067] 2) The adaptive knowledge transfer method for multi-task Bayesian optimization disclosed in this invention can dynamically adjust the selection priority of auxiliary tasks, effectively reduce the negative impact of interfering tasks on optimization performance, and ensure selective cooperation between tasks. Attached Figure Description

[0068] Figure 1 This is a flowchart of the adaptive knowledge transfer method based on multi-task Bayesian optimization disclosed in this invention.

[0069] Figure 2 This is a flowchart of the task similarity measurement based on the kernel method disclosed in this invention.

[0070] Figure 3 This is a flowchart of the heuristic rules based on pheromone concentration disclosed in this invention. Detailed Implementation

[0071] The present invention will be further described below with reference to the accompanying drawings and embodiments. This invention is described using the task of neural network hyperparameter tuning as an example. This task is widely present in the training process of deep learning models, such as the hyperparameter configuration problems involving learning rate, batch size, regularization coefficient, and network structure often involved in image classification tasks. Each optimization task corresponds to the hyperparameter tuning process of a specific network architecture on different datasets, and its optimization objective is to minimize the validation set error, thereby obtaining the optimal model performance. There is structural or data similarity between different tasks, which is suitable for knowledge transfer to improve tuning efficiency. In this embodiment, to verify the effectiveness of the proposed method, experiments were conducted using a standard multi-task optimization test function set (including functions such as Rastrigin, Sphere, and Rosenbrock) to simulate multiple related hyperparameter optimization tasks, so as to systematically evaluate the proposed algorithm's ability to transfer learning and its optimization efficiency across different tasks. Experimental results show that, under the condition of a fixed number of actual evaluations of 60 tasks (i.e., controlling the computational resources and time costs in the optimization process), the method proposed in this invention significantly outperforms the standard single-task Bayesian optimization method and the multi-task Bayesian optimization method based on implicit knowledge transfer mechanism in obtaining the optimal solution of the function in multiple optimization tasks. Specifically, under the same evaluation budget, this method can get closer to the global optimal solution of each task, demonstrating higher optimization accuracy and convergence efficiency.

[0072] An adaptive knowledge transfer method using multi-task Bayesian optimization according to an embodiment of the present invention includes the following steps:

[0073] Suppose there exist K interrelated black-box functions, each function f i (x) are all considered as a task to be optimized, where i = 1, 2, ..., K. Specifically, in image classification applications, f i (x) represents the validation set loss function of the model built using the same neural network (e.g., MobileNetV2) on the i-th image classification dataset. Here, the input variable x represents the hyperparameter configuration of the model. By unifying all tasks into a minimization problem, the optimization objective of this invention is to simultaneously find the optimal solution set for all tasks. and satisfy:

[0074]

[0075] Where f i :X i →R d Represents the i-th task T i The objective function of dimension d, X i For T iThe search space. During the optimization process, each task can be used as the primary task, and auxiliary tasks can be selected to migrate the optimal solution.

[0076] As one example, image classification datasets include datasets such as CIFAR-10, SVHN, and Fashion-MNIST.

[0077] To handle potential nonlinear relationships between tasks during explicit knowledge transfer, this invention introduces a kernel-based autoencoder mechanism. Assume two optimization tasks T1 and T2, with initial solution sets S1 and S2, each containing N solutions. From the perspective of a denoising autoencoder, a solution to one problem domain can be viewed as a corrupted version of a solution to another problem domain. Thus, a link between the two problem domains is constructed through the denoising process of the autoencoder. Essentially, this approach builds a cross-problem-domain mapping using a single-layer denoising autoencoder, employing a single-level mapping M:R. d →R d Reconstructing the damaged input to minimize the squared reconstruction loss yields a closed-form solution that can be expressed as a least-squares method. Simultaneously, kernel methods are used to map the data to a high-dimensional feature space to handle the nonlinear relationship between the two datasets. Specifically, assuming the solution set S1 is mapped by the nonlinear mapping function Φ to the regenerable kernel Hilbert space H, the updated reconstruction loss function can be obtained. Based on the mathematical properties of the regenerable kernel Hilbert space, the mapping M can be derived. k Closed-form solution:

[0078] M k =S2K(S1,S1) T (K(S1,S1)K(S1,S1) T ) -1 (2)

[0079] Wherein, the (m,n)th element of the kernel matrix K(S1,S1) is κ(s 1m ,s 1n ). s 1m and s 2m Let S1 and S2 represent the m-th solutions, respectively, and k(·,·) be the kernel function. To map the optimal solution set PS of T1 to the problem domain of T2, denoted as MPS after mapping, the calculation method is as follows:

[0080] MPS = M k K(S1,PS) (3)

[0081] In this invention, both similarity measurement and knowledge transfer utilize nonlinear mapping relationships between tasks. Figure 1As shown in this embodiment, an adaptive knowledge transfer method based on multi-task Bayesian optimization is implemented, comprising the following steps:

[0082] S1. Initialize the pheromone concentration between tasks, randomly generate n initial solutions for each task, and establish a separate Gaussian process model.

[0083] As one embodiment, this embodiment introduces a heuristic rule based on pheromone concentration to dynamically adjust the priority of auxiliary task selection. max and Phe min These represent the upper and lower limits of pheromone concentration, respectively. For each auxiliary task, the initial pheromone value is set to (Phe max +Phe min ) / 2. For each task, generate n initial solutions randomly and build separate Gaussian process models for each.

[0084] If the i-th task T i Given a finite dataset (X,y), the Gaussian process model can be represented using the prior mean function m(·,·) and the kernel function k(·,·):

[0085] f(x)~GP(m(x),k(x,x′))(4)

[0086] Where x, x′∈X, f(x) represents the i-th task T i The function expression is GP(·), which represents Gaussian process modeling, x represents the input variable, and X represents the set of sampled input variables.

[0087] The mean and variance of candidate point x′ can be calculated using the following expression:

[0088] μ(x′)=k(x′;X)K(X;X) -1 y (5)

[0089] σ 2 (x′)=k(x′;x′)-k(x′;X)K(X;X) -1 k(x′;X) T (6) where μ(x′) represents the mean of candidate point x′, σ 2 (x′) represents the variance of candidate point x′, and k(x′;x′) calculates the autocorrelation of x′. k(x′;X) calculates the correlation between x′ and each x in X using a kernel function, and K(X;X) is a covariance matrix containing information about the internal correlations of X. The Matérn kernel is chosen as the kernel function for Gaussian processes:

[0090]

[0091] Where v is the smoothing coefficient, l is the length scale parameter, and K... v It is the modified Bessel function, and k(x,x′) is the kernel function.

[0092] S2. Begin the evolutionary process with a period of G. In the g-th generation, perform Bayesian optimization for each task and update the optimal solution. After completing the initial modeling, enter the optimization process with a period of G, and the current generation is g. Perform the traditional Bayesian optimization process for each task, using LogEI as the data acquisition function. The specific expression is as follows:

[0093]

[0094] Where x represents the input variable, The function represents the logarithmic sampling function that takes the logarithm of the desired improvement at candidate point x. log_h(.) is a function that performs a logarithmic transformation on a portion of the calculations in the desired improvement. μ(x) represents the mean of candidate point x, and σ(x) represents the standard deviation of candidate point x. The current optimal value of the function is log_h, which can be calculated using the following expression:

[0095]

[0096] Where c1 = log(2π) / 2, log1me(z) = log(1-exp(z)), and ∈ represents the numerical precision. c2=log(π / 2) / 2, erfcx(z)=exp(z 2 erfc(z). φ(z) is the probability density function of the standard normal distribution, z is the standardized variable in the standard normal distribution, Φ(z) is the cumulative distribution function of the standard normal distribution, and exp(·) is the exponential function. The definitions of φ(u) and u are the same as those of φ(z) and z, only the variable names are different. erfcx(·) represents the scaled form of the complementary error function, and erfc(·) is the complementary error function.

[0097] S3. Whenever g is divisible by the set constant interval, update the mapping matrix between any two tasks, calculate the similarity between each pair of tasks, select a suitable auxiliary task for each main task and migrate its optimal solution after mapping, and update the pheromone concentration.

[0098] First, a kernel-based autoencoder mechanism is used to update the mapping matrix between every two tasks, ensuring that the accuracy of the mapping improves with the number of actual evaluations of the tasks. There are two types of mapping matrices, recording the mappings between the decision spaces and the joint space, respectively. The joint space is formed by concatenating the decision space and the target space; that is, for two datasets D...A =(X A ,y A ) and D B =(X B ,y B The mapping is established by [X] A →X B [Upgraded to [X] A ,y A →X B ,y B ]. Where X A and X B Let y be the set of input variables. A and y B This is the set of target values ​​corresponding to the input variables.

[0099] Secondly, for each task, randomly select kN test points on the function's domain, and retain N sample points that are close to the known points and the current optimal point. Let D... candidatds ={x s |x s ~f i (x),s=1,2,…,kN},D candidates Let x be the set of candidate points obtained by randomly sampling on the domain of the function for task i. For each candidate point x∈D candidates The posterior standard deviation was calculated using a Gaussian process model. Let be the function value predicted by the Gaussian process model at candidate point x for task i, and Var[·] represent the variance estimate of the Gaussian process model at candidate point x. A smaller posterior standard deviation means a closer distance between the candidate point and the known point. Mapping the posterior standard deviation to a deterministic score yields:

[0100]

[0101] Where ∈ is a small positive number, used to avoid the denominator being zero. C certainty (x) is the deterministic score of candidate point x, P certainty (x) represents the importance of the prediction confidence of the Gaussian process model at candidate point x. For each candidate point x∈D candidates Calculate the current optimal point x from x to task i. cb The Euclidean distance d(x,x) cb )=||xx cb || 2 Similarly, distance is mapped to distance weight W. distance (x), and calculate the normalized probability:

[0102]

[0103] Where P distance (x) reflects the importance of x among all candidate points based on distance, W distance (x′) represents the distance weight of candidate point x′.

[0104] Introducing a weighting coefficient ω>0, the final joint probability of candidate point x is calculated:

[0105]

[0106] Where P combined (x) represents the final joint probability of candidate point x; P tmp (x) and P tmp (x′) represent the provisional joint probabilities at candidate points x and x′, combining prediction confidence and distance-based importance, respectively. According to the joint probability P... combined (x), from the candidate point set D candidates We will perform sampling without replacement to select N valid sampling points for task i to be used for similarity measurement:

[0107] D selected ={x s |x s ~P combined (x), i=1,2,…,N} (16)

[0108] D selected Let x represent the final set of selected sampling points. s This represents the s-th selected sampling point.

[0109] Finally, by evaluating these sampling points using a Gaussian process model, a similarity measurement dataset for task i can be obtained. For any two tasks T1 and T2 in a multi-task scenario, assume that the corresponding similarity measurement datasets for T1 and T2 are P and Q, respectively, with T2 being the primary task. Based on the existing real evaluation datasets for these two tasks, the mapping matrix of the joint space of P→Q can be calculated using a kernelized autoencoder mechanism, mapping P to another problem domain P′. This invention uses a multivariate Gaussian probability model to approximate the distributions of P′ and Q, and calculates the similarity between them using Kullback–Leibler divergence.

[0110]

[0111] Where Σ represents the covariance matrix and μ represents the mean vector. tr(·) and det(·) represent the rank and determinant of the matrix, respectively. P′ and Q represent two probability distributions, d represents the dimension of the data, and KLD(P′||Q) measures the difference between the two probability distributions.

[0112] Combining kernel-based task similarity measurement with pheromone concentration-based heuristics, the algorithm selects a suitable auxiliary task for each task and performs knowledge transfer. Let ξ(T) m ,T a ) as the main task T m And auxiliary tasks T a The similarity between them is obtained by taking the reciprocal of the Kullback-Leibler divergence between them and normalizing it:

[0113]

[0114] Where T represents T m Auxiliary task set. KLD(T) a ||T m ) represents the main task T m And auxiliary tasks T a Kullback–Leibler divergence between them, KLD(t||T) m ) represents the main task T m The Kullback–Leibler divergence between the auxiliary task t and the auxiliary task t. Let τ(T) m ,T a ) is T a Support T m The pheromone concentration can be used to obtain T. a Support T m The probability formula for knowledge transfer:

[0115]

[0116] Preset probability P select , used for tk P kt A secondary task is randomly selected from the main task. `rand(0,1)` returns a random floating-point number between 0 and 1. Thus, for the main task T... m The selection of auxiliary tasks for knowledge transfer can be represented as:

[0117]

[0118] Where KT(T) m ) represents the main task T m In the selection of auxiliary tasks during knowledge transfer, P select This represents the probability of selecting one task from the current tk tasks as an auxiliary task.

[0119] Determine the auxiliary task T for knowledge transfer a After that, T a The current optimal solution is transferred to the problem domain of the main task through a mapping. Note that the mapping here only includes T.a The decision space. The mapped solution will use T. m The task function is realistically evaluated and added to T. m In the solution set, let α d and α r This represents the reward and decay rate of pheromones. When knowledge transfer occurs, all pheromone concentrations will evaporate in some form. If the mapped solution is less than T... m If the current optimal solution is better, then T is considered better. a For T m It is beneficial, as it increases the concentration of pheromones. The specific formula is as follows:

[0120]

[0121] S4. If the termination condition is met, terminate and return the current optimal solution for each task; otherwise, proceed to step S2.

[0122] The above description of the preferred embodiments is quite specific and detailed, but it merely illustrates one feasible implementation of the present invention and is not intended to limit the scope of the invention. It should be noted that those skilled in the art can add several modifications or improvements based on these preferred embodiments within the framework of the present invention, but these are all within the protection scope of the present invention. The protection scope of the present invention should be determined by the appended claims.

Claims

1. An adaptive knowledge transfer method using multi-task Bayesian optimization, characterized in that, Includes the following steps: S1. Initialize the pheromone concentration between tasks, randomly generate n initial solutions for each task, and establish a separate Gaussian process model. S2. Begin evolution with a period of G. In the g-th generation, perform Bayesian optimization for each task and update the optimal solution. S3. If the current algebra g is divisible by the set threshold, update the mapping matrix between any two tasks, calculate the similarity between each pair of tasks, select a suitable auxiliary task for each main task and migrate its optimal solution after mapping, and update the pheromone concentration. S4. If the termination condition is met, terminate and return the current optimal solution for each task; otherwise, proceed to step S2.

2. The adaptive knowledge transfer method using multi-task Bayesian optimization according to claim 1, characterized in that, In step S1, the pheromone concentration is initialized as follows: τ(T m ,T a )=(Phe max +Phe min ) / 2 (1); Among them Phe max and Phe min τ(T) represents the upper and lower limits of pheromone concentration, respectively. m ,T a ) for auxiliary task T a Main Task Support T m The concentration of pheromones.

3. The adaptive knowledge transfer method using multi-task Bayesian optimization according to claim 1, characterized in that, In step S1, for the i-th task T i The expression for establishing the Gaussian process model is: f(x)~GP(m(x),k(x,x′))(2); where m(·) is the prior mean function, k(·) is the kernel function, and let the i-th task T be denoted as f(x)~GP(m(x),k(x,x′))(2); i Given a finite dataset (X, y), then x, x′∈X, and f(x) represents the i-th task T. i The function expression, GP(·) represents Gaussian process modeling, x represents the input variable, X represents the set of sampled input variables, and T i Let represent the i-th task in multi-task optimization, y represent the set of observations, and the mean and variance of the candidate point x′ are calculated using the following expressions: μ(x′)=k(x′;X)K(X;X) -1 y (3); σ 2 (x′)=k(x′;x′)-k(x′;X)K(X;X) -1 k(x′;X) T (4); where μ(x′) represents the mean of candidate point x′, σ 2 (x′) represents the variance of candidate point x′, k(x′;X) calculates the correlation between x′ and each x in X using a kernel function, and K(X;X) is the covariance matrix containing information about the internal correlations of X. The kernel function chosen for the Gaussian process is the Matérn kernel. Where v is the smoothing coefficient, l is the length scale parameter, and K... v It is the modified Bessel function, and k(x,x′) is the kernel function.

4. The adaptive knowledge transfer method of multi-task Bayesian optimization according to claim 1, characterized in that, In step S2, Bayesian optimization uses LogEI as the acquisition function, and the specific expression is as follows: Where x represents the input variable, LogEI y *(x) represents the sampling function that takes the logarithm of the desired improvement at candidate point x, log_h(.) is the function that performs a logarithmic transformation on a portion of the calculations in the desired improvement, μ(x) represents the mean of candidate point x, σ(x) represents the standard deviation of candidate point x, and y * The current optimal value of the function is log_h, which can be calculated using the following expression: Where c1 = log(2π) / 2, log1me(z) = log(1-exp(z)), ∈ represents the numerical precision, and φ(z) = c2=log(π / 2) / 2, erfcx(z)=exp(z 2 )erfc(z), φ(z) is the probability density function of the standard normal distribution, z is the standardized variable in the standard normal distribution, Φ(z) is the cumulative distribution function of the standard normal distribution, exp(·) is the exponential function, erfcx(·) represents the scaled form of the complementary error function, and erfc(·) is the complementary error function.

5. The adaptive knowledge transfer method of multi-task Bayesian optimization according to claim 1, characterized in that, In step S3, when g is divisible by interval, the specific process of adaptive knowledge transfer is as follows: S31. Calculate the mapping matrix between two tasks through a kernel-based self-encoding mechanism. There are two types of mapping matrices: one records the mapping between decision spaces and the other records the mapping between joint spaces. S32. For each task, randomly select kN test points on the domain of the function and retain N sampling points that are close to the known points and the current best point. Use Gaussian process regression to evaluate these sampling points to obtain the similarity measurement dataset of the task. S33. For each pair of auxiliary tasks and main tasks, the similarity measurement dataset of the auxiliary tasks is mapped to the problem domain of the main task. The distribution of the similarity measurement dataset of the auxiliary tasks and the main tasks is approximated by a multivariate Gaussian distribution model, and the similarity between the two is calculated by Kullback-Leibler divergence. S34. Combining kernel-based task similarity measurement and pheromone concentration-based heuristic rules, select auxiliary tasks for each task and transfer the current optimal solution of the auxiliary tasks to the problem domain of the main task through mapping. S35, Update pheromone concentration.

6. The adaptive knowledge transfer method of multi-task Bayesian optimization according to claim 5, characterized in that, In step S31, the mapping M obtained by the kernel-based automatic encoding mechanism k The closed-form solution is: M k =S2K(S1,S1) T (K(S1,S1)K(S1,S1) T ) -1 (8); Where S1 and S2 are the corresponding initial solution sets of the two optimization tasks T1 and T2, and the (m,n)th element of the kernel matrix K(S1,S2) is k(s1,S2). 1m ,s 1n ), κ(·,·) are kernel functions; s 1m and s 2m Let S1 and S2 represent the m-th solutions, respectively. If we want to map the optimal solution set PS of optimization task T1 to the problem domain of optimization task T2, denoted as MPS after mapping, the calculation method is as follows: MPS=M k K(S1,PS) (9)。 7. The adaptive knowledge transfer method of multi-task Bayesian optimization according to claim 5, characterized in that, In step S32, obtaining the similarity measurement dataset for the task includes the following steps: S321, Note D candidates ={x s |x s ~f i (x),s=1,2,…,kN},D candidates Let x be the set of candidate points obtained by randomly sampling on the domain of the function for task i, where x s Let f represent the s-th candidate point. i (x) represents the objective function of task i, and x represents the input variable of task i; For each candidate point x∈D candidates The posterior standard deviation was calculated using a Gaussian process model. Let Var[·] be the function value predicted by the Gaussian process model at candidate point x for task i, and Var[·] represent the variance estimate of the Gaussian process model at candidate point x. Mapping the posterior standard deviation to the deterministic score, we get: Where ∈ is a positive number, used to avoid the denominator being zero; C certainty (x) is the deterministic score of candidate point x, C certainty(x′) P is the deterministic score of candidate point x′. certainty (x) represents the importance of the prediction confidence of the Gaussian process model at candidate point x; S322, For each candidate point x∈D candidates Calculate the current optimal point x from x to task i. cb The Euclidean distance d(x,x) cb )=||xx cb || 2 Similarly, the distance is mapped to a distance weight W. distance (x), and calculate the normalized probability: Where P distance (x) reflects the importance of x among all candidate points based on distance, W distance (x′) represents the distance weight of candidate point x′; S323. Introduce a weighting coefficient ω>0 to calculate the final joint probability of candidate point x: P combined (x) represents the final joint probability of candidate point x; P tmp (x) and P tmp (x′) represent the provisional joint probability at candidate points x and x′, combining prediction confidence and distance-based importance, respectively; S324, According to the joint probability P combined (x), from the candidate point set D candidates We will perform sampling without replacement to select N valid sampling points for task i to be used for similarity measurement: D selected ={x s |x s ~P combined (x),i=1,2,…,N} (16); D selected Let x represent the final set of selected sampling points. s This represents the s-th selected sampling point.

8. The adaptive knowledge transfer method of multi-task Bayesian optimization according to claim 5, characterized in that, The mapping in step S33 is a mapping between joint spaces, which are formed by concatenating the decision space and the target space. That is, for two datasets D... A =(X A ,y A ) and D B =(X2,y B ), X A and X B Let y be the set of input variables. A and y B The mapping is established for the set of target values ​​corresponding to input variables by [X] A →X B Upgraded to [X] A ,y A →X B ,y B Furthermore, the formula for calculating the Kullback–Leibler divergence is as follows: Where Σ represents the covariance matrix, μ P′ μ Q All represent the average vector, tr(·) and det(·) represent the rank and determinant of the matrix, respectively, P′ and Q represent two probability distributions, d represents the dimension of the data, and KLD(P′||Q) measures the difference between the two probability distributions; The main task T is obtained by taking the reciprocal and normalizing. m And auxiliary tasks T a The similarity ξ(T) between them m ,T a ): Where T represents the main task T m Auxiliary task set, KLD(T) a ||T m ) represents the main task T m And auxiliary tasks T a Kullback–Leibler divergence between them, KLD(t||T) m ) represents the main task T m Kullback–Leibler divergence between the auxiliary task t and the auxiliary task t.

9. The adaptive knowledge transfer method of multi-task Bayesian optimization according to claim 5, characterized in that, The method for selecting the auxiliary task in step S34 is as follows: Where τ(T) m ,T a ) for auxiliary task T a Main Task Support T m pheromone concentration, P kt (T m ,T a ) is T a Support T m The probability of knowledge transfer, ξ(T) m ,T a ) as the main task T m And auxiliary tasks T a The similarity between them; T represents the main task T m Auxiliary task set; pre-set probability P select Used in tk P kt Randomly select an auxiliary task from the main task; rand(0,1) returns a random floating-point number between 0 and 1; for the main task T m The auxiliary task selection for knowledge transfer is represented as: Where KT(T) m ) represents the main task T m In the selection of auxiliary tasks during knowledge transfer, P select This represents the probability of selecting one task from the current τk tasks as an auxiliary task.

10. The adaptive knowledge transfer method of multi-task Bayesian optimization according to claim 5, characterized in that, The method for updating pheromone concentration in step S35 is as follows: Where τ(T) m ,T a ) for auxiliary task T a Main Task Support T m pheromone concentration, α d and α r The reward and decay rates of pheromones.

Citation Information

Patent Citations

  • Aerial power component image classification method based on knowledge transfer learning

    CN110472545A

  • Smart financial market trend prediction method based on federal multitask optimization

    CN120031664A

  • Photovoltaic system performance optimization method, system, equipment and medium

    CN120235030A

  • Systems and methods for black-box optimization

    WO2018222203A1

  • Multi-task hyperparameter optimization method for deep neural network, and device

    WO2020252766A1