Multi-target Bayesian optimization method and device, electronic equipment and storage medium
By using dynamic threshold classification and weighted cross-entropy loss training classifiers in the multi-objective Bayesian optimization method, the problem of difficulty in fully covering the Pareto frontier and high computational cost in the prior art is solved, and more efficient multi-objective optimization is achieved.
Patent Information
- Application Number
- CN202510105724.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-23
AI Technical Summary
When dealing with multi-objective optimization problems, the existing multi-objective Bayesian optimization method is difficult to fully cover the Pareto frontier, and the calculation cost is high, so it is impossible to effectively utilize better acquisition functions.
The observation point set is classified using dynamic thresholds to ensure that the positive class completely covers the Pareto set, and the classifier is trained through weighted cross-entropy loss, so that the density ratio estimation method is suitable for obtaining functions of any utility.
It achieves comprehensive coverage of the Pareto frontier, reduces computational costs, and can effectively utilize better acquisition functions, improving the efficiency and performance of multi-objective Bayesian optimization.
Smart Images

Figure CN120067851A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a multi-objective Bayesian optimization method, apparatus, electronic device, and storage medium. Background Art
[0002] Bayesian Optimization (BO) is an efficient optimization method for expensive black-box functions, which can find an input strategy that optimizes the black-box function. During the optimization process, BO will use a set of initial observations to obtain the prior information of the function f. And through a surrogate model to capture the prior distribution p(y|x,D n ), and provide it to the acquisition function for calculating the benefit brought by different inputs x to the model, so as to obtain the next observation point x i+1 . The new observation point is used to update and p(y|x,D n+1 ). After a certain number of updates, the surrogate model can be used to estimate the black-box function f and estimate the input strategy x that maximizes f. Since BO uses the prior information of the observation points during the optimization process and continuously uses the acquisition function to evaluate the utility of the input strategy, compared with black-box optimization methods such as grid search and random search, BO has a faster optimization speed and better optimization ability.
[0003] BO models the set of observation points through a surrogate model to obtain the prior distribution of the objective function and applies it to the evaluation of the acquisition function for any input strategy. The role of the acquisition function is to select an optimal input strategy from the current perspective based on the prior distribution. The two commonly used acquisition functions are PI (Probability of Improvement) and EI (Expectation of Improvement). As shown below, PI and EI respectively evaluate the probability of improvement and the expected value of improvement of the function value relative to the threshold. Where p(y|x) is the conditional probability distribution of y with respect to x, τ is the threshold, and U(y) is the utility function used to evaluate the quality of y in the optimization problem. The utility function of PI represents the probability of improvement of y relative to y * , and the utility function of EI represents the expected improvement of y relative to τ.
[0004]
[0005] In a multi-objective optimization problem, the optimization goal refers to Pareto optimality, that is, calculating an optimal hyperplane. In a min F problem, for any point on the optimal hyperplane P + on There must not exist a more optimal point y such that where o is the dimension of y. The hypervolume (HV) between the optimal hyperplane and the reference point is usually used to evaluate the optimization result. To make the optimal hyperplane cover more solutions, the optimization problem is transformed into maximizing the hypervolume of the objective space. Therefore, a common Bayesian optimization idea is to use a scalar to evaluate the output of multiple objectives such as the size of the hypervolume, and then use a single-objective optimization method for optimization. Correspondingly, some acquisition functions applicable to multi-objective optimization problems have also been proposed. For example, the combination of hypervolume with EI and PI, EHVI and PHVI, are defined as follows, where is the Pareto set, and HV(·) represents the hypervolume constructed by a point set and the reference point.
[0006]
[0007]
[0008] As Figure 1 shown, the related technology adopts a Bayesian optimization framework based on density ratio estimation, which was proposed in 2022. It performs scalarization operations on multi-dimensional optimization objectives, extending the density ratio estimation method from single-objective problems to multi-objective problems. This framework directly estimates the acquisition function through the classifier π(x), successfully solving the problem of high computational cost caused by GPR (Gaussian Process Regression) in traditional methods, as well as the problem of error accumulation and high-dimensional inadaptability that may occur in TPE (Tree Parzen Estimator). It demonstrates the rationality of directly estimating the density ratio using a classifier from a theoretical level, enabling a classifier to simultaneously undertake the functions of a surrogate model and an acquisition function. Compared with other multi-objective Bayesian optimization methods (such as maximum entropy search), this framework uses simple estimator training to replace the complex acquisition function calculation process, greatly improving the optimization efficiency.
[0009] Before starting the optimization, the framework first samples a set of observation points from the target space. Given the multi-dimensional nature of the output, a scalarizer is used to simplify the original problem into a single-objective optimization problem. Subsequently, a preset threshold is used to classify the set of observation points, and the two types of observation points are input into the classifier to fit the sample data. The classifier integrates the surrogate model and the acquisition function, and can directly estimate the density ratio to determine the optimal x value. Subsequently, the set of observation points is continuously expanded, and the optimization of the multi-objective problem is achieved by modeling a more accurate probability distribution. In this process, the core step is to ensure the estimation accuracy of the classifier probability for the acquisition function, that is, the acquisition function can be represented by the density ratio, and the density ratio can be estimated by the posterior probability output by the classifier. The whole process is shown in Equation (1).
[0010]
[0011] where and respectively represent the conditional probability density of x with respect to y, τ is the threshold, and γ is the positive and negative sample ratio divided according to it.
[0012] However, the framework itself faces several limitations that restrict its performance in ATS scenarios and architecture parameter optimization tasks. First, this framework is directly extended from single-objective optimization methods, so it still depends on the selection of the threshold; second, the framework depends on the classifier's estimation of the acquisition function. Although it gives a theoretical proof of this process, in its conclusion, other types of acquisition functions (such as EI) are all treated as PI, which means that the framework cannot utilize more effective acquisition functions. The following is a detailed description of the two problems.
[0013] (1) Fixed threshold problem
[0014] In Equation (1), the acquisition function is proportional to the conditional probability density ratio of x, and the threshold τ is used to divide the observation points to calculate the conditional probability density and In the context of single-objective optimization, due to the uniqueness of the optimal solution, the classification result can ensure that the optimal solution is covered, and at the same time is used to balance exploration and exploitation. However, when dealing with multi-objective problems, the Pareto front contains multiple solutions. If a single threshold is used for simple division, it may be difficult to fully cover the Pareto front.
[0015] Figure 2 shows a case of observation point classification in a two-dimensional optimization problem. Both dimensions are optimized in the direction of minf. The orange is the Pareto front of the current set of observation points. The hypervolume of each point is calculated and sorted, and the 1 / 3 quantile value is used as the division threshold to classify the observation points. As Figure 2As shown in the left half, when the number of observation points is small, one of the Pareto points cannot be classified into the positive samples; in Figure 2 In the right half, as the number of observation points increases, some non-Pareto points are classified into the positive samples. It can be seen that the division result of observation points with a fixed threshold cannot cover the Pareto set, and the noise in the classification result will cause the convergence effect of the optimization method to deteriorate (or the convergence speed to slow down).
[0016] The classification with a fixed threshold also has deficiencies in terms of efficiency. When new observation points are added to the set, the hypervolume improvement of the original observation points may change accordingly, and it is necessary to re-perform the scalarization process on all points. For high-dimensional tasks with a large number of sub-problems, this repeated scalarization operation is the key to the bottleneck of optimization efficiency.
[0017] (2) Adaptation problem of non-PI acquisition function
[0018] As mentioned above, the Bayesian optimization framework based on density ratio estimation depends on a premise, that is, the density ratio formula represented by Equation (1). However, in single-objective optimization methods, there is an error in the proof process of this formula, and its conclusion only applies to specific types of acquisition functions, that is, the case where the acquisition function is PI. The multi-objective optimization framework inherits this error, which means that this framework cannot use other better acquisition functions. Some studies have pointed out that EHVI can evaluate the improvement expectation of candidate points relative to the threshold, and can better balance the accuracy of the surrogate model and the exploration ability of the optimization method, and is a more effective acquisition function than PHVI.
[0019] Figure 3 Shows the distribution of PHVI and EHVI for a multi-objective optimization problem DTLZ4 (o = 2, d = 2), where x1 and x2 are the two dimensions of the solution space respectively. By randomly selecting 500 observation points, the PHVI values and EHVI values in the solution space are calculated and represented on the Figure 3 (a) and Figure 3 (b) ordinates. It can be seen that near the solution space where x1 = 0, there is a region where the PHVI value is 1 but the EHVI value tends to 0, which means that the candidate points selected there can only extremely limitedly improve the optimization result. The practice of converting different types of acquisition functions into PHVI in the existing framework will limit the optimization performance of the method. Based on the above analysis, the errors in the existing density ratio estimation methods will be described below.
[0020] The existing method divides the observation value set D N into two categories according to the threshold τ, where τ is the γ-th quantile value of the observed y sequence, that is, γ = Φ(τ) = p(y ≤ τ). Thus, two probability density functions and The calculation process of the acquisition function with EI as the utility is as follows:
[0021]
[0022] where α(x; D N , τ) is the acquisition function, and the numerator and denominator of Equation (2) are respectively:
[0023]
[0024] Therefore, we have:
[0025]
[0026] However, in the process of calculating the numerator above, p(x|y, D N ) is regarded as a quantity independent of the distribution of y, and this value is directly extracted outside the equation. In fact, as shown in Equation (6), due to the existence of the utility function (τ - y), the distribution of y in D N has been disrupted, so directly extracting p(x|y, D N ) as is not advisable, and this error has also been mentioned in some papers. It should be noted that the above reasoning process holds when PI is used as the utility function because there is no term (τ - y) in the proof process of PI, so it will not change the distribution of y in d N .
[0027] Summary of the Invention
[0028] The main objective of the embodiments of the present invention is to propose a multi-objective Bayesian optimization method, device, electronic device, and storage medium, which can ensure that the classification results can comprehensively cover the Pareto front and reduce the calculation cost to a certain extent.
[0029] To achieve the above objective, on the one hand, the embodiments of the present invention propose a multi-objective Bayesian optimization method, including the following steps:
[0030] In the multi-objective Bayesian optimization process, a dynamic threshold is used to classify the set of observation points, and the points on the Pareto front in the observation points are regarded as positive samples, and the remaining points are regarded as negative samples, so that the positive class completely covers the Pareto set;
[0031] The density ratio estimation is extended to multi-objective Bayesian optimization, and the classifier is trained through weighted cross-entropy loss, so that the density ratio estimation method can be applicable to the acquisition function for any utility;
[0032] Complete the multi-objective Bayesian optimization process and output the Pareto set; and apply the multi-objective Bayesian optimization process to the multi-objective recognition of image processing tasks to obtain the multi-objective recognition results;
[0033] Among them, in the multi-objective Bayesian optimization process, a dynamic threshold is used to classify the set of observation points, including the following steps:
[0034] Sample n initial observation points in the solution space to generate an initial set of observation points;
[0035] Obtain the Pareto set of each observation point in the initial set of observation points;
[0036] Calculate the utility of each observation point;
[0037] Assign a class label to each observation point, and the classification basis is whether it is in the Pareto set;
[0038] The extension of density ratio estimation to multi-objective Bayesian optimization, training a classifier through weighted cross-entropy loss, includes the following steps:
[0039] Use the utility as the weight of the observation point to train the classifier;
[0040] Use the posterior probability provided by the classifier for optimization to determine the next observation point;
[0041] Evaluate the new observation point and add it to the set of observation points.
[0042] In some embodiments, the sampling of n initial observation points in the solution space to generate an initial set of observation points includes the following steps:
[0043] Sample a set of data from the solution space using random sampling or grid sampling The definition formula for the size of the initial observation set is:
[0044] Input each point in the initial observation set into the black box function F to obtain its observation value D n : D n ←{(x i ,F(x i ))}.
[0045] In some embodiments, the obtaining of the Pareto set of each observation point in the initial set of observation points is specifically: by judging the dominance relationship, determining the Pareto set and the Pareto front in the initial set of observation points; among them, for the min F problem, the judgment criterion for its Pareto set is: for any point on the Pareto front P + on it there must not exist a better point y such that where o is the dimension of y.
[0046] In some embodiments, calculating the utility of each observation point includes the following steps:
[0047] If the acquisition function is PHVI, there is no need to perform a scalarization operation, and directly assign a 0-1 utility value to each point according to the Pareto set;
[0048] If it is other acquisition functions, including EHVI, perform a scalarization operation on the observation points within the Pareto set;
[0049] Among them, the utility u of the observation point of PHVI i and the utility u of the observation point of EHVI i are calculated as shown in the following equations:
[0050]
[0051] Among them, P + is the obtained Pareto set, and P - is the non-Pareto set; HV() represents the hypervolume constructed by a point set and a reference point.
[0052] In some embodiments, in the step of assigning a class label to each observation point, there is no need to sort the utility values. Classify the observation point set, classify the observation points with non-zero utility values into the positive class, and classify the observation points with zero utility values into the negative class;
[0053] Training the classifier with the utility as the weight of the observation point includes the following steps:
[0054] Use the classifier to fit the observation point set, where it varies according to the acquisition function adopted;
[0055] According to the following loss function, use the posterior probability output by the classifier as an estimate of the density ratio, so as to realize the estimation of the acquisition function; loss function The expression is:
[0056]
[0057] Among them, θ is the classifier parameter; and respectively represent the conditional probability density of x with respect to y under two y values; U(s) represents a utility function independent of x; β represents the positive sample ratio; K represents a constant;
[0058] Use XGBoost, random forest, and multi-layer perceptron as classifiers. After training, use the obtained posterior probability as an estimate of the acquisition function to evaluate candidate points.
[0059] In some embodiments, optimizing by using the posterior probability provided by the classifier to determine the next observation point includes the following steps:
[0060] Given a set of candidate solutions, output their acquisition function values through the classifier, and determine a point x that maximizes the acquisition function through optimization in the solution space t , and use it as the next observation point of the black-box function;
[0061] Evaluating the new observation point and adding it to the set of observation points includes the following steps:
[0062] Use the black-box function F to evaluate the obtained best observation point x t , add it to the set of observation points, and for each newly obtained observation point, expand the set of observation points through the acquisition function.
[0063] In some embodiments, in the process of classifying the set of observation points using a dynamic threshold, there is no need to perform scalarization processing on each observation point, and only the dominated situation of each observation point needs to be judged; the dominance relationship of the observation points is realized by comparing the magnitudes of each component y i i ∈ (1, 2,..., q);
[0064] In the process of training the classifier through weighted cross-entropy loss, before performing density ratio estimation, first use a scalar to quantify the output of the black-box function F, that is to obtain the set of observed values U(s) represents a utility function independent of x, p(s|x, D N ) represents the prior distribution provided by the surrogate model, P + and P - respectively represent the Pareto set and non-Pareto set of the function output, and the acquisition function α(x; D N ) represents the expectation of each input point x with respect to the utility function, where the utility function represents the degree of improvement of the hypervolume by the addition of candidate points; the weighted cross-entropy loss is used to maximize the log-likelihood of dr(x) to model the density ratio equation including the coefficient w.
[0065] Another aspect of the embodiments of the present invention also provides a multi-objective Bayesian optimization device, including:
[0066] A first module, configured to, in the process of multi-objective Bayesian optimization, classify the set of observation points using a dynamic threshold, regard the points located on the Pareto front in the observation points as positive samples, and the remaining points as negative samples, so that the positive class completely covers the Pareto set;
[0067] A second module is used to extend the density ratio estimation to multi-objective Bayesian optimization, and train a classifier through weighted cross-entropy loss, so that the density ratio estimation method can be applicable to acquisition functions for any utility;
[0068] A third module is used to complete the multi-objective Bayesian optimization process and output the Pareto set; and apply the multi-objective Bayesian optimization process to multi-object recognition in image processing tasks to obtain multi-object recognition results;
[0069] Wherein, in the multi-objective Bayesian optimization process, a dynamic threshold is used to classify the set of observation points, including the following steps:
[0070] Sample n initial observation points in the solution space to generate an initial set of observation points;
[0071] Obtain the Pareto set of each observation point in the initial set of observation points;
[0072] Calculate the utility of each observation point;
[0073] Assign a class label to each observation point, and the classification basis is whether it is in the Pareto set;
[0074] The extension of the density ratio estimation to multi-objective Bayesian optimization and the training of the classifier through weighted cross-entropy loss include the following steps:
[0075] Use the utility as the weight of the observation point to train the classifier;
[0076] Use the posterior probability provided by the classifier for optimization to determine the next observation point;
[0077] Evaluate the new observation point and add it to the set of observation points.
[0078] To achieve the above object, another aspect of the embodiments of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described above is implemented.
[0079] To achieve the above object, another aspect of the embodiments of the present invention provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0080] The embodiments of the present invention also disclose a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device executes the method described above.
[0081] The embodiments of the present invention at least include the following beneficial effects: The present invention provides a multi-objective Bayesian optimization method, device, electronic device, and storage medium. In this solution, during the multi-objective Bayesian optimization process, a dynamic threshold is used to classify the set of observation points. The points on the Pareto front in the observation points are regarded as positive samples, and the remaining points are regarded as negative samples, so that the positive class completely covers the Pareto set. The density ratio estimation is extended to multi-objective Bayesian optimization, and a classifier is trained through weighted cross-entropy loss, so that the density ratio estimation method can be applied to acquisition functions for any utility. The multi-objective Bayesian optimization process is completed, and the Pareto set is output. Then, the multi-objective Bayesian optimization process is applied to the multi-objective recognition in image processing tasks to obtain multi-objective recognition results. The embodiments of the present invention can ensure that the classification results can comprehensively cover the Pareto front and, to a certain extent, reduce the calculation cost. Description of the Drawings
[0082] Figure 1 It is a schematic diagram of a multi-objective Bayesian optimization framework based on density ratio estimation provided by an embodiment of the present invention;
[0083] Figure 2 It is an example diagram of the classification result of observation points using a fixed threshold provided by an embodiment of the present invention;
[0084] Figure 3 It is an example of the difference between the PI and EI acquisition functions provided by an embodiment of the present invention, where Figure 3 (a) represents the PI acquisition function, Figure 3 (b) represents the EI acquisition function;
[0085] Figure 4 It is a flowchart of the overall steps provided by an embodiment of the present invention;
[0086] Figure 5 It is an example diagram of the classification result of observation points using a dynamic threshold provided by an embodiment of the present invention;
[0087] Figure 6 It is a comparison diagram of the optimization performance of different optimization methods for the DTLZ problem provided by an embodiment of the present invention;
[0088] Figure 7 It is a schematic diagram of the cumulative calculation time of multiple Bayesian optimization algorithms provided by an embodiment of the present invention. Detailed Embodiments
[0089] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention detailed in the appended claims.
[0090] It can be understood that the terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present invention and the above accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0091] It should be understood that in the present invention, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0092] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0093] Before the embodiments of the present invention are described in detail, the meanings of some relevant variables involved in the embodiments of the present invention are described as follows:
[0094] Black-box function to be optimized;
[0095] Set of observations, where (x i , y i ) is an observation point and n is the number of observation points in the set;
[0096] p(y|x, D n ): Prior probability distribution of y with respect to x in the set of observation points;
[0097] Acquisition function, used to calculate the benefit brought by different inputs x to the model;
[0098] U(y): Utility function, used to evaluate the quality of y values in the optimization problem;
[0099] τ: Threshold used to evaluate the value of observation points in the acquisition function;
[0100] π(x): Classifier, used to estimate the acquisition function;
[0101] Respectively represent the conditional probability density of x with respect to y;
[0102] γ: Proportion of positive and negative samples of observation points divided by the threshold τ;
[0103] Scalarizer, used to convert high-dimensional points into a scalar;
[0104] P + , P - : Respectively represent the Pareto set and the non-Pareto set;
[0105] D′ N : A newly constructed set of observation points;
[0106] z, z′: Respectively represent the classification labels of two sets of observation points;
[0107] dr(x): Density ratio, used to represent the acquisition function;
[0108] w: Weight coefficient, used during the training of the classifier.
[0109] The multi-objective Bayesian optimization method, device, electronic device, and storage medium provided by the embodiments of the present invention relate to the field of computer technology. The multi-objective Bayesian optimization method provided by the embodiments of the present invention can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the multi-objective Bayesian optimization method, etc., but is not limited to the above forms.
[0110] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0111] Refer to Figure 4 , Figure 4 is a flowchart of the multi-objective Bayesian optimization method applied to a server provided by the embodiments of the present invention. The execution subject of this method can be any of the aforementioned computer devices (including servers or terminals). Refer to Figure 4 , this method may include the following steps:
[0112] During the multi-objective Bayesian optimization process, a dynamic threshold is used to classify the set of observation points. The points located on the Pareto front in the observation points are regarded as positive samples, and the remaining points are regarded as negative samples, so that the positive class completely covers the Pareto set;
[0113] Extend the density ratio estimation to multi-objective Bayesian optimization, and train a classifier through weighted cross-entropy loss, so that the density ratio estimation method can be applied to acquisition functions for any utility;
[0114] Complete the multi-objective Bayesian optimization process and output the Pareto set; and apply the multi-objective Bayesian optimization process to the multi-object recognition of image processing tasks to obtain multi-object recognition results;
[0115] Among them, in the multi-objective Bayesian optimization process, classifying the set of observation points using a dynamic threshold includes the following steps:
[0116] Sample n initial observation points in the solution space to generate an initial set of observation points;
[0117] Obtain the Pareto set of each observation point in the initial set of observation points;
[0118] Calculate the utility of each observation point;
[0119] Assign a class label to each observation point, and the classification basis is whether it is in the Pareto set;
[0120] The extension of density ratio estimation to multi-objective Bayesian optimization, training a classifier through weighted cross-entropy loss, includes the following steps:
[0121] Use the utility as the weight of the observation point to train the classifier;
[0122] Use the posterior probability provided by the classifier to perform optimization to determine the next observation point;
[0123] Evaluate the new observation point and add it to the set of observation points.
[0124] In some embodiments, the sampling of n initial observation points in the solution space to generate an initial set of observation points includes the following steps:
[0125] Sample a set of data from the solution space using random sampling or grid sampling The definition formula for the size of the initial observation set is:
[0126] Input each point in the initial observation set into the black box function F to obtain its observation value D n : D n ←{(x i ,F(x i ))}.
[0127] In some embodiments, obtaining the Pareto set of each observation point in the initial observation point set specifically includes: determining the Pareto set and the Pareto front in the initial observation point set by judging the dominance relationship; wherein, for the min F problem, the judgment criterion for its Pareto set is: for any point on the Pareto front P + there must not exist a more optimal point y such that where o is the dimension of y. where o is the dimension of y.
[0128] In some embodiments, calculating the utility of each observation point includes the following steps:
[0129] If the acquisition function is PHVI, there is no need to perform a scalarization operation, and directly assign 0-1 utility values to each point according to the Pareto set;
[0130] If it is other acquisition functions, including EHVI, perform a scalarization operation on the observation points within the Pareto set;
[0131] where the utility u of the observation point of PHVI i and the utility u of the observation point of EHVI i are calculated as shown in the following equations:
[0132]
[0133] where P + is the obtained Pareto set, and P - is the non-Pareto set; HV() represents the hypervolume constructed by a point set and a reference point.
[0134] In some embodiments, in the step of assigning category labels to each observation point, there is no need to sort the utility values, classify the observation point set, and classify the observation points with non-zero utility values into the positive class and the observation points with zero utility values into the negative class;
[0135] Training the classifier with the utility as the observation point weight includes the following steps:
[0136] Fitting the observation point set using the classifier, where it varies according to the adopted acquisition function;
[0137] According to the following loss function, use the posterior probability output by the classifier as an estimate of the density ratio to estimate the acquisition function; the loss function is expressed as:
[0138]
[0139] where θ is the classifier parameter; and respectively represent the conditional probability density of x with respect to y under two y values; U(s) represents a utility function independent of x; β represents the positive sample ratio; K represents a constant;
[0140] Using XGBoost, random forest, and multi-layer perceptron as classifiers, after training is completed, the posterior probability obtained from the training is used as the estimated value of the acquisition function to evaluate candidate points.
[0141] In some embodiments, the method of using the posterior probability provided by the classifier to perform optimization and determine the next observation point includes the following steps:
[0142] Given a set of candidate solutions, the acquisition function value is output through the classifier, and a point x that maximizes the acquisition function is determined through optimization in the solution space t , and it is used as the next observation point of the black-box function;
[0143] The method of evaluating the new observation point and adding it to the set of observation points includes the following steps:
[0144] Use the black-box function F to evaluate the obtained best observation point x t , add it to the set of observation points, and for each newly obtained observation point, expand the set of observation points through the acquisition function.
[0145] In some embodiments, in the process of classifying the set of observation points using a dynamic threshold, there is no need to perform scalarization processing on each observation point, and only the dominance situation of each observation point needs to be judged; the dominance relationship of the observation points is realized by comparing the magnitudes of each component y i i ∈ (1, 2, …, q);
[0146] In the process of training the classifier through weighted cross-entropy loss, before performing density ratio estimation, first use a scalar to quantize the output of the black-box function F, that is to obtain a set of observed values U(s) represents a utility function independent of x, p(s|x, D N ) represents the prior distribution provided by the surrogate model, P + and P - respectively represent the Pareto set and non-Pareto set of the function output, and the acquisition function α(x; D N ) represents the expectation of each input point x with respect to the utility function, where the utility function represents the degree of improvement of the hypervolume by the addition of candidate points; the weighted cross-entropy loss is used to maximize the log-likelihood of dr(x) to model the density ratio equation including the coefficient w.
[0147] The following describes in detail the implementation process of the method of the embodiments of the present invention in a specific application scenario:
[0148] To address the problems existing in the prior art, an embodiment of the present invention proposes a multi-objective Bayesian optimization method based on dynamic threshold density-ratio estimation (Dynamic Density-Ratio Estimation for Multi-objective Bayesian Optimization, DR.EMO). In this method, a dynamic threshold is first adopted to ensure that the classification results can comprehensively cover the Pareto front while reducing the computational cost to a certain extent. Secondly, a density-ratio estimation method applicable to all types of acquisition functions is developed, and its correctness is strictly proven theoretically.
[0149] (1) Dynamic threshold classification:
[0150] To avoid the problem that the classification results cannot cover the Pareto set, DR.EMO uses a dynamic threshold to classify the set of observation points. During the optimization process, only the points on the Pareto front have non-zero hypervolume contribution values, and the points not in the Pareto set do not participate in the hypervolume calculation. Therefore, DR.EMO regards the points on the Pareto front in the observation points as positive samples and the remaining points as negative samples. During the optimization process, as the set of observation points expands, the proportion of positive and negative samples changes continuously. Therefore, this classification is based on dynamic threshold classification. As Figure 5 shown, the dynamic threshold classification results can completely cover the Pareto set without being affected by the number of observation points, which enables DR.EMO to use the Pareto set as the optimization direction to improve the convergence speed.
[0151] In the actual classification process, DR.EMO does not need to perform scalarization processing on each observation point, but only needs to judge the domination situation of each observation point. Due to the complexity of high-dimensional hypervolume calculation, another advantage of the dynamic threshold is the high efficiency of calculation. There is no polynomial-time complexity algorithm for high-dimensional hypervolume calculation. Although there are many estimation algorithms, such as the WFG algorithm, its time complexity still reaches O(n q-2 logn), where n and q represent the number of points and the dimension of the Pareto set respectively. Therefore, scalarization processing is usually the main reason for the bottleneck of density-ratio estimation efficiency.
[0152] In DR.EMO, the domination relationship of observation points is realized by comparing the magnitudes of each component y i i∈(1,2,…,q). Facing the initial set of observation points, the time complexity of using dynamic threshold classification is O(n 2q); During the optimization process, as the observation point set expands and a new observation point is added, only the dominance relationship between it and each point in the Pareto set needs to be judged to determine its category, and the time complexity is O(n′q), where n′ is the size of the Pareto set. Thus, when using PHVI as the acquisition function, DR.EMO can avoid performing scalarization operations; when using other acquisition functions, DR.EMO can avoid performing scalarization operations on non-Pareto sets to improve the optimization efficiency.
[0153] (2) Non-PI density ratio estimation:
[0154] Inspired by Figure 1 the framework shown, DR.EMO extends density ratio estimation to multi-objective Bayesian optimization. By training a classifier with weighted cross-entropy loss, it avoids the drawback that density ratio estimation is not applicable to non-PI utility functions. The following gives its theoretical derivation process.
[0155] Before DR.EMO performs density ratio estimation, it first uses a scalar to quantify the output of the black-box function F, that is, to obtain the observation value set U(s) represents a utility function independent of x, and p(s|x,D N ) represents the prior distribution provided by the surrogate model, and P + and P - represent the Pareto set and non-Pareto set of the function output respectively. The acquisition function α(x;D N ) represents the expectation of each input point x with respect to the utility function, where the utility function represents the degree of improvement of the hypervolume by the addition of candidate points, as shown in Equation (7).
[0156]
[0157] Referring to Equation (3), the denominator can be easily proven. In DR.EMO, the classification ratio and probability density of observation points are expressed as: γ := p(y ∈ P + ,D N ),
[0158] Since U(s) changes the distribution of s in D N , the p(x|s,D N ) inside the numerator integral cannot be directly taken out. Before calculating the numerator, as shown in Equation (8), a transformation can be performed on it to eliminate the influence of the U(s) term.
[0159]
[0160] Here D′N For a newly constructed set of observation points, there is the following relationship with the original set of observation points:
[0161]
[0162] where K is a constant used to maintain the normalization of p(s|D′ N ). Since x and s appear in pairs in the set of observation points, when y ∈ P + , the probabilities of x in the new and old sets of observation points also have the following relationship:
[0163] p(x|D′ N ) = K · U(s)p(x|D N )#(10)
[0164] Also, since D′ N only changes the distribution of s, the conditional probability of x with respect to s and D′ N will remain unchanged, that is:
[0165] p(x|s,D N ) = p(x|s,D′ N )#(11)
[0166] From equations (9 - 11), the following equalities can be obtained. Equation (12) shows that the transformation of equation (8) does not change the proportion of the Pareto set in the observation points, and equation (13) shows that this transformation will change the conditional probability of x in the Pareto set.
[0167]
[0168] Substituting equations (12 - 13) into the numerator of equation (7), we have:
[0169]
[0170] In summary, the acquisition function can be reconstructed into the density ratio form shown in equation (15). Compared with the results in the related existing technologies, the density ratio proposed in this section uses to characterize the non - 1 utility of the candidate points. By constructing an auxiliary concept D′ N , the density ratio representation of any acquisition function is realized.
[0171]
[0172] On this basis, the class posterior probability is used to represent the density ratio. This section extends this method to multi-objective Bayesian optimization and corrects the error regarding the density ratio. DR.EMO estimates the density ratio using an unfixed threshold, divides the observed points into the Pareto set and the non-Pareto set, and z and z′ represent the classification labels of the two sets of observed points. Naturally, we have: p(x|y∈P + ,D N ) = p(x|z = 1), p(x|y∈P + ,D′ N ) = p′(x|z = 1), and at the same time γ = p(z = 1) = p′(z = 1) also holds. Therefore, the process of representing the density ratio dr(x) through the class posterior probability is as follows:
[0173]
[0174] According to equation (10), we have: Therefore, a coefficient can be set:
[0175]
[0176] to represent In addition, the class-conditional posterior probabilities of the two sets of observed points with respect to x are equal, that is: p′(z = 1|x) = p(z = 1|x), and the proof is as follows:
[0177]
[0178] where p(x|s,D′ N ) = p(x|s,D N ) always holds; when y∈P - , we have p(s|D′ N ) = p(s|D N ); and at the same time Therefore, we have:
[0179]
[0180] According to equations (16 - 18), the density ratio can be expressed as the class-conditional posterior probability with respect to x:
[0181] dr(x) = w·p(z = 1|x)#(19)
[0182] Equation (19) shows that the density ratio can be estimated through the conditional posterior probability. DR.EMO uses a classification function to estimate the density ratio, and this classification function is expressed as where θ is the classifier parameter. To model the density ratio equation containing the coefficient w, the weighted cross-entropy loss is used to maximize the log-likelihood of dr(x), and the loss function is as follows:
[0183]
[0184] Considering that the number of positive and negative samples may be unbalanced, the set of observation points can be enhanced by the sample ratio during actual application. At this time, there is:
[0185]
[0186] where N + , N - respectively represent the numbers of positive and negative samples. Finally, dr(x) can be estimated through , where
[0187] As shown below, when the acquisition function is PI, equation (22) is a conventional unweighted cross-entropy. This also indicates that the original density ratio estimation method is only applicable to PI. When other acquisition functions are adopted, it is the correct approach to estimate the density ratio according to equations (15) and (21).
[0188]
[0189] In summary, the specific process of the method in the embodiment of the present invention is as follows:
[0190] Aiming at the problem that the fixed threshold classification result is difficult to provide a good optimization direction, DR.EMO uses a dynamic threshold for classification, so that the positive class completely covers the Pareto set; aiming at the problem that the non-PI acquisition function cannot be estimated, DR.EMO uses weighted cross-entropy to train the classifier, so that the density ratio estimation method can be applicable to acquisition functions for any utility. Table 1 shows the algorithm process of DR.EMO. After T optimizations, DR.EMO can obtain the Pareto set of problem F in the solution space under T + n groups of finite observation points.
[0191] Table 1 DR.EMO algorithm process
[0192]
[0193] Among them, in step 1, a certain number of observation points are sampled from the solution space to form an initial set of observation points. The specific steps include: (1) Using methods such as random sampling / grid sampling to sample a set of data from the solution space, and the size of the initial observation set can be customized: (2) Input each point in the initial observation set into the black box function F to obtain its observation value: D n ←{(x i , F(x i ))}.
[0194] In step 2, by judging the dominance relationship, the Pareto set and the Pareto front in the initial observation point set are determined. For the min F problem, the judgment criterion for its Pareto set is: for any point + on the Pareto front P there must not exist a more excellent point y such that where o is the dimension of y.
[0195] In step 3, if the acquisition function is PHVI, no scalarization operation is required, and 0-1 utility values are directly assigned to each point according to the Pareto set; if it is other acquisition functions, such as EHVI, only the scalarization operation needs to be performed on the observation points in the Pareto set. The utility calculation of the observation points of PHVI and EHVI is shown in equations (23) and (24), where P + is the Pareto set obtained in step 2, and P - is the non-Pareto set.
[0196]
[0197] In step 4, there is no need to sort the utility values. The observation point set is classified. The observation points with non-0 utility values are classified as the positive class, and the observation points with 0 utility values are classified as the negative class.
[0198] In step 5, a classifier is used to fit the observation point set. Among them, according to the different acquisition functions adopted, the loss function shown in equation (20) is used for training. The posterior probability output by the classifier is used as an estimate of the density ratio, so as to realize the estimation of the acquisition function. DR.EMO is compatible with any advanced classifier. In subsequent experiments, the present invention uses XGBoost, random forest (RF), and multi-layer perceptron (MLP) as classifiers. After training, the posterior probability it provides can be used as an estimate of the acquisition function to evaluate candidate points.
[0199] In step 6, given the candidate solution set, the classifier can output its acquisition function value. On this basis, a point x t that maximizes the acquisition function is determined through optimization in the solution space and is used as the next observation point of the black-box function. DR.EMO supports a variety of optimization methods. For example, random sampling and covariance matrix adaptation evolution strategy (Covariance Matrix Adaptation Evolution Strategy, CMAES) can be used to optimize the candidate point space.
[0200] In step 7, the black-box function F is used to evaluate the best observation point x obtained in step 6 t, add it to the set of observation points; repeat steps 2-7, obtaining a new observation point in each round, and expanding the set of observation points through the acquisition function, thereby improving the fitting accuracy of the classifier and making the estimation of the acquisition function more accurate.
[0201] In step 8, based on the set of observation points, determine its Pareto set and Pareto front as the final multi-objective optimization result.
[0202] The following describes the process of the optimization performance evaluation experiment of the embodiments of the present invention:
[0203] DR.EMO can estimate a more effective acquisition function through a weighted classifier. Therefore, theoretically, it can provide better optimization performance. The purpose of this experiment is to demonstrate the superiority of DR.EMO in optimization performance through comparison with other optimization methods based on benchmark problems.
[0204] In the selection of test problems, the DTLZ problem is used as the benchmark problem, and by setting different input and output dimensions, the performance of DR.EMO in different-dimensional optimization problems is verified. Table 2 shows the combination of benchmark problems and optimization methods used in this experiment.
[0205] Table 2 Benchmark problems and optimization methods used in the experiment
[0206]
[0207] In the selection of baselines, this section selects advanced methods in two directions in the field of multi-objective Bayesian optimization methods; ① MBORE, which was proposed in 2022 and is the first method to apply density ratio estimation to multi-objective Bayesian optimization problems. It uses PHC (Pareto Hypervolumn Contribution) as a scalar, which can ensure that non-Pareto points also have non-zero scalar values, but it still uses a fixed threshold as the division basis and has the aforementioned problem of inaccurate optimization direction; ② MESMO, which was proposed in 2019 and is a multi-objective Bayesian optimization method based on maximum entropy search. It improves the problem of low computational efficiency of PESMO by evaluating the information gain in the output space and improves the evaluation efficiency. MBORE uses PHC as the utility function and also uses XGB / RF / MLP as the classifier. The number of Monte Carlo samplings of the Pareto front in MESMO is 3. The settings of other parameters refer to their respective papers.
[0208] In the specific experimental settings, in this section, DR.EMO is used to estimate PHVI and EHVI respectively, and classifiers such as Xgboost, random forest RF, and multi-layer perceptron MLP are used for estimation. The relative hypervolume is used to evaluate the optimization results to avoid the influence of problem characteristics on the results. In terms of hyperparameter settings, the parameter settings of the classifier in DR.EMO refer to MBORE, and CMAES is used as the acquisition function optimization method. The method combinations participating in the comparison are shown in Table 2.
[0209] For each benchmark problem with a fixed dimension, all methods share the same initial observation point set (n = 20), and each group of methods is tested twice. The average of the two test results is taken as the final result. The final results are as Figure 6 shown, where the horizontal axis represents the number of iterations and the vertical axis represents the relative hypervolume.
[0210] From the results, it can be obtained that: (1) DREMO-EHVI+XGB / RF is better than other methods in each dimension and has a faster convergence speed; at the same time, the standard deviation of the results of DREMO+EHVI in the two experiments is smaller, and its stability is stronger compared to other methods; (2) Due to the underfitting problem caused by the small amount of data, MLP is not suitable for application in density ratio estimation and performs significantly worse than other methods in some dimensions; (3) MESMO only performs well in the case of lower dimensions. In the experiments with o>2 and d>2, its effect lags far behind other methods. The reason for this may be that this method relies on the prior likelihood assumption of the Gaussian process and the Monte Carlo sampling results.
[0211] At the same time, to verify the computational efficiency of DR.EMO, in this section, the computational times of DREMO(PHVI), DREMO(EHVI), MBORE(PHC), and MESMO are respectively counted. The classifier for density ratio estimation is set to XGB, and the sampling function optimization method is to randomly sample and take the maximum value. The scalarization computational time and total time of each iteration are respectively counted, and the results are as Figure 7 shown in Table 3. The experiment is repeated twice, and the final result is the average of the two results.
[0212] Table 3 Average classification / total time in each iteration process of different Bayesian optimization algorithms
[0213]
[0214] The results show that: (1) DREMO takes the least time to estimate PHVI, and the higher the number of iterations / the higher the problem dimension, the more obvious the time advantage; (2) When DREMO estimates EHVI, it is not necessary to calculate the hypervolume contributions of all observation points, so its calculation time is better than that of MBORE. The amount of time saved depends on the proportion of the Pareto set in the observation point set at each iteration. Compared with MBORE using a fixed threshold, it is more efficient; (3) MESMO based on maximum entropy search has the highest calculation cost. At the same time, the original paper shows that its calculation time is proportional to the number of Monte Carlo samplings used. When the number of samplings is 10 or higher, the calculation time is longer.
[0215] Based on the above experiments, it can be considered that: DR.EMO can achieve higher optimization performance and faster convergence speed by estimating a more efficient acquisition function (EHVI); by adopting a dynamic threshold, it can have higher computational efficiency.
[0216] Another aspect of the embodiments of the present invention further provides a multi-objective Bayesian optimization device, including:
[0217] A first module, configured to classify an observation point set by using a dynamic threshold during the multi-objective Bayesian optimization process, regarding the points located on the Pareto front in the observation points as positive samples and the remaining points as negative samples, so that the positive class completely covers the Pareto set;
[0218] A second module, configured to extend the density ratio estimation to multi-objective Bayesian optimization, and train a classifier through weighted cross-entropy loss, so that the density ratio estimation method can be applicable to acquisition functions for any utility;
[0219] A third module, configured to complete the multi-objective Bayesian optimization process and output the Pareto set; and apply the multi-objective Bayesian optimization process to multi-object recognition in an image processing task to obtain a multi-object recognition result;
[0220] Wherein, the classifying the observation point set by using a dynamic threshold during the multi-objective Bayesian optimization process includes the following steps:
[0221] Sampling n initial observation points in the solution space to generate an initial observation point set;
[0222] Obtaining the Pareto set of each observation point in the initial observation point set;
[0223] Calculating the utility of each observation point;
[0224] Assigning a class label to each observation point, and the classification basis is whether it is in the Pareto set;
[0225] The extension of density ratio estimation to multi-objective Bayesian optimization and training of a classifier using weighted cross-entropy loss includes the following steps:
[0226] Using utility as the weight of the observation point to train the classifier;
[0227] Using the posterior probability provided by the classifier for optimization to determine the next observation point;
[0228] Evaluating the new observation point and adding it to the set of observation points.
[0229] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented in the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0230] An embodiment of the present invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above multi-objective Bayesian optimization method. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0231] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented in the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0232] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above multi-objective Bayesian optimization method.
[0233] It can be understood that the content in the above method embodiments is applicable to the storage medium embodiments. The functions specifically implemented in the storage medium embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0234] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.
[0235] It should be noted that in each specific embodiment of the present invention, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present invention need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present invention will be obtained.
[0236] The embodiments described in the embodiments of the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation to the technical solutions provided by the embodiments of the present invention. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.
[0237] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present invention, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0238] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0239] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0240] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0241] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0242] In addition, the functional units in various embodiments of the present invention may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0243] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.
[0244] The preferred embodiments of the embodiments of the present invention have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present invention. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of the rights of the embodiments of the present invention.
Claims
1. A multi-objective Bayesian optimization method, characterized in that: The following steps are involved: In the multi-objective Bayesian optimization process, a dynamic threshold is used to classify the set of observation points. The points on the Pareto front are regarded as positive samples, and the remaining points are regarded as negative samples, so that the positive class completely covers the Pareto set. The density ratio estimation is extended to multi-objective Bayesian optimization. The classifier is trained by weighted cross entropy loss, making the density ratio estimation method applicable to acquisition functions of arbitrary utility. Complete the multi-objective Bayesian optimization process and output the Pareto set; apply the multi-objective Bayesian optimization process to the multi-objective recognition of image processing tasks and obtain the multi-objective recognition results; Wherein, in the multi-objective Bayesian optimization process, the dynamic threshold is used to classify the observation point set, including the following steps: Sample n initial observation points in the solution space to generate an initial observation point set; Obtain the Pareto set of each observation point in the initial observation point set; Calculate the utility of each observation point; Assign a category label to each observation point, based on whether it is in the Pareto set; The density ratio estimation is extended to multi-objective Bayesian optimization, and the classifier is trained by weighted cross entropy loss, including the following steps: Use utility as observation point weight to train the classifier; Use the posterior probability provided by the classifier to optimize and determine the next observation point; Evaluate new observations and add them to the collection of observations.
2. A multi-objective Bayesian optimization method according to claim 1, characterized in that: The step of sampling n initial observation points in the solution space to generate an initial observation point set includes the following steps: Use random sampling or grid sampling to sample a set of data from the solution space The initial observation set size is defined as: Input each point in the initial observation set into the black box function F to obtain its observation value D n :D n ←{(x i ,F(x i ))}.
3. A multi-objective Bayesian optimization method according to claim 1, characterized in that: The method of obtaining the Pareto set of each observation point in the initial observation point set is specifically as follows: determining the Pareto set and the Pareto front in the initial observation point set by judging the dominance relationship; wherein, for the min F problem, the judgment standard of the Pareto set is: for the Pareto front P + Any point on There must not be a better point y such that Where o is the dimension of y.
4. A multi-objective Bayesian optimization method according to claim 1, characterized in that: The calculation of the utility of each observation point includes the following steps: If the acquisition function is PHVI, there is no need to perform scalarization operations, and a 0-1 utility value is directly assigned to each point according to the Pareto set; If it is other acquisition functions, including EHVI, the scalarization operation is performed on the observation points in the Pareto set; Among them, the observation point utility u of PHVI is i And the observation point effect of EHVI i The calculation is shown in the following equation: Among them, P + is the obtained Pareto set, P - is a non-Pareto set; HV(·) represents a hypervolume constructed by a point set and a reference point.
5. A multi-objective Bayesian optimization method according to claim 1, characterized in that: In the step of assigning a category label to each observation point, there is no need to sort the utility values, and the observation point set is classified, and the observation points with non-zero utility values are classified into positive classes, and the observation points with zero utility values are classified into negative classes; The method of training a classifier using utility as an observation point weight comprises the following steps: Use the classifier to fit the set of observation points, which differs according to the acquisition function adopted; According to the following loss function, the posterior probability output by the classifier is used as the estimated value of the density ratio, thereby realizing the estimation of the acquisition function; loss function The expression is: Among them, θ is the classifier parameter; l(x) and They represent the conditional probability density of x with respect to y under two values of y; U(S) represents the utility function that is independent of x; β represents the proportion of positive samples; K represents a constant; XGBoost, random forest, and multi-layer perceptron are used as classifiers. After training, the posterior probability obtained from the training is used as the estimated value of the acquisition function to evaluate the candidate points.
6. A multi-objective Bayesian optimization method according to claim 1, characterized in that: The method of using the posterior probability provided by the classifier to perform optimization and determine the next observation point includes the following steps: Given a set of candidate solutions, output the acquisition function value through the classifier, and determine a point x that maximizes the acquisition function by optimizing the solution space. t , taking it as the next observation point of the black box function; The step of evaluating the new observation point and adding it to the observation point set includes the following steps: Use the black box function F to evaluate the best observation point x t , add it to the observation point set, and each time a new observation point is obtained, the observation point set is expanded by obtaining the function.
7. A multi-objective Bayesian optimization method according to claim 1, characterized in that: In the process of classifying the observation point set using dynamic threshold, it is not necessary to perform scalar processing on each observation point. It is only necessary to determine the dominance of each observation point. The dominance relationship of the observation point is obtained by comparing the components y i The size of i∈(1,2,…,q) is realized; In the process of training the classifier with weighted cross entropy loss, before the density ratio estimation, the output of the black box function F is first quantized using a scalar S:y→R, i.e. Get the set of observations U(s) represents the utility function that is independent of x, p(s|x,D N ) represents the prior distribution provided by the surrogate model, P + and P - Represent the Pareto set and non-Pareto set of the function output respectively, and obtain the function α(x; D N ) represents the expectation of each input point x about the utility function, where the utility function represents the degree of improvement of the hypervolume by adding the candidate point; the density ratio equation containing the coefficient w is modeled by maximizing the log-likelihood of the weighted cross entropy loss.
8. A multi-objective Bayesian optimization device, characterized in that: include: The first module is used to classify the set of observation points using a dynamic threshold in the multi-objective Bayesian optimization process, and the points on the Pareto front among the observation points are regarded as positive samples, and the remaining points are regarded as negative samples, so that the positive class completely covers the Pareto set; The second module is used to extend the density ratio estimation to multi-objective Bayesian optimization. The classifier is trained by weighted cross entropy loss, so that the density ratio estimation method can be applied to the acquisition function of any utility. The third module is used to complete the multi-objective Bayesian optimization process and output the Pareto set; The multi-objective Bayesian optimization process is applied to the multi-objective recognition of image processing tasks and the multi-objective recognition results are obtained; Wherein, in the multi-objective Bayesian optimization process, the dynamic threshold is used to classify the observation point set, including the following steps: Sample n initial observation points in the solution space to generate an initial observation point set; Obtain the Pareto set of each observation point in the initial observation point set; Calculate the utility of each observation point; Assign a category label to each observation point, based on whether it is in the Pareto set; The density ratio estimation is extended to multi-objective Bayesian optimization, and the classifier is trained by weighted cross entropy loss, including the following steps: Use utility as observation point weight to train the classifier; Use the posterior probability provided by the classifier to optimize and determine the next observation point; Evaluate new observations and add them to the collection of observations.
9. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Analog circuit optimization algorithm based on multi-objective acquisition function integrated parallel Bayesian optimization
CN110750948A
Cable state evaluation method based on Bayesian optimization and XGBoost
CN119293625A
Optimal Strategies in Security Games
US20130273514A1
KR20240030173A