Soft and hard constraint machine learning method and system based on abstract gradient descent
By adopting an abstract gradient descent method in machine learning, the soft and hard constraints are optimized alternately, which solves the problem of difficulty in combining soft and hard constraints in the existing technology, and improves the logical constraint satisfaction and data fitting effect of the learning model.
Patent Information
- Application Number
- CN202411940369.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The existing technology is difficult to effectively combine soft and hard constraints, resulting in the inability to meet data fit and logical constraints at the same time in machine learning tasks.
The method based on abstract gradient descent is adopted to optimize the soft and hard constraints against alternating optimum, and the unified optimization of soft and hard constraints is achieved by constructing the loss function and logical properties.
It realizes the optimization of soft constraints within the feasible range of meeting hard constraints, ensuring the logical constraint satisfaction and data fitting effect of the learning model.
Smart Images

Figure CN119990237A_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the technical field of machine learning, and in particular to a soft and hard constraint machine learning method and system based on abstract gradient descent. Background Art
[0002] In the machine learning task of data fitting, the learning model needs to be trained to fit the data in the training set. However, in some tasks, the learning model needs to not only fit the data, but also meet some constraints. For example, in the field of autonomous driving, the learning model not only needs to learn the mapping from visual images to vehicle operations so that the vehicle can reach the destination smoothly and quickly, but also needs to comply with relevant traffic regulations to avoid traffic accidents. The former is mainly enabled by training data, while the latter needs to meet artificially defined logical rules. The former is called the "soft constraint" that the learning model needs to meet, and the latter is called the "hard constraint" that needs to be met. Soft constraints refer to constraints that do not need to be strictly met, such as the fitting accuracy of the training set does not need to reach 100%; hard constraints refer to constraints that need to be strictly met. Current research mainly focuses on the learning and training of soft constraints, but there is no mature technical means for hard constraints.
[0003] At present, there are three main implementation ideas for soft and hard constrained machine learning in the existing technology.
[0004] The first is to train first and then verify: that is, without considering constraints first. Directly fit the target ∑ to the training data <x,y>∈D Loss(F α (x), y) to optimize, and then verify whether it can pass after the optimization is completed If not, find a counterexample and put it into the training set D, and then train again. However, the method of training first and then verifying has two serious flaws. First, the training process cannot guarantee 100% fit to the soft constraint D, which makes the newly added counterexamples unable to guarantee the effectiveness of the next training for the hard constraint. Second, even if each training can reach the global optimum, due to the limitation of the counterexample as a single sample, it is difficult to guarantee that the learning model can eventually pass the verification of the logical constraint.
[0005] The second is to verify first and then train: that is, do not consider ∑ <x,y1∈D Loss(F α (x),y), we first find About the satisfyable range of α. Then use gradient descent to optimize within this range∑ <x,y>∈D Loss(F α(x), y). The drawback of this method is that the range of α that can be satisfied may be relatively narrow, which limits the subsequent gradient descent optimization space and ultimately makes it impossible to effectively satisfy the soft constraints.
[0006] The third is a simultaneous approach: it mainly relies on the neural relaxation method to constrain the logic Relaxation is an almost everywhere differentiable function and added to the loss, which becomes an optimization goal. In the relaxation process, the upper bound of the negation of the original constraint can be approximated by relaxation, so that when the final loss is 0, the constraint can be guaranteed to be satisfied. There are two defects in this method: first, the gradient descent may not necessarily make the value of the upper bound approximate differentiable relaxation function drop to 0, so the constraint cannot be guaranteed to be satisfied; second, the approximate upper bound function may not be 0 due to errors, so it is impossible to judge whether the constraint is satisfied.
[0007] From the above, we can see that the method of training first and then verifying cannot guarantee that the learning model can eventually pass the verification of logical constraints after the algorithm terminates; verifying first and then training limits the subsequent gradient descent optimization space, resulting in the inability to effectively satisfy the soft constraints; the neural relaxation gradient descent technology that performs training and verification at the same time cannot guarantee and is difficult to judge the satisfaction of constraints. Summary of the invention
[0008] The technical problem to be solved by the present invention is that, in view of the deficiencies in the prior art, the present invention provides a soft and hard constraint machine learning method and system based on abstract gradient descent, which has a simple principle, is easy to implement, is convenient to operate, and is highly efficient.
[0009] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0010] A soft and hard constraint machine learning method based on abstract gradient descent, comprising:
[0011] Step S1: Establish a learning model;
[0012] Step S2: Construct soft constraints from the training set and construct ∑ <x,y>∈D Loss(F α (x), y) loss function;
[0013] Step S3: Construct hard constraints from logical properties. Construct min b F P (b,α);
[0014] Step S4: Performing alternating optimization between soft and hard constraint optimization targets;
[0015] Step S5: Output the trained model.
[0016] As a further improvement of the present invention: in step S4, the soft constraint optimization target min α ∑ <x,y>∈D Loss(F α (x), y), hard constraint optimization target min b F P (b, α), and optimize the two alternately according to the following process:
[0017] Step S41: Assume that the current value of α is α i , i = 0, let the feasible domain of α be its definition domain;
[0018] Step S42: Calculate and abstract downward Where L(α) = ∑ <x,y>∈D Loss(F α (x), y), we get α i+1 The target set: Any value α in the target set i+1 All of these can make L(α i+1 ) is less than L(α i ) value;
[0019] Step S43: Set Optimize the feasible domain of α in the hard constraint objective and calculate the upward abstraction
[0020] As a further improvement of the present invention: the step S43 includes: if the upper abstraction is an empty set and F P When (b,α)=1, it proves that F P The feasible domain of (b,α) in α is When F P The value of (b,α) will never be less than 1, so it is proved that At this time, continue in the feasible region of α Continue to optimize the soft constraint objective of α.
[0021] As a further improvement of the present invention: the step S43 includes: if the upper abstraction is not an empty set, it means that P(b, α) may be false, and then calculate the lower abstraction If the lower abstraction is empty, refine it until it is proved that the upper abstraction is empty or the lower abstraction is not empty.
[0022] As a further improvement of the present invention: the step S43 comprises: if the following abstract is not an empty set, then we get a set that makes F P The set of α where (b,α)=0 is denoted by Another cumulative set S -is the “infeasible” region of α.
[0023] As a further improvement of the present invention: observe the fitting satisfaction of soft and hard constraints, and analyze S - and S + To sort out and analyze whether there are any contradictions between the soft and hard constraints, and make further modifications to the soft and hard constraints.
[0024] As a further improvement of the present invention: a threshold is set, and when the number of backtracking times is greater than the threshold, some hard constraints are abandoned to terminate the process; after the strategy is terminated, a learning model that satisfies the hard constraints and fits the soft constraints as much as possible is obtained.
[0025] The present invention further provides a soft and hard constraint machine learning system based on abstract gradient descent, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0026] Compared with the prior art, the advantages of the present invention are:
[0027] The soft and hard constraint machine learning method and system based on abstract gradient descent of the present invention have simple principles, simple implementation and high efficiency. The soft and hard constraint adversarial learning technology based on abstract gradient proposed by the present invention can determine that the hard constraint is satisfied when the optimization direction of the hard constraint optimization target is abstracted as an empty set and the value of the optimization target is 1, and the soft constraint is optimized within the feasible range that satisfies the hard constraint; and due to the properties of the abstract gradient, the feasible range that satisfies the hard constraint is always the upper bound of the original feasible range, so it can be ensured that the original optimization space will not be unnecessarily restricted due to satisfying the hard constraint (for example, the method of verifying first and then training, in which the verification area is the lower bound of the original feasible range).
[0028] The present invention uses the optimization capability of abstract gradient for discrete functions to achieve the optimization of soft and hard constraints. Based on abstract gradient descent, the present invention designs an intuitive descent strategy for soft and hard constraints. Since the abstract gradient does not require continuous differentiability, the present invention can directly use the logical semantic discrete function of the constraint as the optimization target. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic diagram of the flow chart of the present invention in a specific embodiment. DETAILED DESCRIPTION
[0030] The present invention is further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.
[0031] The problem addressed by the present invention can be further described in a specific formal way as follows:
[0032] 1. A training set D, where each piece of data is recorded as an input-output pair<x,y> ;
[0033] 2. A learning model with input x and output y, denoted by y = F α (x). Where α is the parameter to be learned.
[0034] 3. A logical constraint C, C can be regarded as a predicate, input a learning model, when the learning model satisfies the constraint, the output Boolean value "true", when the learning model does not satisfy the constraint, the output Boolean value "false". α When is a parameterized function with a fixed structure, the input of the logical constraint can be replaced by the parameter α, i.e. C(α).
[0035] More specifically, the present invention considers that when the logical constraint C is a first-order logic formula, the form Where P is a predicate containing only free variables α, and b is the constraint variable of the universal quantifier. This logical constraint can express many property descriptions, such as "when there are pedestrians in front of the vehicle in the driving direction, the accelerator cannot be stepped on under any circumstances to accelerate".
[0036] To this end, the soft and hard constraint learning problem can be defined as the following constrained optimization problem:
[0037]
[0038] The Loss function is a measure of F α (x) is a function of the difference between (x) and y. When y is a real vector, the vector distance can be used as the difference measure function. When F α When (x) and y are equal, Loss(F α (x), y) is zero.
[0039] The problem to be solved by the present invention is to find the value of α given the above constrained optimization objective so that is true, and ∑ <x,y>∈D Loss(F α The smaller the value of (x), y).
[0040] The present invention mainly utilizes the optimization capability of abstract gradient for discrete functions to realize the optimization of soft and hard constraints. Based on abstract gradient descent, the present invention can design an intuitive descent strategy for soft and hard constraints. Since the abstract gradient does not require continuous differentiability, the logical semantic discrete function of the constraint can be directly used as the optimization target. For example Can be converted to a discrete function min b F P (b,α), where F Pis a semantic discrete function of the predicate P. min and max can just correspond to the semantics of arbitrary and existential quantifiers. Since the goal of the above learning is to find the value of α so that the constraint True. Therefore, it can be converted to the optimization objective max α min b F P (b,α), and as long as the value of the optimization problem is 1, the original constraint can be guaranteed to be satisfied. In addition, the optimization goal of the soft constraint is min α ∑ <x,y>∈D Loss(F α (x), y). Therefore, the present invention adjusts the optimization target and combines the two as follows:
[0041]
[0042] Using abstract gradient descent, we can calculate the abstract gradients for the values of α and b respectively.
[0043] It is worth noting that the optimization direction of b is the complement of the abstract gradient (the complement of the upper abstraction produces effective optimization, and the complement of the lower abstraction is empty and can be used to determine the global optimum).
[0044] When the optimization directions of the three parameters are all empty and F P When the value of (b,α) is 1, the algorithm successfully finds an α that satisfies the hard constraints and sufficiently fits the soft constraints.
[0045] There is still a problem in the above process, that is, the non-convergence problem. Since the optimization directions of parameter b and parameter α for the above formula are opposite, the intersection of the lower abstract range of the abstract gradient descent of the two must be empty, which leads to the only way to inspire exploration through upper abstraction, or to optimize alternately. Because the process is too random, it is impossible to track and analyze the convergence, and it is difficult to weigh and adjust the optimization proportion of soft and hard constraints. Therefore, the present invention proposes a soft and hard constraint adversarial learning, which is used to train a learning model that satisfies hard constraints and fits soft constraints.
[0046] like Figure 1 As shown, the soft and hard constraint machine learning method based on abstract gradient descent of the present invention includes the following steps:
[0047] Step S1: Establish a learning model;
[0048] This embodiment is called a symbolic neural network.
[0049] The symbolic neural network may be a traditional neural network, or may be an arbitrary function formed by the participation of discrete functions, such as a semantic discrete function containing logical connectives.
[0050] The learning model can be any function with parameters, denoted by F.
[0051] Step S2: Construct soft constraints from the training set and construct ∑ <x,y>∈D Loss(F α (x), y) loss function;
[0052] Step S3: Construct hard constraints from logical properties. Construct min b F P (b,α);
[0053] This step only requires writing the corresponding discrete function according to the constraint semantics;
[0054] For example, less than or equal to can be written as F ≤ (x,y), its semantics is to output 1 when the value of x is less than or equal to y, and output 0 in other cases.
[0055] Step S4: Performing alternating optimization between soft and hard constraint optimization targets;
[0056] Step S5: Output the trained model.
[0057] In addition to outputting the learning model as a black box function, since F is no longer required to be continuously differentiable, if the function is established in the form of logical symbols, it can be output in the form of a symbolic expression.
[0058] In a specific application example, in step S4, the soft constraint optimization target min α ∑ <x,y>∈D Loss(F α (x), y), hard constraint optimization target min b F P (b, α), the two can be optimized alternately according to the following process:
[0059] Step S41: Assume that the current value of α is α i , i = 0, let the feasible domain of α be its definition domain;
[0060] Step S42: Calculate and abstract downward Where L(α) = ∑ <x,y>∈D Loss(F α (x), y), we get α i+1 The target set: Any value α in the target set i+1 All of these can make L(α i+1 ) is less than L(α i ) value;
[0061] Step S43: Set Optimize the feasible domain of α in the hard constraint objective and calculate the upward abstraction
[0062] As a preferred embodiment, the step S43 may further include: if the upper abstraction is an empty set and F P When (b,α)=1, it proves that F P The feasible domain of (b,α) in α is When F P The value of (b,α) will never be less than 1, so it is proved that At this point, we only need to continue in the feasible region of α Continue to optimize the soft constraint objective of α.
[0063] As a preferred embodiment, the step S43 may further include: if the upper abstraction is not an empty set, it means that P(b, α) may be false, and the lower abstraction needs to be calculated. If the lower abstraction is empty, refine it until it is proved that the upper abstraction is empty or the lower abstraction is not empty.
[0064] As a preferred embodiment, the step S43 may further include: if the following abstract is not an empty set, then we get a set that makes F P The set of α where (b,α)=0 is denoted by Another cumulative set S - is the “infeasible” region of α. This means that any - The value of α in will make is false. Therefore, the value of α cannot be in S - At this time, the feasible domain of soft constraint optimization is defined as (S - ) c , that is, S - The complement of , and then optimize the soft constraints (i.e., execute step S43). - ) c When is an empty set, it proves that the hard constraint is unsatisfiable.
[0065] In this embodiment, the present invention can further controllably and intuitively observe the fitting satisfaction of soft and hard constraints, and by analyzing S - and S + To sort out and analyze whether there are any contradictions in the soft and hard constraints, so as to further modify the soft and hard constraints.
[0066] For example, a threshold can be set, and when the number of backtracking times is greater than the threshold, some hard constraints are abandoned to terminate the process. On the other hand, after the strategy is terminated, a learning model that satisfies the hard constraints and fits the soft constraints as much as possible can be obtained.
[0067] The abstract gradient technology mentioned above in the present invention can unify the modeling and optimization of the satisfaction of logical constraints and the fitting of training data. That is, for any function F(v) (F(v) may be a discrete function), its abstract gradient function is recorded as
[0068]
[0069] Right now, is a set that contains all the change directions Δv about v that can make the value of F(v-Δv) smaller than the current value.
[0070] For example, for the "less than or equal to" discrete function f ≤ For example, When x = 2 When calculated After that, we only need to sample any change direction and use v:=v-Δv to optimize the function value. When the calculated set is empty, it can be proved that the global optimum has been reached.
[0071] The calculation of abstract gradient descent is generally implemented by backward abstract propagation. Backward abstract interpretation refers to constructing a mapping function from its image to the original image for a function F. (called backward abstraction function), which takes as input a set Y of range and outputs a set of domains Satisfy the following formula:
[0072]
[0073] That is, input any image of the function F and output its corresponding original image.
[0074] The abstract gradient set can be composed of Equivalent calculation.
[0075] The construction of backward abstraction functions has the property of the "chain rule of abstraction":
[0076]
[0077] That is, the backward abstraction function of a composite function can be composed of the composite of the backward abstraction functions of each function. According to the "Abstract Chain Rule", the backward abstraction function of any function can be recursively constructed according to its composite structure. Based on this, any function F can be expressed as a composite of several "layer functions". A layer function refers to a function whose "projection function" in each dimension is an atomic function. The projection function g of a function g in the i-th dimension is i is a function that satisfies the following formula:
[0078]
[0079] Based on this, any function F is always expressed as a composite of n layers (n is greater than or equal to 1), that is:
[0080]
[0081] in this way, The calculation can be converted into each layer function Calculation of the layer function The calculation of can be converted into the backward abstract calculation of the projection function (atomic function) of its various dimensions.
[0082] At the same time, in order to allow the calculation process to terminate, the backward abstract function is not calculated, but the backward "down" abstract function is calculated And the backward "up" abstract function They are:
[0083]
[0084] Backward Abstraction Function The output is the backward abstraction function Subset of; backward abstraction function The output is the backward abstraction function A superset of .
[0085] when When it is not empty, abstract gradient descent can be performed; when When it is empty, it can be proved that it reaches the global optimum. The "Abstract Chain Rule" also applies to backward and downward abstraction.
[0086] The present invention further provides a soft and hard constraint machine learning system based on abstract gradient descent, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0087] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A soft and hard constraint machine learning method based on abstract gradient descent, characterized in that: include: Step S1: Establish a learning model; Step S2: Construct soft constraints from the training set and construct ∑ <x,y>∈D Loss(F α (x), y) loss function; Step S3: Construct hard constraints from logical properties. Construct min b F P (b, α); Step S4: Performing alternating optimization between soft and hard constraint optimization targets; Step S5: Output the trained model.
2. The soft and hard constraint machine learning method based on abstract gradient descent according to claim 1, characterized in that: In step S4, the soft constraint optimization target min α ∑ <x,y>∈D Loss(F α (x), y), hard constraint optimization target min b F P (b, α), and optimize the two alternately according to the following process: Step S41: Assume that the current value of α is α i , i = 0, let the feasible domain of α be its definition domain; Step S42: Calculate and abstract downward Where L(α) = ∑ <x,y>∈D Loss(F α (x), y), we get α i+1 The target set: Any value α in the target set i+1 All of these can make L(α i+1 ) is less than L(α i ) value; Step S43: Set Optimize the feasible domain of α in the hard constraint objective and calculate the upward abstraction 3. The soft and hard constraint machine learning method based on abstract gradient descent according to claim 2, characterized in that: The step S43 includes: if the upper abstraction is an empty set and F P When (b, α) = 1, it proves that F P The feasible domain of (b, α) in α is When F P The value of (b, α) will never be less than 1, so it is proved that At this time, continue in the feasible region of α Continue to optimize α with soft constraints.
4. The soft and hard constraint machine learning method based on abstract gradient descent according to claim 2, characterized in that: The step S43 includes: if the upper abstraction is not an empty set, it means that P(b, α) may be false, and then calculate the lower abstraction If the lower abstraction is empty, refine it until it is proved that the upper abstraction is empty or the lower abstraction is not empty.
5. The soft and hard constraint machine learning method based on abstract gradient descent according to claim 2, characterized in that: The step S43 includes: if the following abstract is not an empty set, then we get a set that makes F P The set of α where (b, α) = 0 is denoted by Another cumulative set S - is the "infeasible" region of α.
6. The soft and hard constraint machine learning method based on abstract gradient descent according to any one of claims 1 to 5, characterized in that: Observe the fitting satisfaction of soft and hard constraints and analyze S - and S + To sort out and analyze whether there are any contradictions between the soft and hard constraints, and make further modifications to the soft and hard constraints.
7. The soft and hard constraint machine learning method based on abstract gradient descent according to claim 6, characterized in that: A threshold is set. When the number of backtracking times is greater than the threshold, some hard constraints are abandoned to terminate the process. After the strategy is terminated, a learning model that satisfies the hard constraints and fits the soft constraints as much as possible is obtained.
8. A soft and hard constraint machine learning system based on abstract gradient descent, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for constructing abstract model for target program
CN118331846A
Systems and Methods for Multi-Objective Evolutionary Algorithms with Soft Constraints
US20170169353A1