Multi-objective optimization partially observable markov decision process system utilizing hierarchical adaptive probabilistic information manifestation mechanism
The hierarchical adaptive mechanism in POMDPs addresses inefficiencies by optimizing multiple objectives, adapting to environmental changes, and handling large-scale and continuous spaces, enhancing decision-making in autonomous systems and other applications.
Patent Information
- Application Number
- JP2025079078
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-10
- Publication Date
- 2025-07-25
AI Technical Summary
Conventional Partially Observable Markov Decision Processes (POMDPs) face challenges in handling multiple competing objective functions, adapting to environmental changes, managing large-scale state spaces, and dealing with continuous state and action spaces, leading to inefficiencies and computational difficulties.
A hierarchical adaptive mechanism is introduced, incorporating multi-objective optimization, adaptive manifestation, hierarchical abstraction, and continuous space extension to optimize multiple goals dynamically, reduce computational complexity, and handle continuous spaces.
The mechanism enables efficient, adaptive decision-making that optimizes multiple competing objectives, handles large-scale and continuous state spaces, and maintains theoretical guarantees, applicable to autonomous systems, autonomous driving, medical diagnosis, and smart grid control.
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, particularly decision-making under uncertainty, and more specifically to a system that extends a probabilistic information revelation mechanism to overcome the incompleteness of information in Partially Observable Markov Decision Processes (hereinafter referred to as "POMDPs") and realizes multi-objective optimization. The present invention is useful in various application fields where decision-making under uncertainty is required, such as autonomous robots, autonomous driving vehicles, medical diagnosis systems, financial engineering, and smart grid control. In particular, the present invention exhibits remarkable effects in situations where it is required to simultaneously optimize multiple competing goals, adaptively respond to environmental changes, and efficiently operate even for large-scale state spaces.
Background Art
[0002] Sequential decision-making under uncertainty is an important issue in the fields of autonomous systems, robotics, and artificial intelligence. Partially Observable Markov Decision Processes (POMDPs) are widely used as a mathematical framework for modeling situations where an agent cannot observe the complete state of the environment. POMDPs were defined in the 1960s for operations research and introduced into artificial intelligence by the seminal papers of Kaelbling, Littman, and Cassandra.
[0003] The main problem in conventional POMDPs is the incompleteness of information resulting from the fact that an agent cannot directly observe the true state of the environment. Due to this incompleteness, the agent needs to maintain a "belief" (probability distribution) about the current state. This belief space is continuous, and it is computationally difficult to find an optimal strategy. In fact, for most objective functions, it has been proven that the problem of determining the existence of an optimal strategy in POMDPs is undecidable.
[0004] The research by Belly et al. (Belly et al., 2025) introduced an "explicitization mechanism" to limit information loss in POMDPs. This mechanism guarantees with probability 1 that the agent will ultimately have complete information about the current state. They defined two properties, "weak explicitization" and "strong explicitization", and provided a decidable algorithm using a finite abstraction called belief support MDP for POMDPs that satisfy these properties.
[0005] The weak explicitization property is the property that, under any strategy, guarantees with probability 1 that the agent will make observations that uniquely identify the current state an infinite number of times. The strong explicitization property is the property that guarantees that any state transition in the underlying MDP may occur with a signal that uniquely identifies the state.
[0006] Belly et al. proved that the analysis of the belief support MDP is sound and complete for parity objective functions in the range of priorities 0, 1, 2 in weakly explicitized POMDPs and for any parity objective function in strongly explicitized POMDPs. This provided a decidable algorithm for these classes of POMDPs.
[0007] However, the research by Belly et al. has the following limitations. 1. Their approach only considers a single objective function (parity objective function) and cannot optimize multiple competing objective functions required in many real-world applications simultaneously. For example, an autonomous vehicle needs to optimize multiple goals such as safety, efficiency, and comfort simultaneously. 2. Their explicitization mechanism is static and cannot adapt to changes in the environment. In many real-world applications, the level of environmental uncertainty changes over time, so the information collection strategy also needs to be dynamically adjusted. 3. Their algorithm depends on the construction of belief-support MDPs, and its computational efficiency decreases when the state space is large. Since the size of belief-support MDPs increases exponentially with respect to the number of states of the original POMDP, it becomes computationally difficult for large-scale problems. 4. Their approach only considers discrete state and action spaces and cannot be directly applied to POMDPs with continuous state or action spaces. In many real-world applications, states and actions are often represented by continuous variables.
[0008] In the prior art, several approaches have been proposed to address the problems of information collection and multi-objective optimization in POMDPs. For example, interpolation in the belief space (Lovejoy 1991), approximation of value functions (Hauskrecht 2000), Monte Carlo tree search approach (Silver and Veness 2010), etc. However, these methods are approximate solutions and do not provide theoretical guarantees.
[0009] Also, subclasses of specific POMDPs for limiting information loss have been studied. For example, Berwanger and Mathew (2017) define a class of partial-information multi-player games in which "manifestation" is guaranteed to occur within a finite number of steps. However, these approaches are applicable only under specific constraints and not to general POMDPs.
[0010] In the field of multi-objective optimization, approaches such as Pareto optimization and weighted sum method for simultaneously optimizing multiple objective functions have been proposed. However, research on applying these methods in the context of POMDPs is still limited.
Prior Art Documents
Non-Patent Documents
[0011]
Non-Patent Document 1
[0012] The present invention aims to expand the concept of the revelation mechanism introduced by Belly et al. and overcome the above limitations. Specifically, it aims to solve the following problems. 1. To provide a framework for simultaneously optimizing multiple competing objective functions. In many real-world applications, it is necessary to simultaneously optimize multiple goals such as safety, efficiency, and comfort. 2. Develop a dynamic manifestation mechanism that can adapt to environmental changes. Since the level of environmental uncertainty changes over time, the information collection strategy also needs to be adjusted dynamically. 3. Design an algorithm that can operate efficiently for POMDPs with large-scale state spaces. In many real-world applications, the state space can be very large, and computational efficiency becomes important. 4. Develop a method that can handle POMDPs with continuous state spaces and action spaces. In many real-world applications, states and actions are often represented by continuous variables. 5. Strike a balance between theoretical guarantees and practicality. Although theoretically accurate algorithms are important, computational efficiency and ease of implementation also need to be considered in actual applications.
Means for Solving the Problem
[0013] The present invention provides a "hierarchical adaptive manifestation mechanism" having the following main technical features. 1. **Multi-objective Optimization Framework** The present invention provides a framework for simultaneously considering multiple objective functions (such as parity, reachability, safety, etc.) and deriving Pareto-optimal strategies. A multi-objective POMDP is defined as a tuple P = (Q, Act, Sig, δ, q0, O). Here, Q is a finite state set, Act is a finite action set, Sig is a finite signal set, δ: Q×Act → D(Sig×Q) is a transition function (D(X) represents the set of probability distributions of X), q0 ∈ Q is the initial state, and O = {O1, O2, ..., Ok} is the set of objective functions. For each objective function Oi, a weight wi representing its importance can be set, and a strategy that maximizes the weighted sum Σi wi·Oi can be obtained. Also, the Pareto optimization approach can be used to explicitly consider the trade-offs between objective functions. The value of strategy σ in a multi-objective POMDP is defined as the vector v(σ) = (PPσ[O1], PPσ[O2], ..., PPσ[Ok]). Here, PPσ[Oi] is the probability of satisfying Oi under strategy σ. A strategy σ is Pareto optimal if there is no other strategy that dominates v(σ). That is, there is no strategy σ' such that PPσ'[Oi] ≧ PPσ[Oi] for all i and PPσ'[Oj] ≦ PPσ[Oj] for at least one j. 2. **Adaptive Manifestation Mechanism** The present invention introduces a mechanism for dynamically adjusting the frequency and method of manifestation according to changes in the environment. The adaptive manifestation mechanism is defined as a function E: D(Q) × 2Q^ω → Act that selects a "search action" to promote manifestation based on the current belief state b and the objective function O. Specifically, based on the level of uncertainty H(b) (e.g., the entropy of the belief) in the current belief state b and the current degree of achievement V(b, O) for the objective function O, a balance is struck between search actions and objective achievement actions. When uncertainty is high, search actions are prioritized, and when uncertainty is low, objective achievement actions are prioritized. The adaptive manifestation mechanism selects actions based on the following rules. - When H(b) > θ (θ is a threshold), select the search action E(b, O). - When H(b) ≦ θ, select the optimal action a* = argmax_a V(b, O, a) for objective achievement. Here, V(b, O, a) is the expected value of the objective function O after executing action a in the belief state b. 3. **Hierarchical Abstraction Algorithm** The present invention improves computational efficiency by hierarchically abstracting the state space in the construction of a belief support MDP. The hierarchical abstraction algorithm divides the set of states Q into clusters of semantically similar states and constructs a "hierarchical belief support MDP" that considers transitions between those clusters. The hierarchical belief support MDP is defined as PH = (2C, Act, δH, {C(q0)}). Here, C is the set of clusters of states, C(q) is the cluster containing q, and δH is the transition function at the cluster level. This hierarchical abstraction can reduce the size of the belief support MDP, which increases exponentially with respect to the number of states of the original POMDP, and improve computational efficiency. Clustering is performed based on the "similarity" between states. The similarity between states q, q' ∈ Q is defined based on the following elements. - Similarity of transitions: The similarity of transition probabilities for the same action - Similarity of signals: The similarity of signals observed for the same action - Similarity of objective functions: The similarity of the contribution degrees to the objective function 4. **Continuous Space Extension** The present invention introduces a "continuous manifestation mechanism" using differentiable approximation functions to handle POMDPs with continuous state spaces and action spaces. For a POMDP with a continuous state space X ⊆ Rn and an action space A ⊆ Rm, the continuous manifestation mechanism is defined as a function R: X × A → [0, 1] that calculates the probability that a state x ∈ X is manifested. This function returns the probability R(x, a) that the state x is manifested for a state x and an action a. This probability is calculated based on the observability of the state x and the information-gathering ability of the action a. The continuous manifestation mechanism enables the application of the concept of manifestation to POMDPs with continuous state spaces and action spaces. The continuous manifestation mechanism is designed to satisfy the following properties. - For any state x ∈ X and action a ∈ A, 0 ≤ R(x, a) ≤ 1 - For any strategy σ, an infinite number of manifestations occur with probability 1 - R(x, a) is continuous with respect to x and a 5. **Integrated Algorithm** The present invention provides a "hierarchical adaptive multi-objective manifestation algorithm" that integrates the above four main technical features. This algorithm enables more efficient and effective decision-making by combining multi-objective optimization, adaptive manifestation, hierarchical abstraction, and continuous space expansion. Specifically, the following steps are executed. (1) Define a multi-objective POMDP and set weights for each objective function. (2) Hierarchically cluster the state space and construct a hierarchical belief support MDP. (3) Calculate a strategy that optimizes the weighted objective function for the hierarchical belief support MDP. (4) At runtime, dynamically adjust the balance between exploration actions and goal achievement actions based on the uncertainty of the current belief state. (5) For POMDPs with continuous state spaces and action spaces, apply a continuous manifestation mechanism.
Advantages of the Invention
[0014] The main advantages of the present invention are as follows. 1. **Realization of Multi-Objective Optimization** The present invention can simultaneously optimize multiple competing objective functions and can handle more realistic problem settings. For example, in the control of autonomous vehicles, multiple goals such as safety, efficiency, and comfort can be optimized simultaneously. This can provide a more practical solution compared to conventional approaches that consider only a single objective function. Specifically, by using the multi-objective optimization framework of the present invention, an autonomous vehicle can simultaneously optimize multiple goals as follows. - Safety: Avoid collisions and comply with traffic rules - Efficiency: Reach the destination quickly and minimize fuel consumption - Comfort: Avoid sudden acceleration, deceleration, and sudden steering operations These goals may conflict with each other, but with the Pareto optimization approach of the present invention, an optimal balance between these goals can be found. 2. **Adaptive Information Collection** With the adaptive manifestation mechanism of the present invention, the manifestation strategy can be dynamically adjusted to adapt to changes in the environment. By prioritizing exploration actions in high-uncertainty situations and goal achievement actions in low-uncertainty situations, more efficient information collection becomes possible. This enables more efficient and effective decision-making compared to static manifestation mechanisms. For example, when an autonomous vehicle approaches an intersection, if the uncertainty regarding the intentions of other vehicles (whether to go straight, turn right or left) is high, it can select an exploration action of reducing speed to observe the movements of other vehicles. On the other hand, when the intentions of other vehicles become clear, it can select an optimal action to pass through the intersection safely and efficiently. 3. **Improvement of Computational Efficiency** With the hierarchical abstraction algorithm of the present invention, it can operate efficiently even for POMDPs with a large state space, significantly reducing the computation time and memory usage. In the conventional approach, the size of the belief support MDP increases exponentially with respect to the number of states of the original POMDP, making it difficult to compute for large-scale problems. The hierarchical abstraction of the present invention alleviates this problem and makes it applicable to even larger-scale problems. Specifically, by hierarchically clustering the state space, the size of the belief support MDP can be significantly reduced. For example, in the case of a POMDP with 100 states, the normal belief support MDP has 2^100 - 1 states, but by dividing it into 10 clusters, the hierarchical belief support MDP is reduced to 2^10 - 1 = 1023 states. This can significantly reduce the calculation time and memory usage. 4. **Correspondence to Continuous Spaces** The continuous manifestation mechanism of the present invention can correspond to POMDPs with continuous state spaces and action spaces. Conventional approaches have only considered discrete state spaces and action spaces, but in the present invention, continuous variables can be directly handled. This enables application to a wider range of fields. For example, the continuous manifestation mechanism of the present invention can be applied to a POMDP having continuous state variables such as the position, speed, and acceleration of an autonomous vehicle, and continuous action variables such as the steering angle, accelerator, and brake. This allows the concept of manifestation to be applied to POMDPs with continuous state spaces and action spaces, enabling efficient decision-making. 5. **Achieving Both Theoretical Guarantee and Practicality** The present invention can improve practicality while maintaining the theoretical guarantees provided by Belly et al. By integrating four major technical features: multi-objective optimization, adaptive manifestation, hierarchical abstraction, and continuous space expansion, a theoretically accurate and practical algorithm can be provided. Specifically, the algorithm of the present invention provides theoretical guarantees for parity objective functions within the range of priorities 0, 1, and 2 in weak manifestation POMDPs, and for any parity objective function in strong manifestation POMDPs. At the same time, hierarchical abstraction and adaptive manifestation can improve computational efficiency and practicality. 6. **Expansion of Application Scope** The present invention can be applied to various application fields such as autonomous robots, autonomous driving vehicles, medical diagnosis systems, financial engineering, and smart grid control. In particular, it is effective for complex problems that require simultaneous optimization of multiple goals. For example, in a medical diagnosis system, it is necessary to simultaneously optimize multiple goals such as the accuracy of diagnosis, cost, and patient burden. Specifically, the effects of the present invention are expected in the following application fields. - Autonomous robots: While executing multiple tasks simultaneously, avoid obstacles and optimize energy efficiency - Medical diagnosis systems: Simultaneously optimize multiple goals such as the accuracy of diagnosis, cost, and patient burden - Financial engineering: Formulate investment strategies under uncertain market information while balancing risk and return - Smart grid control: Optimize multiple goals such as cost, stability, and environmental impact while balancing demand and supply
Modes for Carrying Out the Invention
[0015] Embodiments of the present invention will be described in detail below. Basic Definitions and Extensions A partially observable Markov decision process (POMDP) is defined as a tuple P = (Q, Act, Sig, δ, q0). Here, Q is a finite set of states, Act is a finite set of actions, Sig is a finite set of signals, δ: Q×Act → D(Sig×Q) is a transition function (D(X) represents the set of probability distributions of X), and q0 ∈ Q is the initial state. A play in a POMDP is an infinite sequence π = q0a1s1q1a2s2... ∈ (Q·Act·Sig)^ω, and for all i ≧ 0, δ(qi, ai+1)(si+1, qi+1) > 0 holds. A strategy in a POMDP is a function σ: (Act·Sig)* → D(Act) that makes decisions based on the observable history. A belief b is a probability distribution over the state set Q, and the belief support b is the support of the belief (the set of states with positive probability). The belief support is updated by a function B: 2^Q\{}. For b ∈ 2^Q\{}, a ∈ Act, and s ∈ Sig, define B(b, a, s) = {q' ∈ Q | ∃q ∈ b, δ(q, a)(s, q') > 0}. In the present invention, this basic POMDP model is extended as follows. # Multi-objective POMDP A multi-objective POMDP is defined as a tuple P = (Q, Act, Sig, δ, q0, O). Here, O = {O1, O2,..., Ok} is a set of objective functions, and each Oi ⊆ Q^ω is a set of sequences of states. The value of a strategy σ in a multi-objective POMDP is defined as a vector v(σ) = (PPσ[O1], PPσ[O2],..., PPσ[Ok]). Here, PPσ[Oi] is the probability of satisfying Oi under the strategy σ. A strategy σ is Pareto optimal if there is no other strategy that dominates v(σ). That is, there is no strategy σ' such that PPσ'[Oi] ≧ PPσ[Oi] for all i and PPσ'[Oj] > PPσ[Oj] for at least one j. As an approach to solving the multi-objective optimization problem, the weighted sum method is adopted. For each objective function Oi, a weight wi ∈ [0, 1] representing its importance is set, and under the constraint Σi wi = 1, a strategy σ that maximizes the weighted sum Σi wi·PPσ[Oi] is sought. As objective functions in a multi-objective POMDP, the following general objective functions are considered. 1. Parity objective function: For a priority function p: Q → {0, 1, ..., d}, Parity(p) = {q0q1... ∈ Q^ω | lim sup i≧0 p(qi) is even} is defined. This requires that the maximum priority among the states visited infinitely often is even. 2. Reachability objective function: For a set of target states F ⊆ Q, Reach(F) = {q0q1... ∈ Q^ω | ∃i ≧ 0, qi ∈ F} is defined. This requires reaching the target state sometime. 3. Safety objective function: For a set of dangerous states F ⊆ Q, Safety(F) = {q0q1... ∈ Q^ω | ∀i ≧ 0, qi not ∈ F} is defined. This requires always avoiding dangerous states. 4. Expected reward objective function: For a reward function r: Q × Act → R, ExpectedReward(r) = E[Σi≧0 γ^i·r(qi, ai+1)] is defined. Here, γ ∈ [0, 1) is the discount rate. This requires maximizing the expected value of the discounted reward. # Adaptive manifestation mechanism The adaptive manifestation mechanism is defined as a function E: D(Q) × 2^(Q^ω) → Act that selects a "search action" to promote manifestation based on the current belief state b and the objective function O. Specifically, it balances the search action and the goal achievement action based on the level of uncertainty H(b) in the current belief state b and the current degree of achievement V(b, O) for the objective function O. The level of uncertainty H(b) is defined as the entropy of the belief H(b) = -Σq∈Q b(q)·log(b(q)). The adaptive manifestation mechanism selects actions based on the following rules. - If H(b) > θ (θ is a threshold value), select the exploration action E(b, O). - If H(b) ≤ θ, select the optimal action a* = argmax_a V(b, O, a) for achieving the goal. Here, V(b, O, a) is the expected value of the objective function O after executing action a in the belief state b. V(b, O, a) is calculated as follows. V(b, O, a) = Σs∈Sig P(s|b, a)·V(Update(b, a, s), O) Here, P(s|b, a) is the probability of observing signal s after executing action a in the belief state b, and Update(b, a, s) is the updated belief state after executing action a in the belief state b and observing signal s. The exploration action E(b, O) is defined as follows. E(b, O) = argmax_a Σs∈Sig P(s|b, a)·(V(Update(b, a, s), O) + λ·(H(b) - H(Update(b, a, s)))) Here, λ > 0 is a parameter that adjusts the balance between exploration and goal achievement. This equation means selecting an action that maximizes the weighted sum of the expected value of the objective function after executing action a and the amount of reduction in uncertainty (information gain). The threshold value θ of the adaptive manifestation mechanism can also be dynamically adjusted according to changes in the environment. For example, when the urgency of goal achievement is high, the threshold value can be lowered to select goal achievement actions in more situations. Conversely, when long-term performance is important, the threshold value can be raised to select exploration actions in more situations. # Hierarchical Belief Support MDP The hierarchical belief support MDP is constructed by hierarchically abstracting the state space. Specifically, the state set Q is divided into k clusters C = {C1, C2, ..., Ck}. Each cluster Ci is a set of semantically similar states. Clustering is performed based on the "similarity" between states. The similarity between states q, q' ∈ Q is defined based on the following elements. - Similarity of transitions: The similarity of transition probabilities for the same action - Similarity of signals: The similarity of signals observed for the same action - Similarity of objective functions: The similarity of the contribution degrees to the objective function Specifically, the similarity S(q, q') between states q, q' ∈ Q is defined as follows. S(q, q') = α·Strans(q, q') + β·Ssig(q, q') + γ·Sobj(q, q') Here, α, β, γ > 0 are weight parameters, and α + β + γ = 1. Strans(q, q') is the similarity of transitions, Ssig(q, q') is the similarity of signals, and Sobj(q, q') is the similarity of objective functions. The similarity of transitions Strans(q, q') is defined as follows. Strans(q, q') = 1 - max_a Σq''∈Q |P(q''|q, a) - P(q''|q', a)| Here, P(q''|q, a) is the probability of transitioning to state q'' after executing action a in state q. The similarity of signals Ssig(q, q') is defined as follows. Ssig(q, q') = 1 - max_a Σs∈Sig |P(s|q, a) - P(s|q', a)| Here, P(s|q, a) is the probability of observing signal s after executing action a in state q. The similarity Sobj(q, q') of the objective functions is defined as follows. Sobj(q, q') = 1 - Σi wi·|Contrib(q, Oi) - Contrib(q', Oi)| Here, Contrib(q, Oi) is the contribution degree of state q to objective function Oi, and wi is the weight of objective function Oi. As the clustering algorithm, agglomerative hierarchical clustering is used. This algorithm first regards each state as an independent cluster, and constructs a hierarchical cluster structure by sequentially integrating the most similar cluster pairs. The number of clusters k is set according to the complexity of the problem and resource constraints. The hierarchical belief support MDP is defined as PH = (2^C, Act, δH, {C(q0)}). Here, C(q) is the cluster containing q, and δH is the transition function at the cluster level. δH is derived from the transition function δ of the original POMDP. Specifically, for cluster set b ∈ 2^C, action a ∈ Act, and cluster set b' ∈ 2^C, δH(b, a)(b') is defined as the probability of transitioning from the clusters included in b to the clusters included in b' after executing action a. δH(b, a)(b') = P(b'|b, a) = Σs∈Sig P(s|b, a)·[BH(b, a, s) = b'] Here, P(s|b, a) is the probability of observing signal s after executing action a in the set of clusters b, and BH(b, a, s) is the updated set of clusters after executing action a in the set of clusters b and observing signal s. [BH(b, a, s) = b'] is an indicator function that is 1 when BH(b, a, s) = b' and 0 otherwise. P(s|b, a) is calculated as follows. P(s|b, a) = Σq∈∪c∈b c b(q)·Σq'∈Q δ(q, a)(s, q') Here, b(q) is the probability of state q in belief state b. BH(b, a, s) is defined as follows. BH(b, a, s) = {C(q') | q' ∈ Q, ∃q ∈ ∪c∈b c, δ(q, a)(s, q') > 0} Here, C(q') is the cluster containing q'. For the hierarchical belief support MDP, an objective function similar to that of the original POMDP can be defined. For example, in the case of the parity objective function, the priority pH(c) of cluster c ∈ C is defined as the maximum priority of the states contained in c. pH(c) = max_{q∈c} p(q) The hierarchical belief support MDP defined in this way has a significantly reduced state space compared to the original POMDP, so the computational efficiency is improved. Also, by adjusting the number of clusters k, the level of abstraction can be controlled. The smaller k is, the higher the level of abstraction and the better the computational efficiency, but the accuracy may decrease. Conversely, the larger k is, the lower the level of abstraction and the better the accuracy, but the computational efficiency may decrease. # Continuous manifestation mechanism For a POMDP with a continuous state space \(X\subseteq\mathbb{R}^n\) and an action space \(A\subseteq\mathbb{R}^m\), the continuous manifestation mechanism is defined as a function \(R: X\times A\rightarrow[0, 1]\) that calculates the probability that a state \(x\in X\) is manifested. This function returns, for a state \(x\) and an action \(a\), the probability \(R(x, a)\) that the state \(x\) is manifested. This probability is calculated based on the observability of the state \(x\) and the information - gathering ability of the action \(a\).
[0016] Algorithm The present invention provides the following main algorithms. # Multi - objective optimization algorithm The multi - objective optimization algorithm is an algorithm that uses the weighted - sum method to find Pareto - optimal strategies for a multi - objective POMDP. Input: Multi - objective POMDP \(P=(Q, Act, Sig, \delta, q_0, O)\) and weights \(w=(w_1, w_2,\cdots, w_k)\) of the objective functions Output: Strategy \(\sigma\) that maximizes the weighted sum \(\sum_{i}w_i\cdot P_{P,\sigma}[O_i]\) 1. For each objective function \(O_i\), construct the belief - support MDP. 2. Define an appropriate priority function for each belief - support MDP. 3. Calculate the optimal strategy \(\sigma_w\) for the weighted objective function \(O_w=\sum_{i}w_i\cdot O_i\). 4. Lift the strategy \(\sigma_w\) to the original POMDP. The multi - objective optimization algorithm is implemented in the following steps. 1. Receive as input a multi - objective POMDP \(P=(Q, Act, Sig, \delta, q_0, O)\) and weights \(w=(w_1, w_2,\cdots, w_k)\) of the objective functions. 2. Construct the belief - support MDP \(P_B=(2^Q\setminus\{\{\}\}, Act, \delta_B, \{q_0\})\). a. For each belief support \(b\in 2^Q\setminus\{\{\}\}\), action \(a\in Act\), and signal \(s\in Sig\), compute \(B(b, a, s)=\{q'\in Q|\exists q\in b, \delta(q, a)(s, q') > 0\}\). b. For each belief support \(b\in 2^Q\setminus\{\{\}\}\), action \(a\in Act\), and belief support \(b'\in 2^Q\setminus\{\{\}\}\), compute \(\delta B(b, a)(b')=\sum_{s\in Sig}[B(b, a, s)=b']\cdot P(s|b, a)\). 3. For each objective function \(O_i\in O\), define the corresponding objective function \(O_{B_i}\) on the belief support MDP \(P_B\). a. In the case of the parity objective function \(O_i = Parity(p_i)\), let \(O_{B_i}=Parity(p_{B_i})\). Here, \(p_{B_i}(b)=\max_{q\in b}p_i(q)\). b. In the case of the reachability objective function \(O_i = Reach(F_i)\), let \(O_{B_i}=Reach(F_{B_i})\). Here, \(F_{B_i}=\{b\in 2^Q\setminus\{\{\}\}|b\cap F_i\neq\{\}\}\). c. In the case of the safety objective function \(O_i = Safety(F_i)\), let \(O_{B_i}=Safety(F_{B_i})\). Here, \(F_{B_i}=\{b\in 2^Q\setminus\{\{\}\}|b\cap F_i\neq\{\}\}\). 4. Define the weighted objective function \(O_{B_w}=\sum_{i}w_i\cdot O_{B_i}\). 5. Compute the optimal strategy \(\sigma_B\) for the objective function \(O_{B_w}\) on the belief support MDP \(P_B\). a. In the case of the parity objective function, apply the solution algorithm for parity games. b. In the case of the reachability objective function, apply the backward induction algorithm. c. In the case of the safety objective function, apply the attractor computation algorithm. 6. Lift the strategy \(\sigma_B\) to the original POMDP \(P\). Specifically, for each observable history \(h\in(Act\cdot Sig)^*\), let \(\sigma(h)=\sigma_B(B^*(\{q_0\}, h))\). 7. Output the strategy \(\sigma\). # Adaptive Explicitization Algorithm The adaptive explicitization algorithm is an algorithm that dynamically adjusts the balance between exploration actions and goal - achievement actions based on the uncertainty of the current belief state. Input: POMDP P = (Q, Act, Sig, δ, q0), objective function O, uncertainty threshold θ Output: Adaptive strategy σ 1. Set the current belief state b to the initial belief b0. 2. At each step t: a. Calculate the level of uncertainty H(b). b. If H(b) > θ, select the exploration action a = E(b, O). c. If H(b) ≤ θ, select the optimal action a = argmax_a V(b, O, a) for goal achievement. d. Execute the action a and observe the signal s. e. Update the belief state b: b' = Update(b, a, s). f. Set b = b' and proceed to the next step. The adaptive explicitization algorithm is implemented in the following steps. 1. Receive the POMDP P = (Q, Act, Sig, δ, q0), objective function O, and uncertainty threshold θ as inputs. 2. Calculate the initial belief state b0. Usually, b0(q0) = 1, b0(q) = 0 for q ≠ q0. 3. At each step t: a. Calculate the level of uncertainty of the current belief state b: H(b) = -Σq∈Q b(q)·log(b(q)). b. If H(b) > θ: i. For each action a ∈ Act, calculate the information gain IG(b, a) = H(b) - Σs∈Sig P(s|b, a)·H(Update(b, a, s)). ii. Select the action \(a^*=\arg\max_a IG(b, a)\) that maximizes the information gain. c. If \(H(b)\leq\theta\): i. For each action \(a\in Act\), calculate the expected value of the objective function \(V(b, O, a)=\sum_{s\in Sig}P(s|b, a)\cdot V(Update(b, a, s), O)\). ii. Select the action \(a^*=\arg\max_a V(b, O, a)\) that maximizes the expected value of the objective function. d. Execute the action \(a^*\) and observe the signal \(s\). e. Update the belief state \(b\): \(b'(q') = P(q'|b, a^*, s)=\sum_{q\in Q}b(q)\cdot\delta(q, a^*)(s, q') / P(s|b, a^*)\). f. Set \(b = b'\) and proceed to the next step. 4. Output the adaptive strategy \(\sigma\). \(\sigma\) is a function that executes the above steps for each observable history \(h\in(Act\cdot Sig)^*\) and selects the appropriate action. # Hierarchical Abstraction Algorithm The hierarchical abstraction algorithm is an algorithm that hierarchically clusters the state space and constructs a hierarchical belief support MDP. Input: POMDP \(P=(Q, Act, Sig, \delta, q_0)\), number of clusters \(k\) Output: Hierarchical belief support MDP \(P_H\) 1. Partition the state set \(Q\) into \(k\) clusters \(C = \{C_1, C_2,..., C_k\}\). a. Calculate the similarity matrix \(S\) between states. b. Apply a hierarchical clustering algorithm (e.g., agglomerative hierarchical clustering). c. Cut at the number of clusters \(k\) to obtain \(k\) clusters. 2. For each cluster \(C_i\), select a representative state \(r_i\in C_i\). 3. Calculate the transition probabilities between the representative states and construct the hierarchical belief support MDP \(P_H\). a. For a set of clusters $b \in 2^C$, an action $a \in Act$, and a set of clusters $b' \in 2^C$, calculate $\delta_H(b, a)(b')$. 4. Return the hierarchical belief support MDP $PH$. The hierarchical abstraction algorithm is implemented in the following steps. 1. Receive as input a POMDP $P=(Q, Act, Sig, \delta, q_0)$ and the number of clusters $k$. 2. Calculate the similarity matrix $S$ between states. a. For each state pair $q, q' \in Q$, calculate the transition similarity $S_{trans}(q, q')$, the signal similarity $S_{sig}(q, q')$, and the objective function similarity $S_{obj}(q, q')$. b. Calculate the similarity $S(q, q') = \alpha \cdot S_{trans}(q, q') + \beta \cdot S_{sig}(q, q') + \gamma \cdot S_{obj}(q, q')$. 3. Apply the agglomerative hierarchical clustering algorithm. a. First, consider each state $q \in Q$ as an independent cluster. b. Find the pair of clusters with the highest similarity and merge them. c. Repeat step b until the number of clusters becomes $k$. 4. For the resulting $k$ clusters $C = \{C_1, C_2, \ldots, C_k\}$, select a representative state $r_i \in C_i$ from each cluster $C_i$. The representative state is selected as the state with the maximum average similarity to other states within the cluster. 5. Construct the hierarchical belief support MDP $PH=(2^C, Act, \delta_H, \{C(q_0)\})$. a. For each set of clusters $b \in 2^C$, an action $a \in Act$, and a set of clusters $b' \in 2^C$, calculate $\delta_H(b, a)(b')$. b. Define $\delta_H(b, a)(b') = P(b'|b, a) = \sum_{s \in Sig} P(s|b, a) \cdot [B_H(b, a, s) = b']$. c. Calculate P(s|b, a) and BH(b, a, s) according to the above definitions. 6. Output the hierarchical belief support MDP PH. # Continuous Space Algorithm The continuous space algorithm is an algorithm that applies a continuous manifestation mechanism to POMDPs with continuous state spaces and action spaces. Input: POMDP PC with continuous state space X ⊆ Rn and action space A ⊆ Rm, objective function O Output: Continuous strategy σC 1. Define a continuous manifestation mechanism R: X × A → [0, 1]. a. Implement R(x, a) using a function approximator such as a neural network. b. Learn so that R(x, a) satisfies the properties of continuous manifestation. 2. Discretize the continuous state space and action space. a. Divide the state space X into n regions X1, X2,..., Xn. b. Divide the action space A into m representative actions a1, a2,..., am. 3. Apply a multi-objective optimization algorithm and an adaptive manifestation algorithm on the discretized space. 4. Interpolate the discrete strategy into a continuous strategy. a. For each state x ∈ X, interpolate the strategy of the nearest discrete state. 5. Return the continuous strategy σC.
[0017] # Integrated Algorithm The integrated algorithm is a "hierarchical adaptive multi-objective manifestation algorithm" that combines multi-objective optimization, adaptive manifestation, hierarchical abstraction, and continuous space extension. Input: Multi-objective POMDP P = (Q, Act, Sig, δ, q0, O), weight w of the objective function, uncertainty threshold θ, number of clusters k Output: Integrated Strategy σI 1. Apply the hierarchical abstraction algorithm to construct the hierarchical belief support MDP PH. 2. Apply the multi-objective optimization algorithm to calculate the strategy σH that maximizes the weighted sum Σi wi·PPHσ[Oi] on the hierarchical belief support MDP PH. 3. Apply the adaptive realization algorithm to extend the strategy σH to the adaptive strategy σA. 4. For POMDPs with continuous state and action spaces, apply the continuous space algorithm to extend the strategy σA to the continuous strategy σC. 5. Return the integrated strategy σI = σC. The integrated algorithm is implemented in the following steps. 1. Receive as input the multi-objective POMDP P = (Q, Act, Sig, δ, q0, O), the weights w of the objective functions, the uncertainty threshold θ, and the number of clusters k. 2. Apply the hierarchical abstraction algorithm to construct the hierarchical belief support MDP PH. a. Divide the state set Q into k clusters C = {C1, C2,..., Ck}. b. Construct the hierarchical belief support MDP PH = (2^C, Act, δH, {C(q0)}). 3. Apply the multi-objective optimization algorithm to calculate the strategy σH that maximizes the weighted sum Σi wi·PPHσ[Oi] on the hierarchical belief support MDP PH. a. For each objective function Oi ∈ O, define the corresponding objective function OHi on the hierarchical belief support MDP PH. b. Define the weighted objective function OHw = Σi wi·OHi. c. Calculate the optimal strategy σH for the objective function OHw on the hierarchical belief support MDP PH. 4. Apply the adaptive realization algorithm to extend the strategy σH to the adaptive strategy σA. a. For each set of clusters b ∈ 2^C, calculate the level of uncertainty H(b). b. If H(b) > θ, select the exploration action E(b, OHw). c. If H(b) ≤ θ, select the optimal action σH(b) for achieving the goal. 5. For POMDPs with continuous state and action spaces, apply a continuous space algorithm to extend the strategy σA to a continuous strategy σC. a. Define a continuous revelation mechanism R: X × A → [0, 1]. b. Discretize the continuous state and action spaces. c. Apply the strategy σA on the discretized space. d. Interpolate the discrete strategy to a continuous strategy σC. 6. Output the integrated strategy σI = σC. For POMDPs without continuous state and action spaces, let σI = σA.
[0018] Implementation Details The algorithm of the present invention can be implemented with the following pseudo-code. ``` function SolveMultiObjectiveRevealingPOMDP(P, O, w, θ, k): / / P: POMDP, O: set of objective functions, w: weight, θ: uncertainty threshold, k: number of clusters / / Execute hierarchical clustering C = HierarchicalClustering(P.states, k) / / Construct the hierarchical belief support MDP PH = ConstructHierarchicalBeliefSupportMDP(P, C) / / Define the weighted objective function Ow = WeightedObjective(O, w) / / Solve the hierarchical MDP σH = SolveHierarchicalMDP(PH, Ow) / / Extend to an adaptive strategy σA = AdaptiveStrategy(σH, θ) / / Extend to a continuous strategy (if necessary) if IsContinuous(P): σI = ContinuousStrategy(σA) else: σI = σA return σI ``` ``` function HierarchicalClustering(Q, k): / / Q: Set of states, k: Number of clusters / / Calculate the similarity matrix S = ComputeSimilarityMatrix(Q) / / Perform agglomerative hierarchical clustering C = AgglomerativeHierarchicalClustering(Q, S, k) return C ``` ``` function ConstructHierarchicalBeliefSupportMDP(P, C): / / P: POMDP, C: Set of clusters / / Select representative states R = SelectRepresentativeStates(C) / / Calculate the transition function δH = ComputeTransitionFunction(P, C, R) / / Construct the hierarchical belief support MDP PH = (2^C, P.actions, δH, {C(P.initialState)}) return PH ``` ``` function AdaptiveStrategy(σ, θ): / / σ: basic strategy, θ: threshold of uncertainty function SelectAction(b): / / b: current belief state / / Calculate uncertainty H = Entropy(b) if H > θ: / / Select exploration action return ExplorationAction(b) else: / / Select goal - achieving action return σ(b) return SelectAction ``` ``` function ContinuousStrategy(σ): / / σ: discrete strategy function SelectAction(x): / / x: continuous state / / Find the nearest discrete state i = FindNearestDiscreteState(x) / / Apply the discrete strategy return σ(i) return SelectAction ```
Example
[0019] As an example of the present invention, consider the decision-making problem in the traffic environment of an autonomous vehicle. In this problem, the autonomous vehicle has partial observability in that it cannot fully observe the intentions of other vehicles and pedestrians.
[0020] # Example 1: Multi-objective Optimization Consider the following three as the objective functions of the autonomous vehicle. 1. Safety: Avoid collisions (safety objective function) 2. Efficiency: Reach the destination quickly (reachability objective function) 3. Comfort: Avoid sudden acceleration and deceleration (parity objective function) Set weights (0.5, 0.3, 0.2) for these objective functions and apply the multi-objective optimization algorithm of the present invention. Specifically, execute the following steps. 1. Model the environment of the autonomous vehicle as a POMDP. The state is the position and speed of the host vehicle and surrounding vehicles, the actions are acceleration, deceleration, lane change, etc., and the signals are the observed values from the sensors. 2. Define three objective functions. The safety objective function is to avoid collision states, the efficiency objective function is to reach the destination, and the comfort objective function is to minimize the change in acceleration. 3. Set weights (0.5, 0.3, 0.2) for each objective function and find a strategy to maximize the weighted sum. 4. Apply the obtained strategy to the control of the autonomous vehicle. As a result, an optimal driving strategy that balances the three objective functions is obtained. For example, it becomes possible to drive while giving priority to safety and also considering efficiency and comfort. As a specific scenario, consider the situation where an autonomous vehicle is approaching an intersection. There are other vehicles and pedestrians at the intersection, and the autonomous vehicle cannot fully observe their intentions. In this situation, the following decision-making is carried out. - When prioritizing safety, the autonomous vehicle approaches the intersection carefully while maintaining a sufficient distance to avoid collisions with other vehicles and pedestrians. - When prioritizing efficiency, the autonomous vehicle tries to pass through the intersection quickly to reach the destination promptly. - When prioritizing comfort, the autonomous vehicle approaches the intersection with a smooth speed profile to avoid sudden acceleration and deceleration. By using the multi-objective optimization algorithm of the present invention, an optimal driving strategy that balances these objective functions can be obtained. For example, it becomes possible to drive while considering efficiency and comfort while most emphasizing safety. Specifically, while maintaining a sufficient distance to avoid collisions with other vehicles and pedestrians, the autonomous vehicle approaches the intersection at an appropriate speed to reach the destination quickly and maintains a smooth speed profile to avoid sudden acceleration and deceleration.
[0021] # Example 2: Adaptive Manifestation In a traffic environment, when the intentions of other vehicles are unclear (for example, when it is unclear whether to change lanes), the adaptive manifestation mechanism of the present invention selects a search action of temporarily reducing the speed to observe the movements of other vehicles. Specifically, the following steps are executed. 1. Calculate the level of uncertainty H(b) in the current belief state b. For example, calculate the entropy of the belief regarding the intentions of surrounding vehicles. 2. If H(b) exceeds the threshold θ, select a search action. For example, reduce the speed to observe the movements of other vehicles. 3. If H(b) is below the threshold θ, select the optimal action to achieve the goal. For example, proceed safely towards the destination. With this adaptive manifestation mechanism, the intentions of other vehicles are manifested, enabling more appropriate decision-making. For example, if it becomes clear that a surrounding vehicle intends to change lanes, the host vehicle can adjust its speed and respond safely. As a specific scenario, consider a situation where an autonomous vehicle is driving on a highway. There is another vehicle ahead, and it is unclear whether that vehicle will change lanes. In this situation, the following decision-making is carried out. - When the uncertainty regarding the intention of the other vehicle is high (H(b) > θ), the autonomous vehicle selects exploratory behavior. For example, it reduces its speed and observes the movement of the other vehicle. This may clarify whether the other vehicle intends to change lanes. - When the uncertainty regarding the intention of the other vehicle is low (H(b) ≤ θ), the autonomous vehicle selects the optimal behavior for achieving the goal. For example, if it becomes clear that the other vehicle has no intention of changing lanes, the autonomous vehicle can move forward safely. By using the adaptive manifestation mechanism of the present invention, the behavior can be dynamically adjusted according to the uncertainty of the environment. This enables more efficient and effective decision-making.
[0022] # Example 3: Hierarchical Abstraction The state space of the traffic environment is divided into clusters such as "the state of the host vehicle", "the state of surrounding vehicles", "the state of pedestrians", and "the state of signals", and a hierarchical belief support MDP is constructed. Specifically, the following steps are executed. 1. Cluster the state set Q based on similarity. For example, divide it into clusters such as the state of the host vehicle, the state of surrounding vehicles, the state of pedestrians, and the state of signals. 2. Select a representative state for each cluster. For example, in the state cluster of the host vehicle, select a typical combination of position and speed as the representative state. 3. Calculate the transition probabilities between representative states and construct a hierarchical belief-support MDP. 4. Solve the hierarchical belief-support MDP to obtain the optimal strategy at the cluster level. 5. Refine the cluster-level strategy into a state-level strategy. This hierarchical abstraction avoids the problem of the state space growing exponentially and enables efficient computation. For example, the state space of a traffic environment can be extremely large, but the computational complexity can be significantly reduced through hierarchical abstraction. As a specific scenario, consider the situation where an autonomous vehicle is driving in an urban area. There are numerous vehicles, pedestrians, traffic signals, etc. in the urban area, and the state space becomes extremely large. In this situation, the following hierarchical abstraction is performed. - Divide the state set Q into clusters such as "the state of the host vehicle", "the states of surrounding vehicles", "the states of pedestrians", and "the states of signals". - Calculate the similarity between states within each cluster and select representative states. For example, in the "state of the host vehicle" cluster, select a typical combination of position and speed as the representative state. - Calculate the transition probabilities between representative states and construct a hierarchical belief-support MDP. - Solve the hierarchical belief-support MDP to obtain the optimal strategy at the cluster level. - Refine the cluster-level strategy into a state-level strategy. By using the hierarchical abstraction algorithm of the present invention, it can also operate efficiently for POMDPs with a large-scale state space, and significantly reduce the computation time and memory usage. For example, in the case of a POMDP with 100 states, the normal belief-support MDP has 2^100 - 1 states, but by dividing it into 10 clusters, the hierarchical belief-support MDP is reduced to 2^10 - 1 = 1023 states. Thereby, the computation time and memory usage can be significantly reduced.
[0023] # Example 4: Continuous Space Expansion For a POMDP with continuous state variables such as the position, velocity, and acceleration of an autonomous vehicle, and continuous action variables such as the steering angle, accelerator, and brake, the continuous manifestation mechanism of the present invention is applied. Specifically, the following steps are executed. 1. Define a continuous manifestation mechanism R: X × A → [0, 1]. For example, using a neural network, calculate the probability R(x, a) that state x is manifested for state x and action a. 2. Discretize the continuous state space and action space. For example, discretize continuous state variables such as position, velocity, and acceleration within an appropriate range. 3. Apply a multi-objective optimization algorithm and an adaptive manifestation algorithm on the discretized space. 4. Interpolate the discrete strategy into a continuous strategy. For example, for each continuous state, interpolate the strategy of the closest discrete state. By this continuous space expansion, the algorithm of the present invention can also be applied to POMDPs with continuous state spaces and action spaces. For example, the control of an autonomous vehicle is essentially a continuous problem, but it can be effectively addressed by the continuous manifestation mechanism. As a specific scenario, consider the situation where an autonomous vehicle is driving in an environment with continuous state and action spaces. The state of the autonomous vehicle is represented by continuous variables such as position (x, y), velocity v, and direction θ. The action is represented by continuous variables such as acceleration a and steering angle δ. In this situation, the following continuous space expansion is performed. - Define a continuous manifestation mechanism R: X × A → [0, 1]. For example, using a neural network, calculate the probability R(x, a) that state x is manifested for state x and action a. - Discretize the continuous state space X within an appropriate range. For example, discretize the position (x, y) into a 10×10 grid, the velocity v into 5 levels, and the direction θ into 8 directions. - Discretize the continuous action space A within an appropriate range. For example, discretize the acceleration a into 5 levels and the steering angle δ into 7 levels. - Apply a multi-objective optimization algorithm and an adaptive realization algorithm on the discretized space. - Interpolate the discrete strategy into a continuous strategy. For example, for each continuous state, interpolate the strategy of the closest discrete state. By using the continuous space algorithm of the present invention, it is also possible to apply the concept of realization to POMDPs with continuous state spaces and action spaces and make efficient decisions. As a result, it can be applied to a wider range of application fields.
Industrial Applicability
[0024] The present invention is expected to have wide applications in the following fields. 1. **Autonomous Robotics** It can be applied to the control of autonomous robots that need to optimize multiple objectives (such as task completion, energy efficiency, safety, etc.) simultaneously. For example, an autonomous transport robot in a factory needs to optimize efficient path selection and collision avoidance simultaneously. The multi-objective optimization framework and adaptive realization mechanism of the present invention enable efficient decision-making while balancing these goals. Specifically, an autonomous transport robot in a factory needs to optimize the following multiple goals simultaneously. - Task completion: Quickly transport goods to the designated location - Energy efficiency: Minimize battery consumption - Safety: Avoid collisions with workers and other robots - Durability: Avoid sudden acceleration and deceleration and minimize mechanical wear These goals may conflict with each other, but with the multi-objective optimization framework of the present invention, an optimal balance between these goals can be found. Also, since the environment within the factory is constantly changing, with the adaptive manifestation mechanism of the present invention, efficient decision-making can be carried out in adaptation to the environmental changes. 2. **Medical Diagnosis System** It can be applied to formulating an optimal diagnosis and treatment plan based on uncertain symptom information while considering multiple health indicators. For example, a medical diagnosis system needs to simultaneously optimize multiple goals such as diagnostic accuracy, cost, and patient burden. With the multi-objective optimization framework of the present invention, an optimal diagnosis and treatment plan can be formulated while balancing these goals. Specifically, a medical diagnosis system needs to simultaneously optimize multiple goals as follows. - Diagnostic accuracy: Accurately diagnose diseases - Cost: Minimize the cost of examinations and treatments - Patient burden: Minimize invasive examinations and treatments with side effects - Time efficiency: Conduct diagnosis and treatment quickly These goals may conflict with each other, but with the multi-objective optimization framework of the present invention, an optimal balance between these goals can be found. Also, since the patient's condition is constantly changing, with the adaptive manifestation mechanism of the present invention, efficient decision-making can be carried out in adaptation to the changes in the patient's condition. 3. **Financial Engineering** It can be applied to formulating investment strategies under uncertain market information while balancing risk and return. For example, a portfolio management system needs to simultaneously optimize multiple goals such as expected return, risk, and liquidity. With the multi-objective optimization framework and adaptive manifestation mechanism of the present invention, an efficient investment strategy can be formulated while balancing these goals. Specifically, the portfolio management system needs to simultaneously optimize multiple goals as follows. - Expected return: Maximize the return on investment - Risk: Minimize the risk of investment - Liquidity: Enable the asset to be cashed out as needed - Diversity: Diversify investments across different asset classes These goals may conflict with each other, but the multi-objective optimization framework of the present invention can find the optimal balance between these goals. Also, since the market environment is constantly changing, the adaptive manifestation mechanism of the present invention can make efficient decisions adapting to the changes in the market environment. 4. **Autonomous Vehicles** It can be applied to the development of driving strategies that simultaneously optimize multiple objectives such as safety, efficiency, and comfort. For example, autonomous vehicles need to simultaneously optimize multiple goals such as collision avoidance, rapid arrival at the destination, and passenger comfort. The multi-objective optimization framework and the adaptive manifestation mechanism of the present invention can develop an efficient driving strategy while balancing these goals. Specifically, autonomous vehicles need to simultaneously optimize multiple goals as follows. - Safety: Avoid collisions and comply with traffic rules - Efficiency: Reach the destination quickly and minimize fuel consumption - Comfort: Avoid sudden acceleration, deceleration, and sudden steering operations - Adaptability: Adapt to changes in traffic conditions These goals may conflict with each other, but the multi-objective optimization framework of the present invention can find the optimal balance between these goals. Also, since the traffic environment is constantly changing, the adaptive manifestation mechanism of the present invention can make efficient decisions adapting to the changes in the traffic environment. 5. **Smart Grid Control** It can be applied to power management that optimizes multiple goals such as cost, stability, and environmental impact in the face of uncertainties in demand and supply. For example, a smart grid control system needs to simultaneously optimize multiple goals such as power supply stability, cost, and environmental impact. The multi-objective optimization framework and hierarchical abstraction algorithm of the present invention enable efficient power management while balancing these goals. Specifically, a smart grid control system needs to simultaneously optimize the following multiple goals. - Stability: Ensure the stability of power supply - Cost: Minimize the cost of power production and delivery - Environmental impact: Minimize the impact on the environment such as carbon dioxide emissions - Efficiency: Minimize power losses These goals may conflict with each other, but the multi-objective optimization framework of the present invention can find the optimal balance between these goals. In addition, since power demand and supply are constantly changing, the adaptive manifestation mechanism of the present invention can make efficient decisions adapting to the changes in demand and supply. 6. **Environmental Monitoring** It can be applied to a system that estimates the environmental state from limited sensor information and simultaneously monitors multiple environmental indicators. For example, an environmental monitoring system needs to simultaneously monitor multiple indicators such as air pollution, water quality, and ecosystem health. The multi-objective optimization framework and continuous manifestation mechanism of the present invention enable efficient environmental monitoring while balancing these indicators. Specifically, an environmental monitoring system needs to simultaneously optimize the following multiple goals. - Accuracy: Accurately estimate the environmental state - Comprehensiveness: Simultaneously monitor multiple environmental indicators - Cost: Minimize the cost of sensor placement and operation - Early warning: Detect environmental problems early These goals may conflict with each other, but the multi-objective optimization framework of the present invention can find the optimal balance between these goals. In addition, since the environmental state is constantly changing, the adaptive manifestation mechanism of the present invention can make efficient decisions adapting to the changes in the environmental state. 7. **Manufacturing industry** It can be applied to the optimization of production plans considering multiple goals such as quality, cost, and delivery time. For example, a production planning system needs to optimize multiple goals such as product quality, production cost, and delivery time simultaneously. The multi-objective optimization framework and hierarchical abstraction algorithm of the present invention enable efficient production planning while balancing these goals. Specifically, a production planning system needs to optimize the following multiple goals simultaneously. - Quality: Ensure product quality - Cost: Minimize production cost - Delivery time: Meet the delivery time - Flexibility: Develop a production plan that can respond to changes in demand These goals may conflict with each other, but the multi-objective optimization framework of the present invention can find the optimal balance between these goals. In addition, since demand and production conditions are constantly changing, the adaptive manifestation mechanism of the present invention can make efficient decisions adapting to the changes in demand and production conditions. 8. **Communication network management** It can be applied to network management considering multiple goals such as bandwidth, latency, and reliability. For example, a communication network management system needs to simultaneously optimize multiple goals such as efficient utilization of bandwidth, minimization of latency, and ensuring reliability. The multi-objective optimization framework and hierarchical abstraction algorithm of the present invention enable efficient network management while balancing these goals. Specifically, a communication network management system needs to simultaneously optimize the following multiple goals. - Bandwidth: Efficient utilization of bandwidth - Latency: Minimize the latency of data transfer - Reliability: Ensure the reliability of the network - Security: Ensure the security of the network These goals may conflict with each other, but the multi-objective optimization framework of the present invention can find the optimal balance between these goals. Also, since network traffic is constantly changing, the adaptive manifestation mechanism of the present invention can make efficient decisions adapting to traffic changes. 9. **Energy Management** It can be applied to energy management considering multiple goals such as energy consumption, cost, and environmental impact. For example, a building energy management system needs to simultaneously optimize multiple goals such as minimization of energy consumption, reduction of cost, and ensuring comfort. The multi-objective optimization framework and adaptive manifestation mechanism of the present invention enable efficient energy management while balancing these goals. Specifically, a building energy management system needs to simultaneously optimize the following multiple goals. - Energy consumption: Minimize energy consumption - Cost: Minimize energy cost - Comfort: Ensure comfort such as temperature, humidity, and lighting - Environmental impact: Minimize the environmental impact such as carbon dioxide emissions These goals may conflict with each other, but the multi-objective optimization framework of the present invention can find the optimal balance between these goals. Also, since the usage situation of the building and the external environment are constantly changing, the adaptive manifestation mechanism of the present invention can make efficient decisions adapting to the changes in the usage situation and the external environment. 10. **Agriculture** It can be applied to agricultural management considering multiple goals such as yield, quality, and resource efficiency. For example, a precision agriculture system needs to simultaneously optimize multiple goals such as maximizing yield, improving quality, and efficiently using resources such as water and fertilizers. The multi-objective optimization framework and the continuous manifestation mechanism of the present invention enable efficient agricultural management while balancing these goals. Specifically, a precision agriculture system needs to simultaneously optimize the following multiple goals. - Yield: Maximize the crop yield - Quality: Improve the crop quality - Resource efficiency: Efficiently use resources such as water, fertilizers, and pesticides - Environmental impact: Minimize the environmental impact such as soil erosion and groundwater pollution These goals may conflict with each other, but the multi-objective optimization framework of the present invention can find the optimal balance between these goals. Also, since weather conditions and the growth status of crops are constantly changing, the adaptive manifestation mechanism of the present invention can make efficient decisions adapting to the changes in weather conditions and the growth status of crops.
Claims
1. A multi-objective optimization system using a hierarchical adaptive manifestation mechanism in a partially observable Markov decision process (POMDP), hierarchical clustering means for dividing a state set into a plurality of clusters, means for constructing a hierarchical belief support Markov decision process (MDP) based on the clusters, means for defining a weighted objective function based on weighting of a plurality of objective functions and calculating an optimal strategy for the weighted objective function in the hierarchical belief support MDP, an adaptive manifestation mechanism for calculating a level of uncertainty in a current belief state, selecting a search action for promoting manifestation when the level of uncertainty exceeds a predetermined threshold, and selecting an action for optimizing the weighted objective function when the level of uncertainty is below the threshold, A multi-objective optimization system characterized by comprising the above.
2. In the multi-objective optimization system according to Claim 1, the plurality of objective functions include at least one of a parity objective function, a reachability objective function, and a safety objective function, the hierarchical clustering means includes means for calculating a similarity between states, means for performing agglomerative hierarchical clustering based on the similarity, and means for generating a plurality of clusters by cutting at a predetermined number of clusters A multi-objective optimization system characterized by the above.
3. In the multi-objective optimization system according to Claim 1 or 2, further comprising means for implementing a continuous manifestation mechanism for calculating a probability that a state is manifested for a POMDP having a continuous state space and action space, the continuous manifestation mechanism is implemented using a neural network, discretizes a continuous state space and action space, and interpolates a discrete strategy into a continuous strategy A multi-objective optimization system characterized by the above.
Citation Information
Cited By
Task processing method and device, electronic equipment and computer storage medium
CN121234198A
Task processing method and apparatus, electronic device, and computer storage medium
CN121234198B