A context-aware state representation and reinforcement learning decision method and system

CN122819355APending Publication Date: 2026-09-25FUJIAN UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611310853.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

欧氏空间在嵌入这类具有潜在层次结构的数据时,会面临严重的维度灾难与扭曲问题

Benefits of technology

[0033]决策框架的几何重构:首次构建了上下文双曲编码→双曲状态表征→测地线距离→策略分布→双曲空间价值函数这一完整的、完全根植于双曲几何的强化学习决策闭环。决策的每一个核心环节都利用了流形的负曲率特性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819355A_ABST
    Figure CN122819355A_ABST
Patent Text Reader

Abstract

The application discloses a state representation and reinforcement learning decision method and system based on context perception, and aims to solve the problem that the existing Euclidean space decision method is difficult to effectively depict the interactive context hierarchy structure. The core lies in that the decision closed loop is completely constructed on a hyperbolic manifold: firstly, historical context is encoded into a state representation point on the manifold by a hyperbolic gated recurrent unit, and the logarithm and exponential mapping is used to update between the manifold and its tangent space recursively; then, the geodesic distance between the state point and each action embedding point is directly calculated, and the strategy probability is generated through a monotonically decreasing nonlinear mapping; the state value is also estimated after the representation point is mapped to the tangent space. All manifold parameters are updated by Riemannian gradient optimization. The application makes the decision boundary naturally adapt to the hierarchical clustering of context, and significantly improves the convergence speed and strategy quality in the hierarchical sequential decision task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and reinforcement learning technology, specifically relating to a reinforcement learning method and system that combines context-aware state representation with decision-making in non-Euclidean geometric space, and is particularly suitable for sequential decision-making tasks that require capturing the hierarchical structure of interaction history and long-term environmental dependencies. Background Technology

[0002] In reinforcement learning, an agent makes decisions based on its current state. Traditional methods often use the current observation as the state or encode historical information in Euclidean space using recurrent neural networks. However, many real-world sequential decision-making processes, such as dialogue management and robot hierarchical task planning, naturally exhibit power-law distributions and tree-like hierarchical structures in their state transitions and context structures. When embedding such potentially hierarchical data into Euclidean space, it faces severe problems of the curse of dimensionality and distortion.

[0003] On the other hand, existing context-aware methods, such as Transformer-based large-scale decision models, while capable of capturing long-term dependencies, suffer from large parameter sizes, high computational costs, and their state representations are still based on Euclidean geometry assumptions, failing to demonstrate the ability to leverage the hierarchical geometric structure within the context to simplify policy learning and value estimation. Hyperbolic geometry has been shown to continuously embed hierarchical data with extremely low dimensionality and minimal distortion, but the deep integration of hyperbolic space as a representation carrier with reinforcement learning decision mechanisms remains a technological gap. Summary of the Invention

[0004] This invention aims to provide a context-aware hyperbolic state representation and reinforcement learning decision-making method and system. The method encodes historical interaction context into a hyperbolic manifold with negative curvature, directly constructs the policy and value function using order-preserving geodesic distances within this space, and optimizes them through Riemann gradient descent, thereby achieving significant improvements in decision efficiency, representation compactness, and convergence speed.

[0005] To achieve this objective, the present invention provides a context-aware state representation and reinforcement learning decision-making method, comprising the following steps:

[0006] Step S1: Collect the interaction history between the agent and the environment, and combine the observations and actions within a preset window length in chronological order to form a context sequence;

[0007] Step S2: Input the context sequence into the pre-constructed hyperbolic state coding network to generate the state representation point located on the hyperbolic manifold with negative curvature at the current time; the hyperbolic state coding network includes a hyperbolic gated recurrent unit, which pulls the point on the manifold to the tangent space for gating operation through logarithmic mapping, and then pushes the result back to the manifold through exponential mapping to recursively update the hidden state;

[0008] Step S3: Obtain the learnable action points that correspond one-to-one with each discrete action and are directly embedded on the hyperbolic manifold;

[0009] Step S4: On the hyperbolic manifold, the geodesic distance between the state representation point and each action embedding point is calculated using the Möbius summation method and the inverse hyperbolic tangent operation to form a distance vector;

[0010] Step S5: Using the negative value of the distance vector as input, perform nonlinear mapping through the flexible maximum transfer function to generate the selection probability of all actions, wherein the probability value is negatively correlated with the geodesic distance;

[0011] Step S6: Based on the action, select a probability distribution to sample and execute an action, obtain the immediate reward and the next observation from the environmental feedback, and update the context sequence;

[0012] Step S7: The state representation points are pulled to the origin tangent space through logarithmic mapping, and the state value estimate is output through the feedforward network. Based on the near-end policy optimization method, a composite objective function including truncated policy loss and value error loss is constructed.

[0013] Step S8: Calculate the Riemann gradient for the parameters on the hyperbolic manifold, calculate the Euclidean gradient for the parameters in the tangent space, and use the Riemann optimizer to update all network parameters uniformly until the policy converges.

[0014] Preferably, the state update process of the hyperbolic gated loop unit in step S2 is as follows:

[0015] First, transform the hyperbolic hidden state of the previous time step and the hyperbolic embedding of the current input into the origin tangent space through the origin logarithmic mapping to obtain the corresponding tangent vectors;

[0016] Calculate the update gate, reset gate, and candidate tangent vectors within the tangent space;

[0017] The candidate tangent vectors are reconstructed into candidate hidden state points through the origin exponential mapping;

[0018] Finally, using the previous hidden state point as the base point and the candidate hidden state point as the endpoint, geodesic interpolation is performed on the manifold using the update gate as the interpolation coefficient to obtain the hyperbolic hidden state at the current time.

[0019] Preferably, the calculation of the geodesic distance in step S4 is based on the Poincaré sphere model, whose expression includes applying an inverse hyperbolic tangent function to the norm of the Möbius summation result, and is parameterized by the curvature constant and the dimension.

[0020] Preferably, the flexible maximum transfer function in step S5 includes a learnable or preset temperature coefficient τ to control the concentration of the probability distribution.

[0021] Preferably, the feedforward network in step S7 includes at least one nonlinear activation layer. The input of the feedforward network is the logarithmic mapping coordinates of the state representation point in the tangent space at the origin, and the output is the value estimate of the state.

[0022] Preferably, the Riemann optimizer in step S8 updates each manifold parameter as follows: first, the Riemann gradient of the parameter is calculated; based on the Riemann gradient, the adaptive moment estimate update amount is calculated in the tangent space; and then, the update amount is projected from the tangent space back to the hyperbolic manifold through an exponential mapping to ensure that the updated parameter is always within the domain of the manifold.

[0023] To achieve this objective, the present invention also provides a context-aware state representation and reinforcement learning decision-making system, comprising:

[0024] The context collection module is used to continuously acquire and assemble observation and action history of a preset length to form a context sequence;

[0025] The hyperbolic state encoding module, with a built-in hyperbolic gated loop unit, is used to encode the context sequence into state representation points on the hyperbolic manifold;

[0026] An action embedding storage module is used to maintain and output the learnable embedding point of each discrete action on the hyperbolic manifold;

[0027] The geometric decision module connects the hyperbolic state encoding module and the action embedding storage module, and is used to perform the operations described in steps S4 and S5 to generate an action probability distribution.

[0028] The environment interaction module is used to execute the selected action and return the reward and the next observation;

[0029] The value function evaluation module is used to map the state representation points to the tangent space and then output the state value through a feedforward network.

[0030] The parameter update module uses a Riemann optimizer to perform gradient updates on the parameters on the hyperbolic manifold and performs standard Euclidean optimization on the parameters in the tangent space.

[0031] To achieve this objective, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a context-aware state representation and reinforcement learning decision-making method as described in any of the preceding claims.

[0032] Beneficial effects

[0033] A geometric reconstruction of the decision-making framework: For the first time, a complete reinforcement learning decision-making loop was constructed, fully rooted in hyperbolic geometry, consisting of contextual hyperbolic encoding → hyperbolic state representation → geodesic distance → policy distribution → hyperbolic space value function. Each core step of the decision-making process utilizes the negative curvature property of the manifold.

[0034] The pure distance-driven policy mechanism determines the policy probability entirely by the geodesic distance between the state and action on the manifold, abandoning the traditional approach that relies on inner products. This gives the decision boundary a natural hierarchical adaptability, making policy switching smoother and more structured when the context jumps between different abstract clusters.

[0035] Recursive state updates on manifolds: The design of hyperbolic gated loop units makes the dynamic organization of historical context itself conform to hyperbolic geometry, allowing hidden states to move toward the boundary on the manifold as information accumulates, dynamically encoding the ever-deepening context structure.

[0036] Compared to existing Euclidean space-based decision-making methods, this method, due to the matching of state representation and decision geometry, achieves faster convergence, higher final reward, and lower policy variance for decision-making tasks with hierarchical contexts under the same network capacity. Its decision-making process is more interpretable because the action probability is directly related to its relative position within the hierarchical context structure. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0038] Figure 1 This is an overall flowchart of the method of the present invention;

[0039] Figure 2 This is a schematic diagram of the module architecture of the system of the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] The technical problem this invention aims to solve is to overcome the fundamental mismatch between the Euclidean geometry assumptions of reinforcement learning state representations and the latent hierarchical structure of sequential decision contexts in existing technologies. In many practical tasks, such as intent transitions in human-computer dialogues and hierarchical planning in robotic tasks, the interaction history naturally constitutes an implicit tree-like or power-law structure from abstract goals to concrete operations. Forcibly embedding such a structure in Euclidean space not only produces high distortion but also forces the policy to learn extremely complex nonlinear boundaries. Although hyperbolic geometry has been proven to be a natural carrier of hierarchical data, how to extend it from representation learning to a complete, end-to-end reinforcement learning decision loop, especially by directly defining policies and value functions on hyperbolic manifolds, remains an unsolved core problem.

[0042] To address this problem, this invention provides a context-aware hyperbolic state representation and reinforcement learning decision-making method and system. Its core concept is to construct the entire decision-making process—from historical context encoding, state representation, action space organization, to policy evaluation and value estimation—on a hyperbolic manifold with constant negative curvature. In this manifold, the geodesic distance increases exponentially from the root node to the leaf node. This geometric property is directly utilized by this invention as a natural measure of decision probability, allowing action clustering based on similar contexts and state separation at different abstraction levels to emerge naturally without complex parameter learning.

[0043] To achieve this objective, the method of the present invention includes the following steps, the overall flowchart of which is shown below. Figure 1 :

[0044] Step S1: Context Sequence Construction. Continuously collect interaction tuples between the agent and the environment, each tuple containing observations and actions. Based on a set time window length, assemble these tuples into a context sequence in chronological order. Each element in this sequence is formed by concatenating the current observation features with the action features from the previous time step, providing local, real-time input for subsequent encoding.

[0045] Step S2: Hyperbolic State Encoding Network Generates Representation. The context sequence obtained in Step S1 is input into a pre-constructed hyperbolic state encoding network. The core of this network is a hyperbolic gated recurrent unit. Its design principle is as follows: at each time step, the previous hidden states on the manifold and the current input point are first pulled back to the tangent space of the origin through a logarithmic mapping, where the gated signal is calculated and mixed with the candidate states; then, the new mixed state vector is pushed back to the hyperbolic manifold through an exponential mapping, completing the recursive update of the state. After element-wise processing of the entire context sequence, the network finally outputs a point on the manifold at the current time as the hyperbolic state representation. This process ensures that the hierarchical information of the entire long-range context is compressed and encoded into the geometric position of the representation point, and its "depth" on the manifold often corresponds to the abstraction level of the context.

[0046] Step S3: Action Embedding Initialization and Maintenance. For each executable action in the discrete action space, a learnable point, called the action embedding point, is directly maintained on the same hyperbolic manifold. These points share the same manifold geometry with the state representation, forming the reference frame for decision-making.

[0047] Step S4: Geodesic Distance Calculation. Calculate the geodesic distance between the state representation points generated in Step S2 and each action embedding point in Step S3. This distance is not a Euclidean straight-line distance, but is calculated strictly along the shortest path of the hyperbolic manifold. Its mathematical expression involves Möbius summation and inverse hyperbolic tangent operations. This results in a distance vector composed of multiple geodesic distance values, where each component directly quantifies the compatibility of the current context state with each candidate action in the hierarchical geometry.

[0048] Step S5: Distance-based policy probability generation. The distance vector obtained in Step S4 is transformed using a nonlinear, monotonically decreasing transformation to generate a probability distribution over all executable actions. This transformation ensures that the probability of an action being selected is strictly negatively correlated with its geodesic distance to the current state representation—that is, actions closer to the current state receive higher probability quality. This replaces the linear decision boundary based on inner product in traditional methods, allowing the decision boundary to naturally resemble a layered von Neumann sphere on the manifold.

[0049] Step S6: Action Execution and Context Update. The agent samples an action from the probability distribution generated in step S5 and executes it, receives an immediate reward from the environment and the next observation, and then pushes the new interaction tuple into the context window to form the context sequence for the next time step, completing the closed-loop interaction with the environment.

[0050] Step S7: Hyperbolic-driven value assessment and loss construction. The hyperbolic state representation points obtained in Step S2 are pulled back to the tangent space of the manifold origin through a logarithmic mapping and flattened into a vector. This vector is then passed through a feedforward network containing a nonlinear activation function, outputting a scalar as the value estimate for that state. Based on a complete interaction trajectory, a composite reinforcement learning objective function, including truncated policy loss and value error loss, is jointly constructed using the advantage function estimation method.

[0051] Step S8: Parameter Update for Hybrid Geometry. All learnable parameters of the system are updated by category. For parameters located on the hyperbolic manifold, such as action embedding points and biases in hyperbolic gated units, their Riemann gradients are calculated; for parameters located in Euclidean tangent space, such as gating matrices, standard Euclidean gradients are calculated. Finally, a unified Riemann adaptive moment estimation optimizer is used to complete parameter iteration in their respective geometric spaces. This step ensures that throughout the training process, all manifold parameters are strictly maintained within the domain defined by the manifold through the shrinkage effect of the exponential mapping.

[0052] Correspondingly, the modular architecture of the system of the present invention refers to Figure 2 It includes: a context collection module for executing step S1; a hyperbolic state encoding module with a built-in hyperbolic gated loop unit that implements the logic described in step S2; an action embedding storage module for managing learnable action points in step S3; a geometric decision module for executing steps S4 and S5 and outputting action probabilities; an environment interaction module for executing step S6; a value function evaluation module for executing the value output part in step S7; and a parameter update module for executing the hybrid geometric gradient update in step S8.

[0053] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and formulas.

[0054] Example 1

[0055] The hyperbolic space model used in this invention is the Poincaré sphere model. The curvature constant is defined. The dimensions of the Poincaré sphere are The manifold is defined as:

[0056]

[0057] in, For a Poincaré spherical manifold, The curvature constant; For manifold dimension; Let x be the L2 norm of vector x.

[0058] The core nonlinear operations defined above are as follows:

[0059] Möbius method: for ,

[0060]

[0061] in, Let be any point on the manifold. Let be any other point on the manifold;

[0062] This operation ensures that the result remains within the Poincaré sphere manifold.

[0063] Exponential mapping: mapping points Vectors in tangent space Map back to the point on the manifold:

[0064]

[0065] in, For conformal factor, ; A vector in the tangent space;

[0066] Logarithmic mapping: mapping points on a manifold Mapped to The tangent space in which it is located.

[0067]

[0068] in, It is the inverse hyperbolic tangent function;

[0069] Based on the above calculations, the geodesic distance between two points has the following nonlinear analytical form:

[0070]

[0071] The steps of the method are described in detail:

[0072] Step S1: Constructing the context sequence

[0073] Set the context window length to At that moment Collecting Pre-processing An interactive tuple, observation sequence Action sequence Each pair of observations is concatenated with the vector from the previous action. If it's the initial step, the action is set to zero, forming the first concatenation. Original features of the time step: ; The entire context sequence: .

[0074] Step S2: Hyperbolic state coding network generates representations

[0075] The core of the hyperbolic state-coding network is the hyperbolic gated recurrent unit (HyperGRU), whose workflow is as follows: first, it passes through a linear layer... Each Transform to the tangent space at the origin, and then embed the manifold using an exponential mapping:

[0076]

[0077] in, The input is the embedding matrix; The i-th step inputs the embedding point on the manifold; The origin of the Poincaré ball.

[0078] Initialize hidden state ,for Perform the following hyperbolic state transition:

[0079] 1. Logarithmic mapping pulls back to the tangent space of the origin:

[0080] ,

[0081] Let be the logarithmic mapping of the previous hidden state in the tangent space at the origin. For the first The hidden state point of the step, It is a logarithmic mapping embedded in the tangent space at the origin for the current input.

[0082] 2. Update door, reset door:

[0083] ,

[0084] To update the gate vector, To reset the gate vector, It is a sigmoid activation function. To update the weight matrix of the hidden state of the gate, To reset the weight matrix of the hidden state of the gate, To update the weight matrix at the gate input, To reset the weight matrix at the gate input, To update the gate bias vector, To reset the gate bias vector, It is a logarithmic mapping embedded in the tangent space at the origin for the current input.

[0085] 3. Candidate tangent vector:

[0086]

[0087] Candidate tangent vector, For element-wise multiplication, The hidden state weight matrix is... The input weight matrix, This is the bias vector.

[0088] 4. Candidate hidden state points: , These are candidate hidden state points.

[0089] 5. Manifold geodesic interpolation updates hidden states: ,

[0090] After window iteration is complete, retrieve the representation at the current time step: , Let be the hyperbolic state representation point at time t.

[0091] Step S3: Action Embedding

[0092] Discrete Actions Learnable embedding points within the manifold Initialize the random cutting vector. The mapping is obtained.

[0093] Step S4: Calculate the geodesic distance vector

[0094] Distance vector , No. dimension:

[0095]

[0096] This is the geodesic distance vector. For the state point to the Geodesic distance between action points For the first Hyperbolic embedding point of each action.

[0097] Step S5: Generate action probability distribution

[0098] Introduce learnable or fixed temperature hyperparameters The strategy is determined by converting distance into a nonlinear mapping of probabilities:

[0099]

[0100] In the state Select action The probability, For temperature hyperparameters, The size of the discrete action space.

[0101] This distribution satisfies the following condition: the closer the action embedding is to the state representation point, the higher the probability of it being selected. The entire decision boundary is a von Neumann sphere in hyperbolic space, exhibiting hierarchical adaptability.

[0102] Step S6: Execution and Interaction

[0103] Sampling action Interact with the environment to earn rewards for each step. Next observation ;tuple Store the sequence in the replay buffer and construct the context sequence for the next time step using a sliding window.

[0104] Step S7: Value Estimation and Loss Construction

[0105] To estimate the state value, the representation points are pulled back to the Euclidean space through a logarithmic mapping from the origin:

[0106]

[0107] Let be the flattened vector of the state point in the tangent space at the origin.

[0108] Subsequently, the value network outputs through a feedforward network containing a hidden layer with a non-linear activation function such as ReLU. Output scalar value , for The value estimate, It is a feedforward network for value assessment.

[0109] A variant of the proximal policy optimization algorithm is used to construct the loss. The advantage value is calculated using generalized advantage estimation. The strategy loss is:

[0110]

[0111] in, To cut off strategy losses, This is the estimate of the generalized advantage. The probability ratio between the old and new strategies. For the set of policy network parameters, , This is the cutoff constant. Value loss is calculated using the mean squared error:

[0112]

[0113] To cut off strategy losses, To achieve the target return value, This is the set of parameters for the value network.

[0114] An additional policy entropy regularization term is added to encourage exploration.

[0115] Step S8: Hybrid Riemann gradient update

[0116] System parameters are divided into two categories: parameters defined in Euclidean space and parameters located on Poincaré spherical manifolds. For learnable parameters on the Poincaré sphere, and for Euclidean parameters, the standard Euclidean gradient is calculated directly. And apply the Adam update rule for the manifold parameters. First, calculate the Euclidean gradient. Then, the Riemann gradient is obtained through the inverse conformal factor transformation:

[0117]

[0118] It is the Euclidean gradient. For the Riemann gradient, Let be any learnable parameter on the manifold.

[0119] The parameters are then updated using a Riemann adaptive moment estimator optimizer, whose update formula applies an exponential mapping as a contraction mapping:

[0120]

[0121] For the updated manifold parameters, The learning rate; this process ensures that all manifold parameters remain constant during training. Inside.

[0122] Repeat steps S1 to S8 until the policy converges. In the discrete action maze and hierarchical recommendation simulation environment of this embodiment, compared with the policy gradient method based on Euclidean LSTM, the present invention reduces the number of convergence steps by about 30% and has a lower final score variance, confirming its inventive and beneficial effects.

[0123] The above technical solutions only embody the preferred technical solutions of the present invention. Any modifications that may be made by those skilled in the art to certain parts thereof embody the principles of the present invention and fall within the protection scope of the present invention.

Claims

1. A context-aware state representation and reinforcement learning decision-making method, characterized in that, Includes the following steps: Step S1: Collect the interaction history between the agent and the environment, and combine the observations and actions within a preset window length in chronological order to form a context sequence; Step S2: Input the context sequence into the pre-constructed hyperbolic state coding network to generate the state representation point located on the hyperbolic manifold with negative curvature at the current time; the hyperbolic state coding network includes a hyperbolic gated recurrent unit, which pulls the point on the manifold to the tangent space for gating operation through logarithmic mapping, and then pushes the result back to the manifold through exponential mapping to recursively update the hidden state; Step S3: Obtain the learnable action points that correspond one-to-one with each discrete action and are directly embedded on the hyperbolic manifold; Step S4: On the hyperbolic manifold, the geodesic distance between the state representation point and each action embedding point is calculated using the Möbius summation method and the inverse hyperbolic tangent operation to form a distance vector; Step S5: Using the negative value of the distance vector as input, perform nonlinear mapping through the flexible maximum transfer function to generate the selection probability of all actions, wherein the probability value is negatively correlated with the geodesic distance; Step S6: Based on the action, select a probability distribution to sample and execute an action, obtain the immediate reward and the next observation from the environmental feedback, and update the context sequence; Step S7: The state representation points are pulled to the origin tangent space through logarithmic mapping, and the state value estimate is output through the feedforward network. Based on the near-end policy optimization method, a composite objective function including truncated policy loss and value error loss is constructed. Step S8: Calculate the Riemann gradient for the parameters on the hyperbolic manifold, calculate the Euclidean gradient for the parameters in the tangent space, and use the Riemann optimizer to update all network parameters uniformly until the policy converges.

2. The context-aware state representation and reinforcement learning decision-making method according to claim 1, characterized in that, The state update process of the hyperbolic gated loop unit in step S2 is as follows: First, transform the hyperbolic hidden state of the previous time step and the hyperbolic embedding of the current input into the origin tangent space through the origin logarithmic mapping to obtain the corresponding tangent vectors; Calculate the update gate, reset gate, and candidate tangent vectors within the tangent space; The candidate tangent vectors are reconstructed into candidate hidden state points through the origin exponential mapping; Finally, using the previous hidden state point as the base point and the candidate hidden state point as the endpoint, geodesic interpolation is performed on the manifold using the update gate as the interpolation coefficient to obtain the hyperbolic hidden state at the current time.

3. The context-aware state representation and reinforcement learning decision-making method according to claim 1, characterized in that, The calculation of the geodesic distance in step S4 is based on the Poincaré sphere model, whose expression includes applying an inverse hyperbolic tangent function to the norm of the Möbius summation result, and is parameterized by the curvature constant and the dimension.

4. The context-aware state representation and reinforcement learning decision-making method according to claim 1, characterized in that, The flexible maximum transfer function described in step S5 includes a learnable or preset temperature coefficient τ, which is used to control the concentration of the probability distribution.

5. The context-aware state representation and reinforcement learning decision-making method according to claim 1, characterized in that, The feedforward network described in step S7 includes at least one nonlinear activation layer. The input of the feedforward network is the logarithmic mapping coordinates of the state representation point in the tangent space at the origin, and the output is the value estimate of the state.

6. A context-aware state representation and reinforcement learning decision-making method according to any one of claims 1 to 5, characterized in that, The Riemann optimizer in step S8 updates each manifold parameter as follows: First, the Riemann gradient of the parameter is calculated. Based on the Riemann gradient, the adaptive moment estimate update amount is calculated in the tangent space. Then, the update amount is projected from the tangent space back to the hyperbolic manifold through an exponential mapping to ensure that the updated parameter is always within the domain of the manifold.

7. A system utilizing the context-aware state representation and reinforcement learning decision-making method according to any one of claims 1 to 6, characterized in that, include: The context collection module is used to continuously acquire and assemble observation and action history of a preset length to form a context sequence; The hyperbolic state encoding module, with a built-in hyperbolic gated loop unit, is used to encode the context sequence into state representation points on the hyperbolic manifold; An action embedding storage module is used to maintain and output the learnable embedding point of each discrete action on the hyperbolic manifold; The geometric decision module, which connects the hyperbolic state encoding module and the action embedding storage module, is used to perform the operations of steps S4 and S5 as described in claim 1 to generate an action probability distribution. The environment interaction module is used to execute the selected action and return the reward and the next observation; The value function evaluation module is used to map the state representation points to the tangent space and then output the state value through a feedforward network. The parameter update module uses a Riemann optimizer to perform gradient updates on the parameters on the hyperbolic manifold and performs standard Euclidean optimization on the parameters in the tangent space.

8. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a context-aware state representation and reinforcement learning decision-making method as described in any one of claims 1 to 6.