Spatiotemporal fusion reasoning and lifelong cognitive learning methods for behavioral evolution in open environments
Through the collaborative scene migration and hierarchical decision-making model framework, a lifelong cognitive learning method in an open environment is constructed, which solves the problem of insufficient adaptability of robot systems in open environments and realizes the efficient learning and complex environment adaptation of intelligent robots.
Patent Information
- Application Number
- CN202211300756.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-24
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2042-10-24
AI Technical Summary
Existing robotic systems lack a unified cognitive architecture in open environments, making it difficult for them to effectively adapt to environmental changes and conduct lifelong learning, especially in complex and unknown environments where it is difficult to extract adaptive information from past experience.
A model framework of collaborative scenario migration, analogical reasoning and hierarchical decision-making is adopted. A lifelong cognitive learning method in an open environment is constructed through real-time observation and semi-Markov decision-making models. Conditional random fields and Markov Monte Carlo chains are used to generate joint strategies, adjust the degrees of freedom and confidence distribution of autonomous learning, and build a hierarchical machine learning architecture.
It improves the adaptability and learning efficiency of intelligent robots in open environments, realizes multi-target global perception and multi-dimensional decision-making, forms sustainable lifelong learning capabilities, adapts to complex environmental changes and optimizes action strategies.
Smart Images

Figure CN115526270B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence robotics technology, and specifically relates to a spatiotemporal fusion reasoning and lifelong cognitive learning method for behavior evolution in an open environment. Background Art
[0002] In recent years, artificial intelligence (AI) has been widely researched in applications such as Go and autonomous driving, where it aims to solve complex problems through real-time interaction with the environment. Furthermore, it has been used in robotics to autonomously learn skills such as jumping and grasping. Neuroscience suggests that the essence of intelligence lies in two key aspects: the ability to learn and adapt to the environment; and the ability to cognitively anticipate unknown changes. For AI and robotic systems to successfully operate and adapt in the real world, they rely on a hierarchy of cognitive mechanisms, including compositional abstraction and predictive processing.
[0003] Sustainable lifelong learning in open environments faces two challenges: how to thoroughly explore sparse environments and abstract them into actionable tasks, and how to evaluate the credibility of current state transitions. Evolutionary biology suggests that the development of advanced problem-solving skills in biological agents is a process of continuous integration of cognitive mechanisms over a long evolutionary process, adapting to the environment and improving skills, thereby acquiring intelligent action patterns characterized by adaptability and lifelong learning. When solving complex cognitive tasks, humans first intuitively recall their vast memory and reason carefully about previously learned experiences based on temporal evidence. Therefore, humans' ability to adapt to changing environments remains superior to many AI systems. Research results indicate that many existing reinforcement learning-based methods, which attempt to address zero-shot problem solving and transfer learning, are currently only applicable to tasks that are identical or similar to previous tasks or tasks in simple generative domains. However, a unified cognitive architecture for open environments is still lacking to integrate all evolutionary mechanisms that can generate environmentally adaptive actions. Summary of the Invention
[0004] The present invention is designed to solve the above-mentioned problems. Its purpose is to provide a spatiotemporal fusion reasoning and lifelong cognitive architecture for learning behavioral evolution in open environments, as well as a lifelong cognitive learning method based on this unified cognitive architecture. In view of the difficulty in extracting information that can be used to adapt to changing environments from past experience in continuous learning problems, a model framework of collaborative scene migration, analogical reasoning, and hierarchical decision-making is proposed, which reduces the large-scale storage requirements of experience information in the learning program and can smoothly and continuously adapt to environmental changes. This research program combines theoretical algorithm design with physical simulation system verification, providing a new comprehensive perspective for constructing lifelong cognitive learning in an open environment. It can form a unified architecture to achieve rapid and effective adaptation to the environment, guide hierarchical machine learning architectures inspired by more complex cognitive structures, and further apply it to typical application areas such as collaboration between robots, self-game action learning, adaptive dynamic scenes, autonomous driving, and intelligent unmanned systems.
[0005] The present invention adopts the following technical solutions:
[0006] The present invention provides a method for spatiotemporal fusion reasoning and lifelong cognitive learning of behavior evolution in an open environment, which is characterized by comprising the following steps:
[0007] In step S1, each agent in the system observes the open environment in real time through its computer vision device. Based on the real-time observation results and the semi-Markov decision model, a cumulative reward action library is obtained. Monte Carlo sampling is performed on the cumulative reward action library to obtain an observation-action history sequence.
[0008] Step S2: At a predetermined time, the agent replays the real-time observation results and the observation-action history sequence to obtain a historical action-observation experience sequence, and generates an n-step joint system-level strategy for the spatiotemporal fusion perspective of the target in the open environment based on the sequence, wherein the historical action-observation experience sequence includes the confidence of the conditional random field of the open environment;
[0009] Step S3, each of the intelligent agents evaluates the confidence distribution level of the conditional random field in real time, and adjusts the degree of freedom of its autonomous learning based on the confidence distribution level;
[0010] Step S4: during the autonomous learning process, each of the intelligent agents extracts action patterns with confidence and reward higher than a predetermined value, maps the action patterns to tasks in the open environment, and constructs a hierarchical dominant joint strategy based on the conditional random field in the joint space of the intelligent agents;
[0011] Step S5, repeating steps S1 to S4, inferring the intrinsic motivation drives between different levels based on the joint system-level strategy and the layer-dominant joint strategy, and constructing the optimal joint strategy of the system in the current open environment.
[0012] The spatiotemporal fusion reasoning and lifelong cognitive learning method for behavior evolution in an open environment provided by the present invention may also have such a technical feature, wherein, in step S1, a semi-Markov decision process is used for modeling, and the six-tuple of the semi-Markov decision model is represented as Among them, S is the environment state space; A is the action space; for……, Generate external system-level strategies based on historical sequence Ω Among them (Ω×A→[0,1])(Ω→[0,1]), the initial set The internal task-level policy π depends on the current state (S×A→[0,1])(S→[0,1]), and β is t=T terminal or Termination conditions when is the state transition probability matrix; is the immediate reward function; γ∈[0,1] is the discount factor.
[0013] The spatiotemporal fusion reasoning and lifelong cognitive learning method for behavior evolution in an open environment provided by the present invention may also have the following technical features: Step S2 includes the following sub-steps:
[0014] Step S2-1, at time t, the local view state s based on probabilistic reasoning o , the distributed joint view state s of its own observation sg , a stable playback perspective state s of a high-level representation of past experience er determining an observable state in the open environment;
[0015] Step S2-2: extract the observation-action history sequence from time 0 to time t, which is recorded as:
[0016]
[0017] Where, It is to extract and integrate the sampling sequence of state i→j in O(s,a);
[0018] Step S2-3, replay the observation-action history sequence from time 0 to time t, and at time c in the replay phase, replay the experience action library combined with the local observation utility function The historical action-observation experience sequence is obtained, which is recorded as:
[0019]
[0020] In the formula, confidence
[0021] Step S2-4, establish a Markov Monte Carlo chain, and perform multiple random walks after it reaches a stable distribution, so as to generate an n-step joint system-level strategy of spatiotemporal fusion perspective around the target in the open environment:
[0022]
[0023] The spatiotemporal fusion reasoning and lifelong cognitive learning method for behavior evolution in an open environment provided by the present invention may also have the following technical features, wherein step S3 includes the following sub-steps:
[0024] Step S3-1, state changes based on real-time observation Combined with the system-level strategy Define the degrees of freedom of the agent's autonomous learning:
[0025]
[0026] Where a i→j are learnable parameters used for evolution;
[0027] Step S3-2: defining the confidence b of the observation-action history sequence based on the state transition probability of the conditional random field, so that the expected reward of each action is not affected by the noise in the open environment. The confidence b is iterated by continuously observing the open environment in real time, matching it with the best responses at different levels, and constructing a non-a priori perfect Bayesian condition.
[0028] Step S3-3, calibrating the confidence interval of the joint system-level strategy using a regression model, the joint system-level strategy The conditional probability distribution of Expressed as:
[0029]
[0030] Where, T L It is the confidence ratio factor obtained by combining the transfer probability distribution of the joint advantage strategy under risk activation. L Reduce unknown risks accordingly The conditional probability distribution of
[0031] The spatiotemporal fusion reasoning and lifelong cognitive learning method for behavior evolution in an open environment provided by the present invention may also have such a technical feature, wherein, in step S3-1, the learning parameter a i→jThe iterative update process is expressed as:
[0032]
[0033] Where, ∑ i The sum of is to traverse all the comparable references in the shared observations, ∑ g The sum of is to filter all possible finite joint system-level strategies in turn.
[0034] The spatiotemporal fusion reasoning and lifelong cognitive learning method for behavior evolution in an open environment provided by the present invention may also have the following technical features, wherein step S4 includes the following sub-steps:
[0035] Step S4-1, design adjustable factors According to the difficulty of distinguishing the utility functions brought by different actions, the corresponding joint system-level strategy is given Add different weights;
[0036] In step S4-2, a stochastic gradient ascent method is used to design an indicator for evaluating the expected return, and the indicator is used to update the parameters of the joint advantage strategy to maximize the probability of the best response action. The gradient of the joint advantage strategy is expressed as:
[0037]
[0038] Where θ is the policy parameter that traverses all action and state spaces;
[0039] Step S4-3, in the joint space (u t |τ t ), constructing a hierarchical leading joint strategy based on the conditional random field distribution observed in real time,
[0040] When the action observation history is not retrievable in the current observation-action history sequence, that is, When , the updated low-level task-level strategy is recorded as:
[0041]
[0042] Where, δ is the
[0043] When t≤T and When , the task is divided into k-order subtasks and their advantage functions are extracted respectively. The action sequence that can obtain the maximum immediate reward is selected to update the high-level task-level strategy, which is recorded as:
[0044] π h (u t |τ t )=argmaxk [A k (s,u k ∣τ t ,u t ),k].
[0045] The spatiotemporal fusion reasoning and lifelong cognitive learning method for behavior evolution in an open environment provided by the present invention may also have the following technical features: wherein, in step S5, the optimal joint strategy obtained by reasoning based on the joint system-level strategy and the hierarchical dominant joint strategy is expressed as:
[0046]
[0047] Where, the first term is the potential state confidence feature obtained by real-time observation of the open environment, which is used to characterize the probability distribution of the response actions of the real-time observation sequence at different levels.
[0048] The second item is to establish a transfer Gaussian kernel function under different risk perception conditions, use cross-level response actions and confidence intervals to integrate the risk coefficient into the joint advantage strategy, fine-tune the reasoning to update the real-time strategy and degrees of freedom, and thus achieve a long-term stable cognitive autonomous learning model.
[0049] Functions and effects of the invention
[0050] Based on the present invention's method for spatiotemporal fusion reasoning and lifelong cognitive learning of behavioral evolution in open environments, a deductive lifelong learning architecture featuring "multi-objective global perception and multi-dimensional decision-making deployment" is constructed, improving the efficiency of intelligent robots' risk exploration and cognition of unknown scenarios. This approach provides a new paradigm, leveraging cross-level optimal response actions and conditional random field confidence intervals to promote the effectiveness of autonomous learning. This framework constructs a deductive reasoning framework corresponding to combinatorial abstraction, thereby defining the degrees of freedom of autonomous learning and the iterative momentum of the system. Through the high rewards of simple subtasks, hierarchical joint action strategies that adapt to the current environment are gradually formed. Furthermore, dynamic simulation and adjustment of challenging environmental tasks are conducted in open systems, and mechanisms for gradually increasing environmental complexity are designed, enabling intelligent agents to develop sustainable lifelong learning capabilities through learning and evolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 Schematic diagram of a method for spatiotemporal fusion reasoning and lifelong cognitive learning of behavioral evolution in an open environment according to an embodiment of the present invention;
[0052] Figure 2 is a flow chart of a method for spatiotemporal fusion reasoning and lifelong cognitive learning of behavior evolution in an open environment according to an embodiment of the present invention;
[0053] Figure 3 Schematic diagram of the structure of the intelligent robot in an embodiment of the present invention. DETAILED DESCRIPTION
[0054] In order to make the technical means, creative features, objectives and effects of the present invention easy to understand, the following is a detailed description of the behavior evolution learning method in an open environment based on a general cognitive architecture of the present invention in combination with embodiments and drawings.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention pertains. The preferred embodiments used herein in the specification of the present invention are merely illustrative and not restrictive, and should not be used to limit the scope of protection of the present invention.
[0056] <Example>
[0057] Figure 3 Schematic diagram of the structure of the intelligent robot in this embodiment.
[0058] like Figure 3 As shown, in this embodiment, the intelligent robot assembly kit under the Jetson Nano system platform is selected as an example application scenario for model training and fine-tuning reasoning. The open environment is a room with multiple obstacles and targets, and the intelligent robot has no prior information about the room.
[0059] This intelligent robot features a stable, low-latency system that supports multiple sensor connections and shares a common Mecanum wheel omnidirectional chassis. The omnidirectional chassis primarily consists of a chassis body, four Mecanum wheels, and a chassis motion controller. It works with an image transmission system (which includes computer vision devices such as radar or cameras) to provide a first-person perspective for autonomous movement control. The omnidirectional chassis is equipped with a flexible, retractable, wrap-around gripper and a gimbal transmitter. The gimbal transmitter accurately, stably, and continuously emits infrared beams, enabling navigation based on these beams. This allows for intelligent obstacle avoidance and environmental perception, while enabling high-performance servos to drive the gripper for grasping or positioning.
[0060] In this embodiment, there are three intelligent robots, each equipped with an autonomous learning module. Their autonomous learning tasks are to learn how to intelligently avoid obstacles, locate, and grasp predetermined targets: a card with a heart-shaped pattern (cardboard), a rectangular cardboard box (regular object), and a plastic water bottle (irregular object). Images of the target objects are fed into the autonomous learning module. These three autonomously learning intelligent robots constitute a lifelong cognitive learning architecture (or system).
[0061] Figure 1 、 Figure 2 They are respectively the principle diagram and flow chart of the spatiotemporal fusion reasoning and lifelong cognitive learning method of behavior evolution in an open environment in this embodiment.
[0062] like Figure 1-2 As shown, the spatiotemporal fusion reasoning and lifelong cognitive learning method for behavior evolution in an open environment of this embodiment specifically includes the following steps:
[0063] Step S1, opportunistic exploration:
[0064] Each intelligent agent in the system observes the open environment in real time through its computer vision device, obtains a cumulative reward action library based on the real-time observation results and the semi-Markov decision model, and performs Monte Carlo sampling on the cumulative reward action library to obtain an observation-action history sequence.
[0065] Specifically, the semi-Markov decision process is used to model the actions and tasks of the agent. The six-tuple of the model can be expressed as Where S is the environment state space; A is the action space; For options, the state will change after multiple interactions with the environment (a series of actions). Generally, this series of actions is regarded as an option. Generate external system-level strategies based on historical sequence Ω Among them (Ω×A→[0,1])(Ω→[0,1]), the initial set The internal task-level policy π depends on the current state (S×A→[0,1])(S→[0,1]), and β is t=T terminal or Termination conditions when represents the state transition probability matrix; represents the immediate reward function; γ∈[0,1] represents the discount factor.
[0066] In real-time observation, based on the images obtained in real time by the computer vision device, the real-time observation O(s,a) is obtained based on the above model and feature extraction algorithm, and the cumulative reward is obtained through the external system-level strategy μ to generate a reward action library, which is then subjected to Monte Carlo sampling to obtain the observation-action history sequence.
[0067] Step S2, memory playback:
[0068] At a predetermined moment, the agent replays the real-time observation results and the observation-action history sequence to form a target perception of the current open environment and generate an n-step joint system-level strategy with a spatiotemporal fusion perspective around the target.
[0069] The target is an observed environmental object that requires action. In this embodiment, the target includes both objects and obstacles in the environment. The target perception process determines whether the target object / obstacle exists at the corresponding location by whether the action is rewarded.
[0070] Step S2 includes the following sub-steps:
[0071] Step S2-1, classify the observable states in the open environment into three categories.
[0072] At any time t, the observable state in the open environment is deterministic knowledge or statistical probability, and no further exploration is required, only the local perspective state s of probabilistic reasoning is needed o ; Distributed joint view state s based on self-observation sg and a stable replay perspective state s of old memories or high-level representations of past experiences that are similar to the new learning content er .
[0073] Step S2-2: extract the observation-action history sequence from time 0 to time t, which is recorded as:
[0074]
[0075] Where, It is to extract and integrate the sampling sequence of state i→j in O(s,a).
[0076] Step S2-3, replay the observation-action history sequence from time 0 to time t. At time c in the replay phase, the occasional replay experience action library is combined with the local observation utility function The historical action-observation experience sequence is obtained, which is recorded as:
[0077]
[0078] In the formula, confidence That is, according to the conditional random field transition probability distribution of the open environment Enter the next state s t+1 (s t+1 ∈S).
[0079] Step S2-4, establish a Markov Monte Carlo chain (MCMC) to obtain the conditional random field transition probability distribution of the next state within n steps After reaching a stationary distribution, a random walk is equivalent to the high-dimensional distribution sampling of Gibbs sampling, thereby generating an n-step joint system-level strategy around the spatiotemporal fusion perspective of the target in the open environment, which is recorded as:
[0080]
[0081] For a strictly ergodic process, an irreducible, non-periodic, and normally recurrent Markov chain, a unique stationary distribution exists, and the limiting distribution of transition probabilities is the stationary distribution of the Markov chain. Under this condition, the n-step non-periodic Markov Monte Carlo transition probability distribution condition is proposed as a necessary condition for the formation of an ergodic random process in a non-stationary environment.
[0082] By establishing a random conditional field transfer probability distribution Make the experience learnable t (s,u t |τ t ) can be combined with the confidence of the cumulative return This generates appropriate responses under non-stationary conditions By obtaining the state transition probability distribution of different levels of dimensions, it can be predicted that a random walk after the Markov Monte Carlo chain reaches stability is equivalent to the high-dimensional distribution sampling of Gibbs sampling.
[0083] Step S3, evaluate the degrees of freedom:
[0084] Each agent evaluates the confidence level of the conditional random field in real time and adjusts its degree of freedom for autonomous learning based on this confidence level. Through real-time strategy, the learning task of exploring a goal in an open environment is broken down into multiple small tasks, each of which is reinforced one by one. Learning tasks are gradually completed through a combination of conditioned reflex actions at different levels.
[0085] Step S3 specifically includes the following sub-steps:
[0086] Step S3-1, defining the degrees of freedom of autonomous learning of the intelligent agent based on the real-time observed state changes and the joint system-level strategy of the system.
[0087] In order to accelerate the adaptation to the complex external environment, the degree of freedom that individuals can learn depends on the observed state changes. Joint system-level strategy for learning and systems Therefore, the degree of freedom of autonomous learning of this mechanism is recorded as:
[0088]
[0089] Among them, a i→j For the learnable parameters used for evolution, given a compact space Any input within, and z i→j =x i W ij +b ij ,in and are the trainable element-wise weight and bias parameters at order k, respectively.
[0090] Accelerable learning parameter a i→j The iterative update process is:
[0091]
[0092] Where, ∑ i The sum of is traversing all the comparable references in the shared observations, and ∑ g This parallel gradient update strategy is to explore the high-dimensional policy gradient direction. Therefore, it has no impact on the time complexity of forward and backward propagation.
[0093] In this embodiment, as described above, the system includes three autonomously learning intelligent robots. The system has an external execution strategy and an internal sequential strategy for each subtask / substep.
[0094] In step S3-2, the confidence b of the observation-action history sequence is defined based on the state transition probability of the conditional random field, so that the expected reward of each action is not affected by the noise in the open environment. This confidence is iterated by continuously observing the open environment in real time, matching the optimal responses at different levels, and constructing a non-a priori perfect Bayesian condition.
[0095] Specifically, the iterative formula of confidence b is:
[0096]
[0097] Among them, Pr can be expressed as:
[0098]
[0099] In step S3-3, the confidence interval of the joint system-level strategy is further calibrated using a regression model for an open environment.
[0100] Space-time fusion strategy The conditional probability distribution of It can be expressed as:
[0101]
[0102] Among them, T L It is the confidence ratio factor obtained by combining the transfer probability distribution of the joint advantage strategy under risk activation. It can be seen intuitively from the above formula that unknown risks are usually highly uncertain. L The spatiotemporal fusion strategy under this unknown risk can be reduced accordingly The conditional probability distribution of Then improve the adaptive action configuration to adapt to the environment.
[0103] Step S4, Combinatorial Abstraction:
[0104] During its autonomous learning process, each agent extracts action patterns with high confidence and high rewards, maps them to specific tasks in an open environment, and completes related learning tasks based on the semi-Markov decision model.
[0105] Step S4 includes the following sub-steps:
[0106] Step S4-1: To improve the adaptability to unknown environments, design adjustable factors According to the difficulty of distinguishing the utility functions brought by different actions, the corresponding joint system-level strategy is given Adding different weights effectively increases the proportion of highly confident joint advantage strategies in random sampling, thereby making the returns of high-level strategies more stable and reliable.
[0107] In this embodiment, the intelligent robot locates objects in the current open environment and locates the real-time distance of the object. It changes the adjustable factor based on real-time observation and establishes autonomous learning cognition, thereby learning to control the opening and closing degree of the robotic arm and gripper to complete the migration task of grasping different objects.
[0108] In step S4-2, a stochastic gradient ascent method is used to design an indicator for evaluating the expected return, and the indicator is used to update the policy parameters to maximize the probability of the best response action.
[0109] Specifically, according to the discount factor assumption d given by Sutton for the ergodic state distribution π (s), then the joint advantage strategy gradient is represented as follows:
[0110]
[0111] Among them, θ is the policy parameter that traverses all action and state spaces, that is, if If it is greater than zero, it means that the direction of change in unknown environmental risks will increase the probability of the current joint advantage strategy in the current state. If the return has marginal increasing benefits, the larger the amplitude of the gradient update, the greater the probability of the current best response action occurring.
[0112] Step S4-3, in the joint space of multiple agents (u t |τ t ), a hierarchical dominant joint strategy is constructed based on the conditional random field distribution of real-time observations.
[0113] Specifically, the strategy when the action observation history is not retrievable in the current real-time observation sequence (i.e., the above observation-action history sequence) is a low-level task-level strategy, i.e., when When the intelligent robot cannot match the real-time observed target to the inherent pattern in the historical experience action library, the updated low-level task-level strategy is recorded as:
[0114]
[0115] Where, δ is the
[0116] When t≤T and When , it means that after satisfying the current risk perception, the intelligent robot cuts the task into k-order subtasks and extracts their advantage functions respectively, and selects the action sequence that can obtain the maximum immediate reward to update the high-level task-level strategy, which is recorded as:
[0117] π h (u t |τ t )=argmax k [A k (s,u k ∣τ t ,u t ),k]
[0118] Step S5: Construct the action evolution of dynamic interaction and fine-tuning reasoning in an open environment:
[0119] Through collaborative interaction with the external open environment, each agent establishes high-level experience and efficient exploration mechanisms from a cognitive perspective (i.e., repeating steps S1 to S4 above). Based on this, they infer the intrinsic motivations at different levels and construct the optimal joint strategy for achieving stable returns in the current environment. In other words, based on the joint system-level strategy obtained in step S2 and the hierarchical dominant joint strategy obtained in step S4, they infer the optimal joint strategy for multiple agents in the system.
[0120] A binary equilibrium potential function is used to integrate the deductive lifelong learning cognitive framework. Based on the countable conditional random field transfer matrix of the Gibbs distribution in step S2-4, the mapping of the high-order dominant strategy (i.e., the updated high-level task-level strategy) under local observation at each time t can be used to obtain the fine-tuned reasoning response of the current hierarchical combination, which can be expressed as:
[0121]
[0122] Among them, the first is the latent state confidence feature obtained through opportunistic exploration of the open environment, which is used to characterize the probability distribution of the observation sequence to the response actions of different levels; and the second is the transfer Gaussian kernel function established under the conditions of different risk perceptions. The risk coefficient is integrated into the joint advantage strategy using cross-level response actions and confidence intervals, and the real-time strategy and degrees of freedom are updated through fine-tuning and reasoning, thereby realizing an autonomous learning model with long-term stable cognition.
[0123] In this embodiment, three intelligent robots perform autonomous lifelong cognitive learning based on the above method. In the initial state, the three intelligent robots are randomly positioned. The intelligent robot at the center walks randomly, and the two intelligent robots on both sides use infrared lasers to detect their relative distance to the intelligent robot at the center in real time. They use their Mecanum wheels to turn and move, and the three intelligent robots form a relative V shape.
[0124] Furthermore, as the three intelligent robots continuously adjusted their positions and orientations, and the other two intelligent robots also joined the random walk, they were able to quickly locate the position of the center point and form a global space-time perspective using the historical action library. Gradually, they spontaneously formed a variety of freeze-frame actions based on the relative positions of the robots at the center, such as forming circle arrays and herringbone arrays.
[0125] In this embodiment, through autonomous learning, the intelligent robot can locate the positions of obstacles and targets in an open environment, design the shortest path to approach the target through its Mecanum wheel, and then control the opening and closing angles and steering angles of its robotic arm and gripper to grasp the above-mentioned target. After testing, it can achieve good positioning and grasping of three types of targets (patterned cardboard sheets, regular objects, and irregular objects).
[0126] Example Function and Effect
[0127] According to the spatiotemporal fusion reasoning and lifelong cognitive learning method for behavioral evolution in an open environment provided by this embodiment, the intelligent agent can efficiently optimize the hierarchical action response and confidence interval based on the risk exploration and cognition of unknown scenes, and configure the transition probability distribution of the joint advantage strategy to achieve the effect of adapting to the environment. Through the process of continuous autonomous learning to adapt to the environment, the multivariate Gaussian kernel is estimated in the gradient direction using the maximum likelihood estimation method, that is, the higher the risk confidence of the adopted combined joint strategy, the overall expected reward under this condition can also obtain stable and high returns in the environment. In this way, the correlation sampling of the current high-level joint advantage strategy combination under the current environmental risk is increased, and the combination configuration is optimized to make it feasible and effective.
[0128] Therefore, the spatiotemporal fusion reasoning and lifelong cognitive architecture for learning behavioral evolution in open environments proposed in this embodiment effectively constructs a deductive lifelong learning architecture with "multi-objective global perception and multi-dimensional decision-making deployment," improving the efficiency of intelligent robots' risk exploration and cognition in unknown scenarios. This embodiment provides a new paradigm that utilizes cross-level best response actions and conditional random field confidence intervals to promote the effectiveness of autonomous learning. It constructs a deductive reasoning framework corresponding to combinatorial abstraction, and uses this to define the degrees of freedom of autonomous learning and the system's iterative momentum. Through high rewards from simple subtasks, a hierarchical joint action strategy that adapts to the current environment is gradually formed. Furthermore, dynamic simulation and adjustment of challenging environmental tasks in the open system are conducted, and a mechanism for gradually increasing environmental complexity is designed, enabling the intelligent agent to develop sustainable lifelong learning capabilities through learning and evolution. The inherent quantification between the reasoning fine-tuning joint strategy and the lifelong cognitive learning paradigm under random sparse reward feedback in open environments is demonstrated, making it influential in the field of artificial intelligence robotics.
[0129] The above embodiments are only used to illustrate specific implementations of the present invention, and the present invention is not limited to the description scope of the above embodiments.
[0130] In the above embodiments, the method of the present invention is applied to the tasks of multiple intelligent robots locating and grasping targets. In fact, the method of the present invention can also be applied to other tasks of other intelligent bodies, such as collaboration between multiple industrial intelligent robots, automatic driving of intelligent cars, etc.
Claims
1. A method for spatiotemporal fusion reasoning and lifelong cognitive learning of action evolution in an open environment, characterized by: The following steps are involved: In step S1, each agent in the system observes the open environment in real time through its computer vision device. Based on the real-time observation results and the semi-Markov decision model, a cumulative reward action library is obtained. Monte Carlo sampling is performed on the cumulative reward action library to obtain an observation-action history sequence. Step S2: At a predetermined time, the agent replays the real-time observation results and the observation-action history sequence to obtain a historical action-observation experience sequence, and generates an n-step joint system-level strategy for the spatiotemporal fusion perspective of the target in the open environment based on the sequence, wherein the historical action-observation experience sequence includes the confidence of the conditional random field of the open environment; Step S3, each of the intelligent agents evaluates the confidence distribution level of the conditional random field in real time, and adjusts the degree of freedom of its autonomous learning based on the confidence distribution level; Step S4: during the autonomous learning process, each of the intelligent agents extracts action patterns with confidence and reward higher than a predetermined value, maps the action patterns to tasks in the open environment, and constructs a hierarchical dominant joint strategy based on the conditional random field in the joint space of the intelligent agents; Step S5, repeating steps S1 to S4, inferring the intrinsic motivation drives between different levels based on the joint system-level strategy and the layer-dominant joint strategy, and constructing the optimal joint strategy of the system in the current open environment.
2. The method for spatiotemporal fusion reasoning and lifelong cognitive learning of action evolution in an open environment according to claim 1, characterized in that: in, In step S1, a semi-Markov decision process is used for modeling, and the six-tuple of the semi-Markov decision model is represented as Among them, S is the environment state space; A is the action space; For options, Generate external system-level strategies based on historical sequence Ω Among them, Ω×A→[0,1], Ω→[0,1], the initial set The internal task-level policy π depends on the current state, S×A→[0,1], S→[0,1], and β is t=T terminal or Termination conditions when is the state transition probability matrix; is the immediate reward function; γ∈[0,1] is the discount factor.
3. The method for spatiotemporal fusion reasoning and lifelong cognitive learning of action evolution in an open environment according to claim 2, characterized in that: in, Step S2 includes the following sub-steps: Step S2-1: At time t, the observable states in the open environment are divided into three categories, namely, local view state s based on probabilistic reasoning and o , the distributed joint view state s of its own observation sg , a stable playback perspective state s of a high-level representation of past experience er ; Step S2-2: extract the observation-action history sequence from time 0 to time t, which is recorded as: Where, It is to extract and integrate the sequence of state i→j in O(s,a); Step S2-3, replay the observation-action history sequence from time 0 to time t, and at time c in the replay phase, replay the observation-action history sequence in combination with the local observation utility function The historical action-observation experience sequence is obtained, which is recorded as: In the formula, confidence Step S2-4: Establish a Markov Monte Carlo chain to obtain the conditional random field transition probability distribution of the next state within n steps After reaching a stationary distribution, multiple random walks are performed to generate an n-step joint system-level strategy with a spatiotemporal fusion perspective around the target in the open environment:
4. The method for spatiotemporal fusion reasoning and lifelong cognitive learning of action evolution in an open environment according to claim 3, characterized in that: in, Step S3 includes the following sub-steps: Step S3-1, state changes based on real-time observation Combined with the system-level strategy Define the degrees of freedom of the agent's autonomous learning: Where a i→j are learnable parameters used for evolution; Step S3-2: defining the confidence b of the observation-action history sequence based on the state transition probability of the conditional random field, so that the expected reward of each action is not affected by the noise in the open environment. The confidence b is iterated by continuously observing the open environment in real time, matching it with the best responses at different levels, and constructing a non-a priori perfect Bayesian condition. Step S3-3, calibrating the confidence interval of the joint system-level strategy using a regression model, the joint system-level strategy The conditional probability distribution of Expressed as: Where, T L It is the confidence ratio factor obtained by combining the transfer probability distribution of the joint advantage strategy under risk activation. L Reduce unknown risks accordingly The conditional probability distribution of 5. The method for spatiotemporal fusion reasoning and lifelong cognitive learning of action evolution in an open environment according to claim 4, characterized in that: in, In step S3-1, the parameter a can be learned i→j The iterative update process is expressed as: Where, ∑ i The sum of is to traverse all the comparable references in the shared observations, ∑ g The sum of is to filter all possible finite joint system-level strategies in turn.
6. The method for spatiotemporal fusion reasoning and lifelong cognitive learning of action evolution in an open environment according to claim 4, characterized in that: in, Step S4 includes the following sub-steps: Step S4-1, design adjustable factors According to the difficulty of distinguishing the utility functions brought by different actions, the corresponding joint system-level strategy is given Add different weights; In step S4-2, a stochastic gradient ascent method is used to design an indicator for evaluating the expected return, and the indicator is used to update the parameters of the joint advantage strategy to maximize the probability of the best response action. The gradient of the joint advantage strategy is expressed as: Where θ is the policy parameter that traverses all action and state spaces; Step S4-3, in the joint space (u t |τ t ), constructing a hierarchical leading joint strategy based on the conditional random field distribution observed in real time, When the action observation history is not retrievable in the current observation-action history sequence, that is, When , the updated low-level task-level strategy is recorded as: Where, δ is the When t≤T and When , the task is divided into k-order subtasks and their advantage functions are extracted respectively. The action sequence that can obtain the maximum immediate reward is selected to update the high-level task-level strategy, which is recorded as: p h (u t |t t )=argmaxk[A k (s,u k |t t ,u t ),k].
7. The method for spatiotemporal fusion reasoning and lifelong cognitive learning of action evolution in an open environment according to claim 6, characterized in that: in, In step S5, the optimal joint strategy is obtained by reasoning based on the joint system-level strategy and the hierarchical dominant joint strategy, and is expressed as: Where, the first term is the potential state confidence feature obtained by real-time observation of the open environment, which is used to characterize the probability distribution of the response actions of the real-time observation sequence at different levels. The second item is to establish a transfer Gaussian kernel function under different risk perception conditions, use cross-level response actions and confidence intervals to integrate the risk coefficient into the joint advantage strategy, fine-tune the reasoning to update the real-time strategy and degrees of freedom, and thus achieve a long-term stable cognitive autonomous learning model.
Citation Information
Patent Citations
Centralized cognitive radio spectrum allocation method based on improved reinforcement learning
CN108809456A
Multi-agent sparse reward environment cooperative exploration method based on internal motivation
CN114169421A