Method and system for determining behavior of artificial intelligence agent on basis of environment-adaptive hierarchical belief state
Patent Information
- Application Number
- PCT/KR2026/004615
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2026-03-19
- Filing Date
- 2026-03-23
- Publication Date
- 2026-09-24
Smart Images

Figure KR2026004615_24092026_PF_FP_ABST
Abstract
Description
Method and System for Action Decision of Environment-Adaptive Hierarchical Belief State-Based Artificial Intelligence Agent
[0001] The present disclosure relates to a method and system for determining the behavior of an artificial intelligence agent, and more specifically, to an artificial intelligence agent technology that generates a hierarchical belief state based on information about the environment, adaptively updates the hierarchical belief state, and determines a behavior for performing a task through simulation of candidate behaviors and evaluation based on information gain.
[0002] With the recent advancement of artificial intelligence technology, active research is being conducted on AI agents (Physical AI) that operate in physical environments, such as robots, autonomous driving, and logistics automation. These agents are required to understand their surroundings based on environmental information acquired from sensors and to determine actions to perform target tasks.
[0003] However, in real-world environments, agents face the challenge of making decisions based on incomplete or uncertain information, as it is difficult to directly observe the overall state of the environment due to limited sensor range, physical obstacles, and dynamic changes. In particular, even for the same task, the location or state of objects and user habits can vary depending on the environment, making it difficult to make effective behavioral decisions relying solely on static rules or predefined knowledge.
[0004] Conventional approaches to address these issues have proposed methods based on reinforcement learning, or methods that directly estimate environmental states and select actions using language or neural network models. However, these methods often rely on policies learned for specific environments, making generalization to new environments difficult, or they have limitations in that they fail to effectively consider the balance between exploration and exploitation due to their inability to explicitly express uncertainty.
[0005] Furthermore, some methods have a problem in that they fail to adequately consider the value of potential information obtainable through action execution or the impact of such actions on task achievement, as they determine a single action based solely on current observations or perform simple score-based selections within a limited set of candidate actions. This leads to repeated unnecessary searches or missed opportunities to acquire important information, resulting in a decrease in overall task execution efficiency.
[0006] For example, when a service robot performs the task of finding a specific object in a home environment, conventional methods may repeat inefficient searches by following a fixed search sequence based on initial assumptions or making immediate decisions based solely on observed information, even though the object's location may vary depending on user habits. In such environments, decision-making is required to balance search and utilization by considering both the value of information obtainable through action execution and the feasibility of task achievement.
[0007] Therefore, a new method and system for behavioral decision-making of an artificial intelligence agent is required that can express beliefs about the environmental state under incomplete environmental information, adaptively update them, and select an optimal action by considering both information acquisition effects and task suitability based on the expected results of performing candidate actions.
[0008] One embodiment of the present disclosure aims to provide a method and system for an artificial intelligence agent to make a decision on behavior that can more accurately reflect environmental conditions even in environments with limited observation by generating a hierarchical belief state composed of a global belief (G), a domain belief (R), and an object belief (S) based on information about the environment, and adaptively updating the belief state based on actual environmental information and a posterior alignment score obtained after the execution of a behavior.
[0009] In addition, one embodiment of the present disclosure aims to provide a method and system capable of efficient decision-making that simultaneously considers uncertainty resolution and prediction reliability by generating a plurality of candidate actions based on the hierarchical belief state, simulating expected environmental information samples expected to be obtained when performing each candidate action, and selecting an action based on an evaluation score obtained by weighted combination of information gain (IG) and pre-alignment scores based on the expected environmental information samples.
[0010] Furthermore, one embodiment of the present disclosure aims to provide a method and system for making behavioral decisions that allows an agent to adapt to various environmental changes and progressively improve performance by determining the appropriateness of a decision based on a post-alignment score calculated after the execution of a behavior, performing active information search including additional information inference, search for external information sources, and re-searching the environment when the decision is determined to be inappropriate, and dynamically modifying and reinforcing inference rules based on the information obtained therefrom.
[0011] However, the technical problems that the various embodiments of the present disclosure aim to solve are not limited to those described above, and there may be other technical problems that can be achieved through the technical means described in this specification.
[0012] One embodiment is,
[0013] A method executed by a computer comprises: obtaining user input instructing a predetermined task through a data input interface; obtaining information about an environment required for performing the task; generating a hierarchical belief state about an environment state for performing the task based on the information about the environment, using at least one artificial intelligence model by at least one processor; generating at least one candidate action based on the hierarchical belief state; and for each of the at least one candidate action, selecting one of the at least one candidate action based on an evaluation that considers expected environment information regarding the performance of the action.
[0014] In another aspect, the method may further include the step of calculating action data corresponding to the selected action; and the step of ingesting the calculated action data to a subsequent processing component.
[0015] In another aspect, the subsequent processing component may be configured to execute the action based on the action data.
[0016] In another aspect, the above hierarchical belief state may include estimated information regarding the probability that an object exists at a specific location within the environment or the state of said object.
[0017] In another aspect, the above-mentioned hierarchical belief state is composed of multiple layers of different levels, wherein the top layer includes comprehensive estimation information regarding the composition of the entire environment and user behavior patterns, the middle layer includes estimation information regarding the possibility of object existence by region within the environment, and the bottom layer may include estimation information regarding the specific location or state of individual objects.
[0018] In another aspect, the step of generating the hierarchical belief state may include the step of generating estimated information regarding the location or state of an object from information about the environment using at least one artificial intelligence model.
[0019] In another aspect, the step of generating at least one candidate action may include: determining a target candidate region associated with the task goal based on estimated information regarding the possibility of existence or location of an object included in the hierarchical belief state; and, for the determined target candidate region, inferring the probability of an object's existence within the target candidate region through at least one artificial intelligence model and generating at least one candidate action corresponding thereto.
[0020] In another aspect, the step of selecting one of the at least one candidate action may include: a simulation step of generating an expected environmental information sample expected to be obtained when performing each candidate action using the at least one artificial intelligence model; a step of calculating an information gain representing the degree of uncertainty reduction of the hierarchical belief state for each of the at least one candidate action based on the expected environmental information sample; and a step of selecting one of the at least one candidate action based on the information gain.
[0021] In another aspect, the simulation step comprises the step of independently generating a plurality of expected environment information samples for each candidate action through the at least one artificial intelligence model, and the step of selecting one action among the at least one candidate action may further include the step of calculating a pre-alignment score based on the degree of interrelationship between the generated plurality of expected environment information samples for each candidate action; and the step of selecting one action among the at least one candidate action based on the information gain and the calculated pre-alignment score.
[0022] In another aspect, the step of selecting one of the at least one candidate action may include: a step of calculating an evaluation score for each of the candidate actions by applying weights to the information gain and the alignment score, respectively; and a step of selecting one of the at least one candidate action based on the evaluation score.
[0023] In another aspect, the step of calculating the evaluation score may include dynamically adjusting the weights according to an uncertainty index calculated based on the entropy or variance of the hierarchical belief state, wherein if the uncertainty index is above a threshold, the weight for the information gain is set relatively high, and if the uncertainty index is below the threshold, the weight for the prior alignment score is set relatively high to calculate the evaluation score.
[0024] In another aspect, the method may further include the steps of: obtaining a sample of actual environmental information based on the execution of the selected action; calculating a post-alignment score based on the correlation between the obtained sample of actual environmental information and the sample of expected environmental information; and determining an update weight to reflect the sample of actual environmental information in the hierarchical belief state based on the calculated post-alignment score, and updating the hierarchical belief state by applying the determined weight.
[0025] In another aspect, the step of updating the hierarchical belief state may include: a step of first updating the existence probability or attribute information for individual objects included in the lowest layer of the hierarchical belief state based on the actual environment information and the update weights; and a step of propagating the updated information of the lowest layer to the upper layer to sequentially update the inference content of the upper layer, including the nature of the area where the object is located or the task-related context.
[0026] In another aspect, the process may include repeating steps from the step of generating candidate behaviors to the step of updating hierarchical belief states, wherein the process may include terminating the repetition step when it is determined that the goal condition cannot be achieved or the environment is identified, such that the hierarchical belief state satisfies the goal condition of a predefined task, or when the state in which the posterior alignment score is above a threshold value persists for a predetermined number of times or more, causing the amount of change in the hierarchical belief state to converge to below the threshold value.
[0027] In another aspect, the method may further include: a step of determining that the decision of the selected action is inappropriate if the calculated posterior alignment score is below a predefined threshold; an active information search step of obtaining auxiliary information by performing an external knowledge source query or additional environment interaction to resolve uncertainty of the hierarchical belief state in the inappropriate decision situation; and a step of dynamically modifying or reinforcing the inference rules that the artificial intelligence model refers to when forming beliefs about the environment or generating the candidate action based on the obtained auxiliary information.
[0028] One embodiment is,
[0029] The system comprises: at least one memory; and at least one processor that reads at least one instruction stored in the at least one memory and performs a method for determining the action of an environment-adaptive hierarchical belief state-based artificial intelligence agent; wherein the at least one instruction includes: a step of obtaining user input directing a predetermined task through a data input interface; a step of obtaining information about an environment required for performing the task; a step in which the at least one processor uses at least one artificial intelligence model to generate a hierarchical belief state about an environment state for performing the task based on the information about the environment; a step of generating at least one candidate action based on the hierarchical belief state; and a step of selecting one of the at least one candidate action based on an evaluation that considers expected environment information regarding the performance of the action for each of the at least one candidate action.
[0030] In another aspect, the system comprises: a plurality of neurons configured in an array including at least one register, at least one programmable logic, and at least one input interface; a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; and at least one routing network that controls the data flow between the plurality of neurons; wherein each of the plurality of neurons is connected to at least one other neuron through the routing network to establish a transmission path for the weights, and may further comprise a Field Programmable Gate Array (FPGA) implementation for a predetermined artificial neural network.
[0031] In another aspect, the system comprises: a plurality of neurons organized into an array including at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; wherein each of the plurality of neurons may further comprise an Application Specific Integrated Circuit (ASIC) for a predetermined artificial neural network that is connected to at least one other neuron through any one of the plurality of synapse circuits.
[0032] A method and system for determining the behavior of an artificial intelligence agent according to one embodiment of the present disclosure can improve the reliability and success rate of task execution by more accurately reflecting the environmental state even in a real environment where observation is limited, by generating a hierarchical belief state composed of a global belief (G), a domain belief (R), and an object belief (S) based on information about the environment, and adaptively updating it based on actual environment information and a posterior alignment score obtained after the execution of a behavior.
[0033] In addition, one embodiment of the present disclosure simulates expected environmental information samples expected to be obtained when performing each candidate action, and selects an action based on an evaluation score obtained by weighting the information gain (IG) calculated through a virtual update and a pre-alignment score based on the correlation between the expected environmental information samples. This enables efficient decision-making that simultaneously considers uncertainty resolution and prediction reliability, and improves overall task execution efficiency by effectively obtaining necessary information while reducing unnecessary searches.
[0034] In addition, one embodiment of the present disclosure can adaptively adjust the balance between search and utilization according to the situation by dynamically adjusting the relative weights of information gain and pre-alignment scores based on the uncertainty index of the belief state, thereby rapidly resolving uncertainty in the early stages of search and prioritizing actions with high predictive reliability in the later stages of search.
[0035] In addition, one embodiment of the present disclosure determines the appropriateness of a decision based on a post-alignment score, and if it is determined to be an inappropriate decision, performs active information search including additional information inference, search for external information sources, and re-search the environment, and by dynamically modifying and reinforcing inference rules based on the information obtained therefrom, the agent can adapt to various environmental changes and gradually improve performance.
[0036] However, the effects obtainable from the present disclosure are not limited to those mentioned above, and there may be other effects that can be clearly understood by a person skilled in the art to which the various embodiments of the present disclosure belong through the configurations described in this specification.
[0037] FIG. 1 illustrates an example of a block diagram of a computing system implementing a method for determining the behavior of an environment-adaptive hierarchical belief state-based artificial intelligence agent according to one embodiment of the present disclosure.
[0038] FIG. 2 briefly illustrates the structure of a neuromorphic circuit that may be included in a processor according to one embodiment.
[0039] FIG. 3 illustrates an example of a block diagram of a computing device implementing an environment-adaptive hierarchical belief state-based artificial intelligence agent action decision method according to one embodiment of the present disclosure.
[0040] FIG. 4 is a block diagram illustrating the internal architecture and data processing pipeline of an artificial intelligence model according to one embodiment of the present disclosure.
[0041] FIG. 5 illustrates an example of a block diagram in another aspect of a computing device implementing an environment-adaptive hierarchical belief state-based artificial intelligence agent action decision method according to one embodiment of the present disclosure.
[0042] FIG. 6 is a block diagram illustrating the data flow and system interaction of the action decision service application process of an environment-adaptive hierarchical belief state-based artificial intelligence agent according to one embodiment of the present disclosure.
[0043] FIG. 7 illustrates the architecture of a Universal Dynamic Multi-Agent System according to one embodiment of the present disclosure.
[0044] FIG. 8 is a conceptual diagram showing an inference loop of a method for determining the behavior of an artificial intelligence agent according to one embodiment of the present disclosure.
[0045] FIG. 9 is a flowchart showing the overall flow of a method for determining the behavior of an artificial intelligence agent according to one embodiment of the present disclosure.
[0046] FIG. 10 is a conceptual diagram of a scenario illustrating a hierarchical belief state-based behavior selection process according to one embodiment of the present disclosure.
[0047] FIG. 11 is a flowchart illustrating the flow of a method for selecting an action based on information gain according to one embodiment of the present disclosure.
[0048] FIG. 12 is a flowchart illustrating the flow of a method for calculating a post-action alignment score and updating a hierarchical belief state after an action execution according to one embodiment of the present disclosure.
[0049] FIG. 13 is a block diagram showing the module configuration of a server computing device according to one embodiment of the present disclosure.
[0050] As the embodiments of the present disclosure are subject to various modifications and may have various embodiments, specific embodiments are illustrated in the drawings and described in detail in the detailed description. The effects and features of the present disclosure and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the drawings. However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various forms. In the following embodiments, terms such as "first," "second," etc., are used not in a limiting sense but for the purpose of distinguishing one component from another. Also, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, terms such as "include" or "have" mean that the features or components described in the specification exist, and do not preclude the possibility that one or more other features or components may be added. Additionally, in the drawings, the size of components may be exaggerated or reduced for convenience of explanation. For example, the size and thickness of each component shown in the drawings are arbitrarily depicted for convenience of explanation, so the embodiments of the present disclosure are not necessarily limited to those depicted.
[0051] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings. When describing with reference to the drawings, identical or corresponding components are given the same reference numerals, and redundant descriptions thereof will be omitted.
[0052]
[0053] - Environment-adaptive hierarchical belief state-based artificial intelligence agent behavior decision system (1000)
[0054] A system (1000) according to one embodiment of the present disclosure can effectively perform various tasks by an artificial intelligence agent operating in an environment where observation is limited, by generating a hierarchical belief state about an environment state using at least one artificial intelligence model and determining an action based thereon.
[0055] In particular, the present system (1000) can perform real-world tasks by updating a hierarchical belief state based on real environment information obtained through a sensor in a Physical AI environment equipped with physical devices such as robot manipulation, autonomous mobile robots, and drone control, determining an action based on the updated belief state, and directly controlling the driving part of the physical device according to the determined action.
[0056] The system (1000) can achieve a high task performance success rate even in a partially observable environment by not simply relying on a pre-learned policy, but by dynamically updating a hierarchical belief state based on environmental information obtained from the actual environment by the agent, and by evaluating the Information Gain and alignment score to select the optimal action.
[0057] When an inappropriate decision is determined during the task execution process, the system (1000) can determine a modified action by autonomously acquiring additional information necessary for task execution through active information search and updating a hierarchical belief state. Through this, the agent can perform an adaptive response in which it identifies the cause and modifies its behavior on its own, even in failure situations caused by unexpected environmental changes or incomplete observational information.
[0058] In addition, the present system (1000) can continuously improve search efficiency in similar situations by dynamically updating information search rules based on new information obtained during the active information search process. The user can assign tasks to the agent using only instructions in the form of natural language, and the agent provides an environment for completing complex tasks step by step through the creation and updating of hierarchical belief states.
[0059] FIG. 1 illustrates an example of a block diagram of a computing system (1000) implementing an environment-adaptive hierarchical belief state-based artificial intelligence agent action decision method according to one embodiment of the present disclosure.
[0060] Referring to FIG. 1, a computing system (1000) for implementing a method for determining the behavior of an artificial intelligence agent in an environment where observation is limited according to one embodiment includes a user computing device (110), a server computing system (130), and a training computing system (150), and each device can communicate through a network (170).
[0061] A method for determining the behavior of an artificial intelligence agent according to one embodiment of the present disclosure may be implemented and provided locally by a user computing device (110), implemented and provided in the form of a web service by a server computing system (130) communicating with the user computing device (110), or implemented and provided by the user computing device (110) and the server computing system (130) in conjunction with each other.
[0062] In this embodiment, the user computing device (110) and / or server computing system (130) may train a machine learning model (120 and / or 140) through interaction with a training computing system (150) that is communicationally connected via a network (170). In the embodiment of the present disclosure, the machine learning model (120, 140) may include an agentic architecture that includes at least one artificial intelligence model beyond a simple prediction model. The artificial intelligence model is responsible for generating a hierarchical belief state from environmental information and selecting an optimal action by evaluating Information Gain and alignment scores. In one embodiment, the artificial intelligence model may include a large language model (LLM) responsible for general reasoning and a small language model (SLM) responsible for task-specific action execution, but is not limited thereto.
[0063] In one embodiment, the machine learning model may be an AI agent having a structure that receives user input and environment information, generates a hierarchical belief state regarding an environment state, simulates expected environment information regarding candidate actions, autonomously determines an optimal action based on information gain and alignment score, and updates the hierarchical belief state according to the result of the action execution.
[0064] In another embodiment, the machine learning model may include an adaptive information search architecture that, when an inappropriate decision is determined, actively searches for relevant information from external information sources and updates hierarchical belief states based on this information to make more accurate behavioral decisions. In this case, the search target is not limited to a fixed query, and an adaptive search method may be adopted that dynamically expands search queries and updates information search rules through active information search in situations of inappropriate decision-making.
[0065] In another embodiment, the machine learning model may be a multi-agent system in which multiple AI agents cooperate to accomplish a specific task. The multi-agent system may have a supervisory pattern in which a central supervisor agent distributes tasks to subordinate specialist agents and aggregates the results. Alternatively, it may have a hierarchical pattern in which a meta-agent acts as an intermediary manager to control and coordinate subordinate agents. It is also possible to include a multi-agent debate pattern in which multiple agents present different action plans and select the optimal action through evaluation.
[0066] In another embodiment, the machine learning model may include a structure combining a Large Language Model (LLM) and a Diffusion Model. The Large Language Model is a model that learns vast amounts of text data to demonstrate excellent capabilities in context understanding, natural language generation, question answering, etc., and may include GPT family, BERT family, T5, PaLM, etc. The Diffusion Model is a probabilistic generative model composed of a forward process that gradually adds noise to the data and a reverse process that gradually removes noise to restore the original data or generate new data, and can generate sophisticated, high-quality samples for various forms of data such as images, videos, voice, and motion trajectories.
[0067] When combining a large language model and a diffusion model, the large language model analyzes user input and contextual information to generate condition embeddings in a format understandable by the diffusion model, while the diffusion model generates target data by incorporating these condition embeddings and performing a noise removal process. Specifically, during the training phase, the diffusion model learns how to remove noise from data with varying levels of added noise, and the large language model learns how to extract meaningful condition information from text into embedding vectors. In the execution (inference) phase, the large language model analyzes user input to generate condition embeddings and passes them to the diffusion model, which then generates high-quality output data that meets the conditions through the noise removal process.
[0068] In one embodiment, the combined structure of a large language model and a diffusion model can be applied to various types of tasks, such as text-to-image generation, text-to-video generation, and multimodal generation. For example, in a text-to-image generation task, the large language model extracts key keywords and detailed conditions from a user's natural language request and converts them into condition embeddings, and the diffusion model generates high-quality images through a noise removal process based on the condition embeddings. Additionally, in a context-aware refinement task that modifies parts of already generated images or videos, the large language model can operate by interpreting the user's modification request to identify the modification target and changes, and the diffusion model can perform partial noise removal and inpainting on the corresponding area. Furthermore, it can be utilized for agent planning-based content generation in large-scale content generation projects, where the large language model establishes the overall scenario and key scene plans, calls the diffusion model for each scene to generate image or video sequences, and if the generation result is unsatisfactory, the large language model directs repainting with modified conditions.
[0069] In this way, the combined structure of a large language model and a diffusion model fuses text-based flexible prompt control with the high-quality generative capabilities of the diffusion model. This allows the large language model to interpret and reinforce linguistic contexts and conditions, thereby inducing the diffusion model to generate detailed and diverse samples, and can be applied to various fields such as art, design, education, and entertainment.
[0070] The server computing system (130) can host AI agents such as those mentioned above, particularly hierarchical belief state-based reasoning requiring complex computations, or a large-scale long-term memory that preserves the agent's behavior history and environmental information for a long period. Additionally, the server computing system (130) includes a Model Context Protocol (MCP) server to manage integration with various external tools, and by relaying communication with external knowledge bases, cloud APIs, search engines, etc., it can efficiently obtain external information necessary for active information search.
[0071] The training computing system (150) can fine-tune the artificial intelligence model by utilizing the agent's success trajectory data and failure trajectory data, or perform iterative learning to gradually improve the agent's performance through a self-reflection mechanism in which the artificial intelligence model evaluates and modifies the results of task execution. Additionally, the training computing system (150) can continuously improve the agent so that it can more effectively update hierarchical belief states and perform tasks in similar situations by dynamically updating information search rules based on new information obtained through active information search.
[0072] The training computing system (150) may be separate from the server computing system (130) or part of the server computing system (130). Additionally, in some embodiments, the training computing system (150) may be separate from the user computing device (110) or part of the user computing device (110).
[0073] The artificial intelligence model can be 1) trained directly locally by a user computing device (110), 2) trained by the server computing system (130) and the user computing device (110) interacting with each other through a network (170), and 3) trained by a separate training computing system (150) using various training and learning techniques. The artificial intelligence model trained by the training computing system (150) can be transmitted to the user computing device (110) and / or the server computing system (130) through the network (170) to be provided and updated.
[0074]
[0075] - User Computing Device (110: User Computing Device)
[0076] The user computing device (110) may include all other types of computing devices such as a smartphone, mobile phone, digital broadcasting device, PDA (personal digital assistants), PMP (portable multimedia player), desktop, wearable device, embedded computing device and / or tablet PC.
[0077] In one embodiment, the user computing device (110) may include a computing device that is mounted on or communicates with a physical device equipped with a driving unit, such as a robot, an autonomous mobile device, or a drone, and may support operation in a Physical AI environment that directly controls the driving unit of the physical device based on an action value determined by an agent.
[0078] In another embodiment, the user computing device (110) may be a user terminal device such as a smartphone, and the server computing system (130) may present actions and action values determined based on hierarchical belief states to the user through a user interface. In this case, the user may confirm, approve, or modify the presented action decision and then direct execution, and may control the execution of the corresponding action only when the user's approval is confirmed. In this way, by combining the user's confirmation step with the agent's autonomous action decision, a Human-Agent Collaboration environment can be provided that prevents the agent from performing unexpected actions and allows the user to directly supervise the progress of the task.
[0079] The user computing device (110) may include at least one processor (111) and memory (112). Here, the processor (111) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions, or a plurality of electrically connected processors.
[0080] In particular, according to the embodiment, the processor (111) may be configured based on a Field Programmable Gate Array (FPGA) implementation and / or an Application Specific Integrated Circuit (ASIC), which is a hardware technology for implementing a predetermined digital circuit.
[0081] Here, a Field Programmable Gate Array (FPGA) can refer to a flexible digital circuit that is programmable according to user needs. A Field Programmable Gate Array implementation may include registers that support the synchronized operation of the FPGA by temporarily storing data and controlling signal flow and timing to maintain intermediate results of operations or state information; programmable logic, which consists of logic circuits configurable according to user needs to program operations within the FPGA to perform specific functions or operations; and input interfaces that serve as channels for receiving data from outside the FPGA, receiving signals from external devices or sensors and transmitting them to internal circuits. Through the combination of these components, a Field Programmable Gate Array implementation can provide flexible and diverse forms of digital circuits.
[0082] An Application-Specific Integrated Circuit (ASIC) may include registers, which are small memory devices that temporarily store and manage data and support rapid processing of ASIC operations by storing intermediate calculation results or state information; a microprocessor, which acts as a central processing unit performing control and computations within the ASIC to coordinate the operation of the entire system by performing various operations or generating control signals when necessary; and an input block, which serves as an interface for receiving data from the outside, receiving data to be processed by the ASIC, transmitting it internally, and receiving various input data through connections with sensors or external devices. Through the combination of these components, the Application-Specific Integrated Circuit can perform specific tasks in an optimized manner. For example, the ASIC may have a structure of a neuromorphic circuit in the form of an array containing multiple neuron circuits, thereby enabling the execution of artificial intelligence computations required for an agent's behavioral decisions with low power consumption and high efficiency.
[0083] FIG. 2 briefly illustrates the structure of a neuromorphic circuit (600) that may be included in a processor (111, 131, 151) according to one embodiment.
[0084] Referring to FIG. 2, for example, a neuromorphic circuit (600) may include a plurality of presynaptic neuron circuits (610), a plurality of presynaptic lines (611) extending laterally from the plurality of presynaptic neuron circuits (610), a plurality of postsynaptic neuron circuits (620), a plurality of postsynaptic lines (621) extending longitudinally from the plurality of postsynaptic neuron circuits (620), and a plurality of synaptic circuits (630) provided at the intersection of the plurality of presynaptic lines (611) and the plurality of postsynaptic lines (621).
[0085] A plurality of free synaptic neuron circuits (610) can transmit signals input from the outside in the form of electrical signals to a plurality of synaptic circuits (630) through a plurality of free synaptic lines (611).
[0086] Additionally, a plurality of post-synaptic neuron circuits (620) can receive electrical signals from a plurality of synaptic circuits (630) through a plurality of post-synaptic lines (621).
[0087] Furthermore, multiple post-synaptic neuron circuits (620) may transmit electrical signals to multiple synaptic circuits (630) through multiple post-synaptic lines (621).
[0088] A plurality of synapse circuits (630) can store weights included in layers constituting a neural network system implemented by a neuromorphic circuit (600) and perform a predetermined operation based on the weights and input data.
[0089] For example, each of the plurality of synaptic circuits (630) may include a resistive memory cell having a variable resistance. In this case, the resistance value of the plurality of synaptic circuits (630) changes by a voltage applied through the plurality of presynaptic neuron circuits (610) or the plurality of postsynaptic neuron circuits (620), and can store weight data according to this resistance change.
[0090] The neuromorphic circuit (600) is formed by mimicking the structure of neurons and synapses, which are essential elements of the human brain. When a deep neural network (DNN) is realized using the neuromorphic circuit (600), the data processing speed can be improved and power consumption can be reduced compared to when the existing von Neumann structure is utilized.
[0091] The memory (112) of the user computing device (110) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory device, magnetic disk, etc., and combinations thereof, and may include web storage of a server that performs memory storage functions on the internet. This memory (112) may store data (113) and instructions (114) necessary for the at least one processor (111) to perform functional operations such as training an artificial intelligence model or performing data filtering through an artificial intelligence model.
[0092] In one embodiment, the user computing device (110) may store at least one machine learning model (120). For example, the user computing device (110) may be various machine learning models, such as multiple neural networks (e.g., deep neural networks) that perform a method of generating and updating hierarchical belief states in an environment where observation is limited and determining actions based on Information Gain and alignment scores, or other types of machine learning models including non-linear models and / or linear models, and may be composed of a combination thereof.
[0093] For example, the machine learning model (120) may be a linear regression, decision tree, random forest, gradient boosting or / and deep learning-based hierarchical belief state-based action decision model, etc. And the neural network may include at least one of feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, Transformer or / and other forms of neural networks.
[0094] Additionally, according to an embodiment, the user computing device (110) may store a model to be used in each process and a prompt template that serves as the basis for input to the model in order to perform at least part of the processes, such as generating hierarchical belief states, simulating candidate behaviors, calculating information gain and alignment scores, and making decisions, through a large language model (LLM) or a small language model (SLM) in an environment where observation is limited.
[0095] In one embodiment, a user computing device (110) receives at least one machine learning model (120) from a server computing system (130) through a network (170), stores it in memory (112), and then executes the stored machine learning model (120) by a processor (111) to perform a method for determining the behavior of an artificial intelligence agent based on hierarchical belief states in an environment where observation is limited.
[0096] In another embodiment, the user computing device (110) can perform operations through a machine learning model (140) including at least one machine learning model (140) in conjunction with a server computing system (130), and can provide the user with an action decision service of an artificial intelligence agent based on a hierarchical belief state in an environment where observation is restricted by communicating related data to the outside.
[0097] For example, a user computing device (110) can provide an action decision service for an artificial intelligence agent based on hierarchical belief states in an environment where observation is restricted, in which a server computing system (130) provides an output for the user's input using a machine learning model (140) via the web.
[0098] Additionally, the artificial intelligence model can be implemented in such a way that at least some of the machine learning models (120 and / or 140) are executed on a user computing device (110) and the rest are executed on a server computing system (130).
[0099] Additionally, the user computing device (110) may include at least one input component (121) that detects user input.
[0100] For example, the user input component (121) may include a touch sensor (e.g., a touch screen and / or a touch pad, etc.) that detects a touch of the user's input medium (e.g., a finger or a stylus), an image sensor that detects the user's motion input, a microphone that detects the user's voice input, a button, a mouse and / or a keyboard, etc.
[0101] Here, the image sensor may include an image processing module. Specifically, the image sensor may process still images or video obtained by an image sensor device (e.g., CMOS or CCD).
[0102] In addition, the image sensor can process a still image or video acquired through the image sensor device using an image recognition process (e.g., OCR, etc.) and / or an image processing module to extract necessary information and transmit the extracted information to a processor.
[0103] Additionally, the input component (121) can receive input from an external controller (e.g., mouse, keyboard, etc.) based on an interface module, and in this case, may include an external output device (e.g., speaker).
[0104] At this time, the interface module may be configured to include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, an earphone port, a power amplifier, an RF circuit, a transceiver, and other communication circuits.
[0105] Additionally, the external output device may include a display system that outputs various information related to an action decision service of an artificial intelligence agent based on a hierarchical belief state in an environment where observation is limited, such as the current hierarchical belief state, information gain and alignment score for each candidate action, selected action, and execution result, as a graphic image.
[0106] Such a display system may be implemented by including at least one of a liquid crystal display (LCD), a thin film transistor-liquid crystal display (TFT LCD), an organic light-emitting diode (OLED), a flexible display, a 3D display, and an e-ink display.
[0107] Meanwhile, the user computing device (110) including the above-described components may further perform at least some of the functional operations performed by the server computing system (130) described later.
[0108]
[0109] -Server Computing System (130: Server Computing System)
[0110] The server computing system (130) can perform a series of processes to provide an action decision service for an artificial intelligence agent based on a hierarchical belief state in an environment where observation is limited.
[0111] Specifically, in an embodiment, the server computing system (130) can provide a behavior decision service for an artificial intelligence agent based on a hierarchical belief state in an environment where observation is limited by exchanging data necessary to drive a behavior decision service process for an artificial intelligence agent based on a hierarchical belief state with an external device such as a user computing device (110).
[0112] More specifically, in an embodiment, the server computing system (130) may provide an environment in which an application can operate to provide an action decision service for an artificial intelligence agent based on a hierarchical belief state on a user computing device (110). To this end, the server computing system (130) may include an application, data and / or commands, etc., for performing functions such as generating and updating hierarchical belief states, generating candidate actions, simulating expected environment information, calculating information gain and alignment scores, making action decisions, calculating action values, active information search, and updating information search rules, and may transmit and receive various data based thereon with the external device.
[0113] A server computing system (130) may include at least one processor (131) and memory (132). Here, the processor (131) of the server computing system (130) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions, or a plurality of electrically connected processors.
[0114] For example, ASICs may have a structure of a neuromorphic circuit in the form of an array containing multiple neuron circuits (see Fig. 2).
[0115] The memory (132) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory (132) may store data (133) and instructions (134) necessary for the processor (131) to perform functional operations such as training at least one artificial intelligence model, or executing methods for determining the behavior of an artificial intelligence agent, such as generating and updating hierarchical belief states in an environment where observation is limited, generating candidate behaviors, simulating expected environment information, calculating information gain and alignment scores, determining behavior, active information search and updating information search rules.
[0116] In one embodiment, the server computing system (130) may be implemented to include at least one computing device. For example, the server computing system (130) may be implemented to operate a plurality of computing devices according to a sequential computing architecture, a parallel computing architecture, or a combination thereof. Additionally, the server computing system (130) may include a plurality of computing devices connected to a network (170).
[0117] Additionally, the server computing system (130) may store at least one machine learning model (140). For example, the server computing system (130) may include at least one artificial intelligence model as a machine learning model (140) that performs the generation and updating of hierarchical belief states, simulation of expected environment information, calculation of information gain and alignment scores, and action decisions, and may include a neural network and / or other multi-layer non-linear model. Exemplary neural networks may include a feed-forward neural network, a deep neural network, a recurrent neural network, and a convolutional neural network.
[0118] In an embodiment, the server computing system (130) may further include a data store computing system (hereinafter, data store) which is a repository for continuously storing and managing raw data such as environmental information, hierarchical belief state data, success trajectory data, failure trajectory data, and information search rules that form the basis of a behavior decision service.
[0119] These data stores may include various forms of data storage, ranging from file systems to cloud storage. For example, a data store may include a relational database that uses a structured query language (SQL) to define and manipulate data, a NoSQL database designed for flexibility and scalability to process unstructured and semi-structured data, a data warehouse optimized for querying and analysis by centralizing large volumes of data from multiple sources, a data lake that stores structured data, semi-structured data, and unstructured data, and at least one database among local storage devices or Network Attached Storage (NAS).
[0120]
[0121] - Training Computing System (150: Training Computing System)
[0122] The training computing system (150) may include at least one processor (151) and memory (152). Here, the processor (151) of the training computing system (150) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions, or a plurality of electrically connected processors. For example, ASICs may have the structure of a neuromorphic circuit in the form of an array including a plurality of neuron circuits (see FIG. 2).
[0123] The memory (152) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory (152) may store data (153) and instructions (154) necessary for a processor (151) to perform at least one artificial intelligence model training, parameter updates of a hierarchical belief state-based action decision model, information search rule updates, etc.
[0124] For example, the training computing system (150) may include a model trainer (160) that trains a machine learning model (120 and / or 140) stored in a user computing device (110) and / or a server computing system (130) using various training or learning techniques, such as backpropagation of error. The model trainer (160) may perform backpropagation updates to one or more parameters of the machine learning model (120 and / or 140) for making behavior decisions of an artificial intelligence agent based on hierarchical belief states in an environment where observations are limited, based on a defined loss function. In some embodiments, performing backpropagation of error may include performing truncated backpropagation through time. The model trainer (160) may perform a number of generalization techniques (e.g., weight devaluation, dropout, and / or knowledge distillation, etc.) to improve the generalization ability of the machine learning model (120 and / or 140) being trained.
[0125] Additionally, the model trainer (160) can train a machine learning model (120 and / or 140) based on a series of training data (161). Here, the training data (161) may include agent success trajectory data, failure trajectory data, environmental information, hierarchical belief state data, history of calculating information gain and alignment scores, additional information obtained through active information search, and updated information search rules, and may include data of different forms such as images, sensor data, text, etc. Such training data (161) may be provided by a user computing device (110) and / or a server computing system (130). When the training computing system (150) trains the machine learning model (120 and / or 140) on specific environmental data of the user computing device (110), the machine learning model (120 and / or 140) may be characterized as a hierarchical belief state-based personalized behavior decision model specialized for that environment.
[0126] The model trainer (160) includes computer logic utilized to provide desired functions and may be implemented as hardware, firmware and / or software controlling a general-purpose processor. In one embodiment, the model trainer (160) includes a program file stored in a storage device, loaded into memory (152), and may be executed by one or more processors (151). In another embodiment, the model trainer (160) includes one or more sets of computer-executable data (153) and instructions (154) stored in a tangible computer-readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.
[0127] In this system (1000), a user computing device (110) and a server computing system (130) may be connected via a wired / wireless network (170) for communication. The network (170) includes, but is not limited to, a 3GPP (3rd Generation Partnership Project) network, an LTE (Long Term Evolution) network, a WIMAX (World Interoperability for Microwave Access) network, the Internet, a LAN (Local Area Network), a Wireless LAN (Wireless Local Area Network), a WAN (Wide Area Network), a PAN (Personal Area Network), a Bluetooth network, a satellite broadcasting network, an analog broadcasting network and / or a DMB (Digital Multimedia Broadcasting) network. Generally, communication through the network (170) can be performed using any type of wired and / or wireless connection through various communication protocols (e.g., TCP / IP, HTTP, SMTP and / or FTP, etc.), encodings or formats (e.g., HTML and / or XML, etc.), and / or protection schemes (e.g., VPN, Secure HTTP and / or SSL, etc.).
[0128] The server computing system (130) may further include a plurality of logically and physically separated specialized engines and repositories to perform a method of determining the behavior of an artificial intelligence agent based on a hierarchical belief state in an environment where observation is limited. In one embodiment, the server computing system (130) may include at least one engine among a reasoning engine that processes user input and environment information to generate and update a hierarchical belief state and determines a behavior based on information gain and alignment score, a search engine that performs active information search, and a rule management engine that manages information search rules. Here, the term "engine" may include not only a set of instructions executed by a processor (131) to perform specific logic, but also dedicated hardware circuits to accelerate said logic.
[0129] Such an engine may run on hardware accelerators optimized to handle the computational load of at least one artificial intelligence model. The hardware accelerator is a processor specialized in matrix operations and vector processing and may include at least one of a Tensor Processing Unit (TPU), a Graphics Processing Unit (GPU), a Field-Programmable Gate Array (FPGA), or an Application-Specific Integrated Circuit (ASIC). Such hardware accelerators can provide technical improvements that distribute the computational load of the artificial intelligence model, such as generating and updating hierarchical belief states, simulating expected environment information, and calculating information gain and alignment scores, and enable real-time action decisions.
[0130] Additionally, the data (133) may include a structured embedding repository to support active information search rather than a simple set of data. The embedding repository stores environmental information, hierarchical belief state data, success trajectory data, failure trajectory data, and high-dimensional vector representations of information search rules, thereby enabling the action decision engine to perform high-speed search based on semantic similarity. This can serve as a technical means to suppress hallucination in the artificial intelligence model and increase the accuracy of action decisions.
[0131] Specifically, the memory (132) of the server computing system (130) may include a structured Knowledge Base Layer to physically support active information retrieval. The Knowledge Base Layer may include an Embedding Repository that stores high-dimensional vector representations of environmental information and knowledge related to task execution, and a Policy Document Repository that stores information retrieval rules and behavior execution constraints. In this case, the behavior decision engine vectorizes environmental information received from the environment or results of inappropriate decision judgments and queries the Embedding Repository to search for context information with high semantic similarity in real time, thereby allowing the agent to technically suppress hallucination phenomena by referring to external verified knowledge rather than relying solely on intrinsic parameters.
[0132] Additionally, the user computing device (110) may include a trigger event detector that provides an interface for interaction with an agent and initiates the operation of the agent. The trigger event detector can detect not only user input but also environmental changes obtained from sensors, reception of results of inappropriate decision-making, etc., and transmit a processing request to a server computing system (130).
[0133] Additionally, the model trainer (160) of the training computing system (150) may include a Supervision Signal Engine. The Supervision Signal Engine may calculate a difference value by comparing the execution result of an action performed by an agent with additional information obtained through active information search, and execute a reinforcement learning process to update a Reward Model or fine-tune at least one artificial intelligence model based on this. In one embodiment, the Supervision Signal Engine may increase the learning efficiency of the artificial intelligence model by utilizing the history of calculating information gain and alignment scores and the results of updating hierarchical belief states as additional supervisory signals.
[0134] As such, in one embodiment, the system (1000) of the present disclosure may be implemented as a technical system in which a specialized hardware accelerator that performs hierarchical belief state-based reasoning, active information retrieval, and information retrieval rule updates, a vectorized data storage, and a physical engine that controls the same are organically combined, rather than being a simple set of software algorithms.
[0135] FIG. 3 illustrates an example of a block diagram of a computing device (100) implementing an environment-adaptive hierarchical belief state-based artificial intelligence agent action decision method according to one embodiment of the present disclosure.
[0136] Referring to FIG. 3, the computing device (100) included in the user computing device (110), server computing system (130), and training computing system (150) may include a plurality of applications (e.g., Application 1 to Application N). Each application may include a machine learning library and one or more machine learning models. For example, the applications may include an image processing application (e.g., Detection, Classification and / or Segmentation, etc.), a sensor data processing application, an environment information collection application, a natural language processing application, a robot control application and / or an agent action decision application, etc.
[0137] Additionally, for example, the computing device (100) may include a model that performs a method for determining the behavior of an artificial intelligence agent based on a hierarchical belief state in an environment where observation is limited, and may include an application that provides related services to the user, namely, a specialized application for generating and updating hierarchical belief states, generating candidate behaviors, simulating expected environment information, calculating information gain and alignment scores, determining behavior, calculating action values, and active information search.
[0138] In an embodiment, the computing device (100) may include a model trainer (160) for training at least one artificial intelligence model, and by storing and operating the trained artificial intelligence model, it may provide hierarchical belief states, behavior decisions, and action values based on user input and environment information as output data.
[0139] Each application of the computing device (100) may communicate with a number of other components of the computing device (100), such as, for example, at least one sensor, a context manager, a device state component, and / or additional components. In one embodiment, each application may communicate with each device component using an API (e.g., a public API). For example, an environment information collection application may communicate with a sensor component to obtain raw data, and an action decision application may communicate with the inference engine of the server computing system (130) via an API to receive hierarchical belief states, action decision results, and action values. In one embodiment, the API used by each application may be specific to that application.
[0140] FIG. 4 is a block diagram illustrating the internal architecture and data processing pipeline of an artificial intelligence model according to one embodiment of the present disclosure.
[0141] Referring to FIG. 4, a computing device may have a pipeline structure that receives input data (202) and generates output data (212) through a series of transformation processes to perform an environment-adaptive hierarchical belief state-based artificial intelligence agent action decision method. This process is performed through a preprocessing module (204), an encoder / embedding model (206), a neural network layer (208), and a decoder / generation head (210).
[0142] First, the preprocessing module (204) receives input data (202) (e.g., text prompt, image, or multimodal signal) from a user or system. The preprocessing module (204) performs tokenization and normalization on the input data to generate a sequence of tokens, which are the smallest units that the model can process.
[0143] Next, the encoder / embedding model (206) receives the generated token as input and converts it into a vector / embedding mapped to a number in a high-dimensional vector space. At this stage, the discrete information of the input data is converted into a continuous numeric matrix, making it a form that can be computed by the machine learning model.
[0144] Next, the neural network layer (208) receives the vector / embedding and performs deep computation. The neural network layer (208) may have a structure in which a plurality of sub-layers (e.g., Layer 1 to Layer N) are stacked. Each layer abstracts and refines input features through an attention mechanism or convolution operation, etc.
[0145] In particular, the final output of the neural network layer (208) is defined as a latent representation. This latent representation has a structure different from the original input data (202) and may correspond to an intermediate representation in which the semantic features of the data are highly compressed and abstracted. This implies that it is not a simple transmission of data, but a technical data structure that is valid only within the system.
[0146] Finally, the decoder / generation head (210) receives the potential representation as a conditioning input. The decoder / generation head (210) interprets the compressed potential representation and reconstructs or generates output data (212) in a form recognizable by the user (e.g., natural language text, image pixels, control codes, etc.).
[0147] This stepwise data transformation structure (token → vector → latent representation → output) can clearly demonstrate that it functions not as a simple sequence of operations, but as a concrete device that technically processes input data to generate useful information.
[0148] FIG. 5 illustrates an example of a block diagram in another aspect of a computing device (200) implementing an environment-adaptive hierarchical belief state-based artificial intelligence agent action decision method according to one embodiment of the present disclosure.
[0149] Referring to FIG. 5, the computing device (200) includes a plurality of applications (e.g., Application 1 to Application N). Each application can communicate with a central intelligence layer. For example, applications may include an image processing application, a text messaging application, an email application, a dictation application, a virtual keyboard application and / or a browser application. In one embodiment, each application can communicate with the central intelligence layer (and a model stored therein) using an API (e.g., a common API across all applications).
[0150] Additionally, in one embodiment of the present disclosure, the application may include an application for determining the behavior of an artificial intelligence agent based on hierarchical belief states in an environment where observation is limited, an application for generating and updating hierarchical belief states, an application for active information retrieval, an energy management application, a logging and analysis application, etc.
[0151] The central intelligence layer may include a number of machine learning models. For example, as illustrated in FIG. 5, at least some of the machine learning models may be provided for each application and managed by the central intelligence layer. In another embodiment, two or more applications may share a single machine learning model. For example, in some embodiment, the central intelligence layer may provide a single model for all applications. In some embodiment, the central intelligence layer may be included within the operating system of the computing device (200) or otherwise implemented.
[0152] In one embodiment, the central intelligence layer may be integrated as part of the operating system or implemented as a separate logical layer, and may perform the role of transmitting input time series data to the corresponding model to return a prediction result.
[0153] The central intelligence layer can communicate with the central device data layer. The central device data layer may be a centralized data store for the computing device (200).
[0154] For example, the central device data layer can integrate and store sensor data, device state information, external environment information, etc., stored within the computing device (200), and provide this as input data required for an artificial intelligence agent's action decision service based on hierarchical belief states in an environment with limited observation, such as generating and updating hierarchical belief states, generating candidate actions, simulating expected environment information, calculating information gain and alignment scores, and making action decisions. Each device component (e.g., sensor, state manager, etc.) can communicate with the corresponding data layer through a private API, etc.
[0155] As illustrated in FIG. 5, the central device data layer can communicate with a number of other components of the computing device (200), such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0156] The technology described herein may refer to servers, databases, software applications, and other computer-based systems, as well as actions taken and information transmitted to or from said systems. It will be recognized that the inherent flexibility of computer-based systems allows for a wide range of possible configurations, combinations, division of tasks, and functionality between and from components. For example, the processes described herein may be implemented using a single device or component or multiple devices or components operating in combination. Databases and applications may be implemented in a single system or in a distributed system across multiple systems. Distributed components may operate sequentially or in parallel.
[0157] FIG. 6 is a block diagram illustrating the data flow and system interaction of the action decision service application process of an environment-adaptive hierarchical belief state-based artificial intelligence agent according to one embodiment of the present disclosure.
[0158] Referring to FIG. 6, the computing system may be configured as an organic data pipeline between a trigger event detector (310), an inference engine (330), and a mobile device screen (350).
[0159] First, the trigger event detector (310) is configured to monitor and detect a signal initiating the operation of the system. The trigger event detector (310) receives at least one of i) a time event indicating the arrival of a specific point in time, ii) a user action resulting from a user's physical input, or iii) a system state indicating a change in internal system data, and transmits an activation signal to the inference engine (330). This means that the service can be actively initiated depending on the situation without an explicit request from the user.
[0160] The reasoning engine (330) performs a multi-stage operation process that transforms raw data into a final result in response to the activation signal. This process is implemented as a series of logically connected prompt chains.
[0161] For example, the first prompt, contextualization (332), allows the inference engine (330) to receive unstructured raw data (e.g., user logs, channel metadata), analyze and refine it, and generate compressed summary information.
[0162] And the second prompt, Content Generation (334), can generate multiple candidate results that match the user's intent or situation using a generative AI model based on the summary information generated above.
[0163] In addition, the third prompt, Verification (336), performs filtering and verification by comparing the generated candidate results with a predefined policy (e.g., safety guidelines, format rules) to derive a reliable final result.
[0164] Finally, the final result generated by the inference engine (330) can be transmitted to and implemented on the mobile device screen (350) through an auto pre-filling (354) operation. Specifically, the system changes the state of the interface by directly writing the final result to a memory address of a specific target field (352) (e.g., text input field, setting value) within the mobile device screen (350). Here, the mobile device screen (350) may be an example of a user computing device (110).
[0165] Such a configuration can go beyond the simple display of information and provide specific technical means for data generated by an external trigger to physically control and complete the input interface of the user terminal.
[0166] FIG. 7 illustrates the architecture of a multi-agent system (Universal Dynamic Multi-Agent System, 400) according to one embodiment of the present disclosure.
[0167] Referring to FIG. 7, the system (400) may be configured around an orchestration engine (410) that determines and executes an optimal agent collaboration structure in real time according to the nature of the user's request or task. The system (400) may operate in an organically combined manner, including a complexity analyzer (405), an agent pool (420), a shared memory fabric (430), and a tool execution interface (440).
[0168] 1. Input Analysis and Dynamic Topology
[0169] The complexity analyzer (405) evaluates the complexity of the user query and the required domain expertise. Based on this evaluation result, the orchestration engine (410) dynamically configures an optimal agent collaboration topology for task resolution.
[0170] For example, in the case of a simple question, the engine (410) activates a single agent mode.
[0171] If a complex plan is required, the engine (410) configures a Supervisor pattern and instantiates one supervisor agent to control subordinate agents.
[0172] When high accuracy is required, the engine (410) configures a debate pattern and sets a control path so that multiple agents perform mutual criticism.
[0173] 2. Agent Pool and Instantiation (Agent Pool)
[0174] The agent pool (420) is a repository of template agents that have prompts and tool sets specialized for specific functions (e.g., web search, code generation, data analysis). The orchestration engine (410) can select and activate the necessary agents at runtime according to the determined topology.
[0175] 3. Shared Memory Fabric
[0176] The shared memory fabric (430) is a data pipeline that synchronizes the state and context between multiple collaborating agents. By mediating the short-term memory of individual agents and the knowledge base of the entire system, it can ensure that the output of Agent A is transferred to the input of Agent B without loss.
[0177] 4. Tool Execution & Feedback Loop
[0178] Each agent communicates with an external API (search engine, calculator, AWS, etc.) through a tool execution interface (440). The execution result generated at this time is fed back to the orchestration engine (410), and the engine can determine the consistency of the result and perform self-correction logic to instruct the agent to rework or proceed to the next step.
[0179] This structure enables the implementation of an adaptive artificial intelligence system that flexibly changes the system's processing structure according to the nature of the input problem, rather than a fixed (static) algorithm.
[0180] Meanwhile, the orchestration engine (410) analyzes the characteristics of the task received from the complexity analyzer (405) (e.g., creativity, logic, whether coding is required) and selects specialized agents waiting in the agent pool (420) to form a dynamic collaboration topology. The topology can be reconfigured into various operation modes as follows depending on the type of task.
[0181] First Operation Mode: Sequential Pattern
[0182] For tasks such as a single-flow creation or report writing, the orchestration engine (410) connects multiple agents in series. For example, it forms a pipeline in the order of user agent -> writer agent -> style agent to control the output of the previous stage to be passed as input to the next stage.
[0183] Second Operation Mode: Supervisor Pattern
[0184] In cases where complex sub-tasks are mixed, the engine (410) forms a centralized star topology in which one agent is designated as a supervisor and the remaining agents (e.g., research agent, math agent) are assigned as workers. The supervisor agent distributes tasks to sub-agents and aggregates the results.
[0185] Third Operation Mode: Hierarchical Pattern
[0186] In cases of high complexity, such as large-scale project management, the engine (410) establishes a command system by forming a tree structure topology in which a meta-agent is placed as the upper manager and multiple specialized agent groups are placed below it.
[0187] 4th Operation Mode: Debate Pattern
[0188] In cases where the correct answer is unclear or high reliability is required (e.g., social issue analysis), the engine (410) constructs a competitive topology in which multiple agents present different perspectives on the same topic and perform mutual criticism and voting to derive the optimal conclusion.
[0189] Fifth Operation Mode: Mixture-of-Agents Pattern
[0190] When parallel processing is required, the engine (410) forms a parallel processing structure that places multiple agents in multiple layers to perform tasks simultaneously and finally integrates the results through an aggregator agent.
[0191] Additionally, the agent pool (420) includes various template agents that can be deployed into the topology. Each agent is instantiated to perform the following core methodologies under the control of the orchestration engine (410).
[0192] For example, at least one agent can execute a loop that repeats reasoning and action to perform reasoning and action (ReAct Paradigm) that refines answers based on results from external tools (e.g., Google Search).
[0193] In addition, at least one agent can perform code-based behavior (CodeAct Paradigm) that involves complex calculations or data analysis by generating and executing executable code (e.g., Python) instead of natural language.
[0194] In addition, at least one agent can perform self-reflection to improve the quality of the output by carrying out a metacognitive process of self-evaluating (Critique) and modifying the generated output.
[0195] Furthermore, the system provides a shared memory fabric (430) to enable individual agents to cooperate organically without disconnection. This prevents context loss by managing the conversation history (Short-Term Memory) between agents and an external knowledge base in an integrated manner. Additionally, the tool execution interface (440) securely connects the agents with external APIs, cloud services, and databases through protocols such as Multi-Channel Processing (MCP).
[0196] Thus, the system (400) of the present disclosure implements an adaptive artificial intelligence platform that is not a fixed single model, but rather dynamically changes the system's processing structure and behavior according to the nature of the input problem.
[0197] Specifically, one embodiment of the present disclosure may be implemented as an Agentic RAG (Retrieval-Augmented Generation) system that supports advanced question-answering by including agent functions. This system is a search-based generation system that retrieves information from external data sources and generates answers based thereon.
[0198] The system's data processing pipeline can consist of a data extraction phase and an Agentic RAG pipeline phase. In the data extraction phase, content in various formats, such as text and images, is collected from designated websites. Subsequently, text and metadata are extracted from the collected content, the text is chunked into small units, and each chunk is vectorized using an embedding model and stored in a vector database.
[0199] The Agentic RAG pipeline can be divided into search and generation phases. In the search phase, when a user query is input, query rewriting and embedding are performed, and highly relevant documents are identified through similarity searches in a vector database. Subsequently, the retrieved results can be ranked based on relevance to construct context. In the generation phase, the user query and the retrieved context are combined to construct a prompt, which is then input into a large language model to generate the final response. During this process, the agent can respond to complex queries by utilizing agentic elements such as memory storage, calling external tools, and planning.
[0200] An AI agent according to one embodiment of the present disclosure can support complex decision-making and task execution through a hierarchical memory structure similar to human memory. The agent's memory can be broadly divided into short-term memory and long-term memory.
[0201] Short-term memory is a temporary memory space focused on the currently ongoing workflow and may include working memory, which manages workflow-specific reasoning and task flows, and cache memory, which provides quick access to frequently used data or results.
[0202] Long-term memory is a memory space based on knowledge and experience that is continuously preserved, and may include episodic memory, which records events or incidents manually saved in a specific workflow; semantic memory, which stores conceptual or factual knowledge; and procedural memory, which stores knowledge of how to perform specific tasks or procedural knowledge.
[0203] This memory structure can be integrated with the language model framework through a central memory controller. Additionally, it can leverage external knowledge or support real-time integration by connecting with external vector databases, semantic databases, or third-party APIs via the MCP server. Through this, the agent can generate user-customized responses that comprehensively reflect past experiences and current context.
[0204] One embodiment of the present disclosure may include various agentic workflows to solve complex problems. These workflows may be designed to suit specific business purposes.
[0205] For example, the system may include a Plan and Execute workflow. In this workflow, a Planner breaks down a single top-level task into multiple sub-tasks, specialized agents process each sub-task, and then integrates the results. This can be utilized for business process automation or data pipeline orchestration.
[0206] As another example, the system may include an Orchestrator-Worker workflow. In this structure, a central orchestrator language model breaks down tasks, distributes them to multiple worker language models for processing, and then integrates the results. This can be used in the implementation of Agentic RAGs or coding agents.
[0207] As another example, the system may include a routing workflow. This workflow is structured to analyze input tasks, classify them into the most suitable ones among several predefined subtasks, and forward them to a specialized language model or path capable of handling the task. This can be applied to customer support agents or multi-agent discussion systems.
[0208] One embodiment of the present disclosure may include a protocol for efficient and secure communication between a plurality of agents or between an agent and an external tool.
[0209] In one embodiment, an Agent2Agent (A2A) protocol may be used for communication between agents. The A2A protocol can enhance security by enabling each agent to communicate without directly sharing their internal data. Through this protocol, multiple agents can share tasks and negotiate, and each agent can operate independently using its own language model, framework, and database.
[0210] In another embodiment, the Multi-Channel Processing (MCP) protocol may be used for communication between an agent and an external function server. MCP has a structure that separates each external function, such as file access, search, and cloud API calls, into a separate server for communication. For example, one agent may communicate with a local file system or a search engine via the MCP protocol, while another agent may communicate with a cloud provider such as AWS or a communication tool such as Slack via the same protocol.
[0211]
[0212] Hereinafter, a method for determining the behavior of an artificial intelligence agent based on a hierarchical belief state in an environment where observation is limited, using at least one artificial intelligence model, is described in detail with reference to FIGS. 8 to 13.
[0213] In one embodiment, the method described below may be performed by at least one processor included in the server computing device (500). However, it is not limited thereto, and at least a part of the method may be performed by a processor provided in an agent, and another part may be performed by a processor of the server computing device (500).
[0214] FIG. 8 is a conceptual diagram showing an inference loop of a method for determining the behavior of an artificial intelligence agent according to one embodiment of the present disclosure.
[0215] Referring to FIG. 8, the action decision method of an artificial intelligence agent according to the present disclosure operates in a circular loop structure leading to an action, a utility estimate, a next action, and a reward, and adopts an iterative reasoning mechanism in which a belief update is performed based on the reward obtained after each action is executed, and the updated belief state is reflected in the next action decision.
[0216] Specifically, the agent can generate candidate actions based on the current hierarchical belief state and select the optimal action by estimating the utility for each candidate action. When the selected action is executed, a reward is received from the actual environment, which updates the hierarchical belief state and serves as the basis for the next action decision. By repeating this cycle of action decision, execution, and belief update, the agent can effectively perform tasks while progressively improving its understanding of the environment state, even in an environment with limited observation.
[0217] In one embodiment, the reward in FIG. 8 refers to actual environment information obtained from the environment after the execution of an action, and may include a posterior alignment score representing the degree of alignment between the actual environment information and the sample of expected environment information simulated by the agent before the execution of the action. The actual environment information is used as the primary input for updating the hierarchical belief state, and the posterior alignment score can be used to determine the update weights that reflect the actual environment information in the belief state. Additionally, a lower posterior alignment score acts as a trigger to determine that the agent's decision was inappropriate, thereby inducing active information search and further updating of the hierarchical belief state.
[0218] FIG. 9 is a flowchart showing the overall flow of a method for determining the behavior of an artificial intelligence agent according to one embodiment of the present disclosure.
[0219] First, at least one processor can obtain user input that directs a predetermined task through a data input interface (S101).
[0220] For example, referring to FIG. 10, the user can input a task to be performed by the agent in natural language, such as "bring me a banana." The data input interface may support various forms of input, such as text, voice, and images. In one embodiment, the acquired user input may be stored in at least one memory, but is not limited thereto, and may be processed by streaming processing or pipeline processing, or converted into other forms such as vector embeddings and utilized. The acquired user input may subsequently be used to generate hierarchical belief states and candidate actions.
[0221] Next, at least one processor can obtain information about the environment required for task execution (S103).
[0222] Information about the environment can be obtained in various ways. In one embodiment, the environment information can be extracted from raw data obtained through at least one sensor, such as a visual sensor or a distance measuring sensor, equipped in an agent. In this case, the server computing device (500) can control a physical device equipped with a sensor to scan the surrounding environment and identify at least one of the type, location, and state of an object from the obtained raw data to use as environment information.
[0223] In other embodiments, environmental information is not limited to sensor data and can be obtained from various sources, such as environmental maps or 3D models stored in a database (DB), state information from a physical simulation environment, real-time environmental data obtained through an external API, or a pre-built knowledge graph. For example, in an autonomous driving environment, environmental information can be constructed by combining a High Definition Map (HD Map) and GPS data, and in a smart factory environment, the current state information of the process line can be obtained as environmental information from the DB of a process control system.
[0224] In any case, only information recognizable within the agent's current observation range is obtained. For example, referring to FIG. 10, when the agent observes the kitchen, environmental information containing only partial information of the entire environment can be obtained, such as "Cabinet 1 in the kitchen, refrigerator confirmed, shelf in the living room confirmed, interior of each storage space unobserved." As such, since the environmental information includes only information recognizable within the agent's current observation range or the range of an accessible data source, it reflects the characteristics of a partially observable environment.
[0225] Next, at least one processor can use at least one artificial intelligence model to generate a hierarchical belief state for an environment state for task execution based on information about the acquired environment (S105).
[0226] Hierarchical belief states represent the environmental state currently perceived by an agent in a structured form; they can include not only directly observed information but also information estimated from environmental states outside the observation range by utilizing common-sense reasoning from at least one artificial intelligence model. As such, hierarchical belief states serve as the result of the agent's comprehensive estimation of the environmental state required for task execution, and subsequently form the basis for generating candidate actions and selecting actions.
[0227] In one embodiment, the hierarchical belief state may be composed of multiple layers of different levels. Referring to FIG. 10, for example, when a task "bring me a banana" is input and currently observable environmental information such as kitchen cabinet 1, refrigerator, and living room shelf is obtained, at least one artificial intelligence model may generate the following hierarchical belief state.
[0228] The top layer (G, Global) contains comprehensive estimated information regarding the overall composition of the environment and user behavior patterns. For example, high-level environmental hypotheses such as "This house has a general structure where food is stored with a kitchen-centered layout, and bananas are likely to be stored in the kitchen storage space" can be expressed as beliefs in the top layer. The broad Gaussian curve shown in the lower left of Fig. 10 represents this initial belief state with high uncertainty.
[0229] The middle level (R, Room-level) contains estimated information regarding the probability of an object's existence in different areas within the environment. For example, beliefs differentiated by area, such as "there is a higher probability of bananas in the kitchen area than in the living room area," can be expressed in the middle level.
[0230] The lowest level (S, Subject-level) contains estimated information regarding the specific location or state of individual objects. For example, specific estimations regarding individual storage spaces can be expressed in the lowest level, such as "It is most likely that the banana is in Cabinet 1, followed by the refrigerator and the living room shelf." This corresponds to the initial belief text in FIG. 10, "The banana will be in the cabinet or the refrigerator."
[0231] In this way, hierarchical belief states represent environmental states in a hierarchical structure of G→R→S, enabling the agent to systematically manage various levels of information, ranging from a comprehensive understanding of the entire environment to the specific location estimation of individual objects.
[0232] In one embodiment, a hierarchical belief state can be generated by at least one artificial intelligence model by considering environmental information and user input together. For example, a large language model (LLM) can combine user input such as "bring me a banana" with currently observed environmental information to perform common-sense reasoning regarding the general storage location of bananas and express this as a hierarchical belief state. In this case, since the hierarchical belief state includes not only directly observed information but also reasoning results based on the artificial intelligence model's prior learning knowledge, it can provide a reasonable initial estimate regarding environmental conditions outside the observation range.
[0233] Next, at least one processor can generate at least one candidate action based on a hierarchical belief state (S107).
[0234] A candidate action is an action corresponding to a search target location, object, or interaction target determined based on estimated information regarding the location or state of an object included in the current hierarchical belief state, and constitutes a candidate set of actions that the agent can next take to perform a task.
[0235] Referring to FIGS. 10 and 13, a candidate action generation module (535) of a server computing device (500) receives a hierarchical belief state from a belief management module (530) and can determine a target candidate region associated with a task goal based on estimation information regarding the location or state of an object generated in step (S105). A method for generating a candidate action corresponding to the determined target candidate region can be implemented in various embodiments.
[0236] In one embodiment, the candidate action generation module (535) may operate in a rule-based manner to generate candidate actions by directly referencing location estimation information of individual objects included in the lowest layer (S) of the hierarchical belief state. For example, based on estimation information such as "Cabinet 1 has the highest estimation probability, followed by the refrigerator and living room shelf," candidate actions such as "move to Cabinet 1," "move to the refrigerator," and "move to Living Room Shelf 1" corresponding to the top k locations of estimation probability can be regularly extracted. This method has the advantage of fast processing speed because it derives candidate actions directly from the belief state without separate model inference.
[0237] In another embodiment, the candidate behavior generation module (535) can generate candidate behaviors from hierarchical belief states using a learned predictor. For example, a prediction model learned from past task performance experiences can receive the current belief state as input and output candidate behaviors with a high probability of completing the task. This method has the advantage of generating candidate behaviors that more richly reflect the task context than simple probability-based extraction.
[0238] In another embodiment, the candidate behavior generation module (535) can use a large language model (LLM) to infer the probability of an object's existence within a target candidate region and generate a corresponding candidate behavior. For example, the LLM can consider the user input "Bring me a banana" together with the current hierarchical belief state to evaluate the probability of each target candidate region through common-sense inference, such as "Cabinet 1 is a kitchen storage space and there is a possibility that a banana is stored there," "The refrigerator is a food storage place and there may be a banana there," and "The living room shelf is a space other than the kitchen, but food may be stored there," and generate a corresponding candidate behavior. This method has the advantage of enabling rich contextual inference using prior learning knowledge.
[0239] In another embodiment, the above methods may be used in combination. For example, the method may operate by first extracting candidates with high belief probabilities using a rule-based method, re-evaluating the probability of the extracted candidates using LLM, or supplementing locations with high search value even if their belief probabilities are low with additional candidates.
[0240] In one embodiment, candidate actions are not limited to locations with a high probability of belief, and locations that have not yet been explored or locations estimated to have a relatively low probability in the belief state may also be included as candidate actions. For example, referring to FIG. 10, a living room shelf estimated to have a relatively low probability of containing a banana in the lowest level (S) of the estimated information may also be included as a candidate action, because the amount of information that can be obtained when exploring that location in the subsequent information gain (IG) calculation step may be high.
[0241] In other embodiments, candidate actions are not limited to movement to a specific location and may include various forms of behavior, such as direct interaction with objects (e.g., opening a door, pulling a drawer), querying external information sources, or additional environmental observation. Through this, the agent can acquire information about the environment and perform tasks in various ways, in addition to physical exploration.
[0242] Next, at least one processor can select one of at least one candidate action for each of at least one candidate action based on an evaluation that considers the expected environment information for performing the action (S109).
[0243] Referring to FIGS. 10 and 11, the action selection is performed by an expected environment simulation module (540) and an action selection module (550), and may include a simulation step for generating expected environment information for each candidate action and an Information Gain (IG) calculation step. A detailed description thereof will be provided later with reference to FIG. 11.
[0244] First, the predicted environment simulation module (540) can perform a simulation using at least one artificial intelligence model to generate a sample of predicted environment information expected to be obtained when performing each candidate action (S1091). In one embodiment, the predicted environment simulation module (540) can independently generate a plurality of samples of predicted environment information for each candidate action.
[0245] At this time, the artificial intelligence model that generates predicted environmental information samples can be implemented in various forms. In one embodiment, the artificial intelligence model may include a large language model (LLM). In this case, the LLM receives current hierarchical belief states and candidate behaviors as input, predicts what environmental information will be obtained at a given location through common-sense reasoning, and can independently generate multiple predicted environmental information samples through stochastic sampling.
[0246] In another embodiment, the artificial intelligence model may include a learned predictor learned from past task performance experiences. In this case, the predictor estimates the probability distribution of environmental information to be obtained when performing a corresponding action based on learned parameters, conditioned on the current belief state and candidate actions, and can generate multiple predicted environmental information samples therefrom. In yet another embodiment, the artificial intelligence model may include a small language model (SLM) or a fine-tuned model specialized for a specific domain, and can generate more accurate predicted environmental information samples that reflect the characteristics of a specific environment.
[0247] For example, referring to FIG. 10, for a candidate action "move to living room shelf 1," at least one AI model can independently generate multiple samples of expected environment information, such as "there will be books," "there will be snacks," and "there will be a remote control." As a living room shelf is a space where various items can be stored, the prediction results may not converge and may be dispersed in various ways even from the perspective of the AI model's common-sense reasoning.
[0248] On the other hand, for the candidate action "move to Cabinet 1," samples of expected environmental information such as "there will be plates and bowls" and "there will be pots and kitchen utensils" can be generated. Cabinet 1 in the kitchen is generally a space where kitchenware is stored, and from the perspective of common-sense reasoning of the AI model, the prediction results may have the characteristic of consistently converging to kitchenware.
[0249] As such, the degree of dispersion of the expected environmental information samples reflects the AI model's common-sense confidence regarding the location, which is subsequently used to calculate the pre-alignment score.
[0250] Next, the information gain calculation module (551) can calculate the information gain for each candidate action based on the generated expected environment information sample (S1093).
[0251] Information gain is an indicator representing how much the uncertainty of the hierarchical belief state will decrease when the corresponding candidate action is performed. It can be calculated by performing a virtual update that sequentially updates the hierarchical belief state from the top level (G) to the bottom level (S) by virtually reflecting expected environmental information samples, and based on the change in uncertainty between the hierarchical belief states before and after the virtual update.
[0252] For example, referring to Fig. 10, let us assume that for a candidate action "move to living room shelf 1," expected environment information samples such as "book," "snack," and "remote control" are generated.
[0253] If we virtually apply these samples to hierarchical belief states, for instance, when "snacks" are found, the current belief state—"bananas are in the cabinet or refrigerator"—is significantly modified to "this user stores food in public places," drastically reducing uncertainty. Similarly, when a "remote control" is found, a new environmental pattern—"the living room shelf is a storage space for household items"—is reflected in the belief state, narrowing the search range for the banana's location and thus reducing uncertainty. As such, searching for the living room shelf is highly likely to significantly update the current belief state regardless of the outcome, which can result in a high information gain.
[0254] On the other hand, the samples of expected environmental information for the candidate action "move to Cabinet 1" converge mostly to kitchenware, such as "plates, bowls" and "kitchen tools." If these samples are hypothetically reflected in a hierarchical belief state, for example, when "plates" are found, the lowest level (S) updates the estimated information to "Cabinet 1 stores cutlery and does not contain bananas," thereby excluding Cabinet 1 from the search candidates; similarly, when "bowls" are found, the estimated information is updated to "Cabinet 1 is a kitchenware storage space," thereby excluding Cabinet 1 from the search candidates. In other words, regardless of which sample is reflected, only the estimated information for Cabinet 1 is updated, while the estimated information for the remaining locations (refrigerator, living room shelf, etc.) remains unchanged, and the overall belief structure itself at the highest level (G), "bananas will be in the kitchen storage space," does not change.
[0255] As such, regardless of the result, the cabinet 1 search only slightly adjusts the probability values for individual locations in the lowest level (S), and the overall belief structure of the highest level (G), "the banana will be in the kitchen storage," does not change.
[0256] On the other hand, if food is found in a space outside the kitchen, such as when snacks are discovered during a search of the living room shelf, the belief of the top tier (G) itself is completely revised in the direction that "this user stores food in a public place," and accordingly, the middle tier (R) and the lowest tier (S) are also drastically updated in a chain reaction.
[0257] As such, since Cabinet 1 search only adjusts the probability of individual locations without structural changes in beliefs, the information gain is relatively low, whereas Living Room Shelf search is more likely to induce a significant update of beliefs across the entire hierarchy, so the information gain can be the highest. Accordingly, as shown in the bar graph of IG scores by behavior in the lower left of Figure 10, the IG score of Living Room Shelf can be the highest.
[0258] In one embodiment, the action selection module (550) can select candidate actions based on information gain (S1095). In this case, the candidate action with the highest information gain is selected as the final action, and a move to the living room shelf can be selected, as shown in the success route at the top of FIG. 10.
[0259] In another embodiment, the action selection module (550) can select a candidate action by considering the prior alignment score together with the information gain.
[0260] The pre-alignment score calculation module (552) can calculate a pre-alignment score based on the correlation between multiple expected environment information samples generated for each candidate action.
[0261] For example, the expected environment information samples for "Move to Living Room Shelf 1," such as "books," "snacks," and "remote control," are distinct from one another, so a low pre-sorting score is calculated, whereas the expected environment information samples for "Move to Cabinet 1," such as "plates, bowls," and "kitchen tools," consistently converge to kitchenware, so a high pre-sorting score can be calculated.
[0262] The pre-alignment score is an indicator representing the reliability of the AI model's prediction for a given location; even if the information gain is high, if the prediction itself is unstable, this can be reflected in the decision-making process.
[0263] Such pre-alignment scores can have a complementary relationship with information gain. For example, moving to living room shelf 1 has high information gain but a low pre-alignment score, whereas moving to cabinet 1 has low information gain but a high pre-alignment score.
[0264] Choosing action based solely on information gain can lead to a so-called "groundless gamble" where locations with low predictive reliability are prioritized for exploration, whereas choosing action based solely on prior alignment scores can result in repeated inefficient searches that only confirm already expected outcomes.
[0265] Therefore, by selecting an action based on an evaluation score that combines information gain and prior alignment scores with a predetermined weight, the agent can select an action that effectively updates the belief state while ensuring a certain level of prediction reliability.
[0266] In another embodiment, the evaluation score calculation module (553) calculates an evaluation score by combining information gain and pre-alignment scores based on a predetermined weight, and can select the candidate action with the highest evaluation score as the final action.
[0267] For example, an evaluation score can be calculated by weightedly combining the information gain and the pre-alignment score as in score(a) = λ·IG(a) + (1-λ)·Align(a) (where a is a candidate action).
[0268] In one embodiment, the weight λ can be dynamically adjusted according to the uncertainty index of the hierarchical belief states. For example, in a situation where uncertainty is high, such as during the initial search, the weight for information gain can be set high to select an action that quickly resolves uncertainty, and in a situation where the search has progressed and the belief states have converged to some extent, the weight for the prior alignment score can be increased to prioritize the selection of an action with high prediction reliability.
[0269] For example, in the initial state of the banana finding task, the uncertainty of the hierarchical belief state regarding the location of the banana is high, so λ can be set high, such as 0.8, to select an action based on information gain. In this case, as shown in Fig. 10, the belief state can be updated quickly by prioritizing the selection of a location with high information gain, such as a living room shelf, where the prediction results are diverse.
[0270] Subsequently, when the belief pattern "this user stores food in a public place" is established through the search of the living room shelf and uncertainty decreases below a certain level, λ can be adjusted to a low value, such as 0.3, to select an action based on the prior alignment score. At this stage, public places such as the countertop or dining table, which the AI model predicts with high confidence, are prioritized, allowing for efficient search for bananas. By dynamically adjusting λ according to the uncertainty indicator in this way, the agent can rapidly align its belief state through active information gathering during the initial stages of search, and perform efficient search based on highly reliable predictions during the later stages.
[0271] In this way, the action selected in step (S109) can be passed to a subsequent processing component and actually executed, and a cyclic loop can be repeated in which the hierarchical belief state is updated based on the actual environment information obtained after the action execution and leads to the next candidate action generation step. A detailed explanation of this will be described later with reference to FIG. 12.
[0272] Below, with reference to FIG. 10, the entire scenario after the living room shelf is selected in step (S109) is divided into a success route and a failure route.
[0273] Referring to the success route (top path) in Fig. 10, the action of moving to living room shelf 1 is executed, and "snack found" can be obtained as an actual environmental information sample. This is a different result from the expected environmental information samples prior to the action execution, such as "book" and "remote control," and a low posterior alignment score is calculated. However, the actual environmental information that "food was found in a space outside the kitchen" is reflected in the hierarchical belief state, and the top level (G) is significantly updated in the direction that "this user stores food in a public place," and accordingly, the middle level (R) and the lowest level (S) are also updated sequentially. Based on the updated belief state, a new candidate action is generated, and as a result of moving to countertop 1, a banana is found, and the task can be successfully completed.
[0274] Referring to the failure route (bottom path) in Fig. 10, assume the case where Cabinet 1, which has low information gain, is selected instead of the living room shelf. An action to move to Cabinet 1 is executed, and "dishes found" may be obtained as actual environmental information. This result matches the expected environmental information sample "dishes, bowls," and thus a high posterior alignment score is calculated; however, as illustrated at the bottom of Fig. 10, no structural change in the hierarchical belief state occurs, moving from "expected kitchen items → no belief change." Subsequently, Countertop 2 is selected and moved, but a knife is found, so the banana is not found, and the task fails or is delayed.
[0275] As such, FIG. 10 intuitively demonstrates that selecting an action with high information gain contributes to efficiently completing a task by rapidly aligning the belief state with the environment. This corresponds to the concept of "Align While Search," which is the core principle of the present disclosure.
[0276] Hereinafter, a process for obtaining actual environment information after the action selected in step (S109) is executed and updating the hierarchical belief state based thereon is described in detail with reference to FIG. 12.
[0277] FIG. 12 is a flowchart illustrating the flow of a method for calculating a post-action alignment score and updating a hierarchical belief state after an action execution according to one embodiment of the present disclosure.
[0278] Referring to FIG. 12, after the action selected in step (S109) is executed, the actual environment information acquisition step (S121), the post-alignment score calculation step (S123), and the hierarchical belief state update step (S125) may be performed.
[0279] First, at least one processor can acquire a sample of actual environment information according to the execution of a selected action (S121).
[0280] The selected action can be executed in various ways. In one embodiment, the selected action can be directly executed by a physical device equipped with a drive unit, such as a robot, an autonomous mobile device, or a drone, based on action data calculated by a server computing device (500).
[0281] Here, action data is a concept encompassing various forms of data for executing selected actions, and can be implemented in diverse forms, such as commands like "go to shelf," output tokens of language models, vectors containing control parameters like joint angles and movement speeds for directly controlling the actuators of physical devices, or code that can be directly executed in a simulation or execution environment. As such, action data can be generated in a suitable form depending on the type of environment in which the agent operates and the characteristics of the subsequent processing components.
[0282] The server computing device (500) can input the generated action data into a subsequent processing component. The subsequent processing component is a component that receives the action data and executes an actual action, and can be implemented in various forms depending on the type of environment in which the agent operates.
[0283] In one embodiment, the subsequent processing component may include a control module of a physical device equipped with a drive unit, such as a robot arm, an autonomous mobile robot, or a drone. In this case, the server computing device (500) generates a control signal to control the drive unit of the physical device based on action data, and the physical device can execute an action by receiving the control signal and driving the drive unit.
[0284] In another embodiment, the subsequent processing component may include an execution module of a physics engine-based simulation environment. In this case, action data is input into a simulator to execute actions in a virtual environment, and state information obtained from the simulation results can be utilized as actual environment information. In yet another embodiment, the subsequent processing component may include industrial control systems, such as a manufacturing process control system of a smart factory, a conveyor control module of a logistics automation system, or an IoT device control interface of a smart home.
[0285] In another embodiment, the subsequent processing component may include a software execution environment such as a web browser, an external API call module, or a database query module, and in this case, action data may be input in the form of API call commands, database query statements, etc., and the result may be utilized as actual environment information.
[0286] In another embodiment, the follow-up processing component may include a user interface module of a user computing device. In this case, the action data is presented to the user through the user interface, and the system may operate in a Human-Agent Collaboration manner where actual execution takes place only when the user approves the action. As such, the follow-up processing component can be implemented in various forms depending on the characteristics of the environment in which the agent operates, and the action determination method according to the present disclosure is applicable regardless of the type of the follow-up processing component.
[0287] In this case, actual environment information is directly obtained from the results of action execution through visual sensors, distance measuring sensors, etc. mounted on the physical device, and can be updated in real time in conjunction with the agent's physical movement.
[0288] In another embodiment, the selected action may be controlled and executed by a server computing device (500) remotely connected to a physical device. In this case, the result of the action execution acquired by a sensor equipped in the physical device may be received via a network (170) and utilized as actual environment information.
[0289] In another embodiment, the selected action may be executed in a physics engine-based simulation environment. In this case, state information corresponding to the result of the action execution may be obtained from the simulation environment as actual environment information.
[0290] In another embodiment, the selected action may be executed through software interactions such as a web browser, external API calls, or database queries, and the response data returned as a result may be obtained as actual environment information.
[0291] The environment information analysis module (555) can obtain and analyze actual environment information from the result of an action execution. For example, actual environment information such as "snack found" can be obtained as a result of the action of moving to living room shelf 1.
[0292] Next, at least one processor can calculate a post-alignment score based on the correlation between the acquired actual environment information sample and the expected environment information sample (S123).
[0293] The post-alignment score calculation module (560) can calculate a post-alignment score by comparing an expected environment information sample generated through simulation before action execution with an environment information sample actually obtained.
[0294] For example, if the expected environmental information sample for "moving to living room shelf 1" was "books", "remote control", etc., but "snacks" were actually found, the correlation between the expectation and the actual is low, so the posterior alignment score may be calculated low. On the other hand, if the expected environmental information sample for the action of moving to cabinet 1 was "plates, bowls" and "plates" were actually found, the expectation and the actual match, so the posterior alignment score may be calculated high.
[0295] Next, at least one processor can determine an update weight that reflects actual environment information samples into the hierarchical belief state based on the calculated posterior alignment score, and update the hierarchical belief state by applying the determined weight (S125).
[0296] The belief state update module (565) can determine update weights based on posterior alignment scores and perform updates by reflecting actual environment information into the hierarchical belief state.
[0297] In one embodiment, the update of the hierarchical belief state can be performed in a bottom-up manner, which is sequentially propagated from the lowest layer (S) to the highest layer (G). First, based on actual environment information and update weights, the estimated information for individual objects included in the lowest layer (S) is updated first, and the updated information of the lowest layer (S) is propagated to the upper layers to sequentially update the estimated information of the intermediate layer (R) and the highest layer (G).
[0298] For example, if a snack is found on living room shelf 1, the lowest level (S) is first updated to "snack on living room shelf 1, no banana," and this information is propagated to the middle level (R) to be updated to "food can be stored in the living room area," and the overall belief can be updated at the top level (G) to "this user stores food in a public place."
[0299] In another embodiment, the update weight may be determined inversely proportional to the posterior alignment score. The lower the posterior alignment score, that is, the greater the discrepancy between the prediction and the actual environmental information, the more strongly the actual environmental information is reflected in the belief state, resulting in a significant update of the belief state; conversely, the higher the posterior alignment score, the weaker the actual environmental information is reflected in the belief state, resulting in a slight reinforcement of the belief state.
[0300] For example, if a completely different snack than expected is found on living room shelf 1, the update weight is set high so that the hierarchical belief state is significantly updated, whereas if a plate as expected is found on cabinet 1, the update weight is set low so that only the estimated information for cabinet 1 of the lowest level (S) is slightly adjusted.
[0301] The updated hierarchical belief state is passed to the candidate action generation module (535) to serve as the basis for generating new candidate actions, and a cyclic loop that repeats from step (S107) of FIG. 9 continues. In one embodiment, the cyclic loop may be terminated when the goal condition of the task is satisfied or when the state in which the posterior alignment score is above a threshold value persists for a predetermined number of times or more, and the amount of change in the hierarchical belief state converges to below the threshold value.
[0302] The reason for not setting the satisfaction of the task's goal conditions alone as the termination condition is as follows. In environments with limited observation, there may be cases where the target object is not found even after all searchable locations have been explored.
[0303] For example, in a banana-finding task, if all searchable locations within the house are explored but the banana is not found, this may imply that the banana does not exist within the current environment or is located outside the agent's observation range. In such a situation, if the termination condition is set solely to the fulfillment of the task objective, a problem may arise where the agent repeats the search indefinitely. Therefore, the circular loop can be terminated by determining that the environment is understood or the objective condition is impossible to achieve when the change in the hierarchical belief state converges below the threshold after the posterior alignment score has remained above a threshold for a predetermined number of times—that is, when the belief state is no longer significantly updated even if additional searches are performed.
[0304] Additionally, if the post-alignment score is below a predefined threshold, the inappropriate decision-making judgment module (570) determines that the decision of the corresponding action is inappropriate and can further update the hierarchical belief state by obtaining additional information through the active information search module (580).
[0305] Below, a process for obtaining additional information and updating the hierarchical belief state through active information search when an inappropriate decision is determined is described in detail with reference to Fig. 13.
[0306] FIG. 13 is a block diagram showing the module configuration of a server computing device (500) according to one embodiment of the present disclosure.
[0307] Referring to FIG. 13, the server computing device (500) may include a user input acquisition module (510), an environment information acquisition module (520), a belief management module (530), a candidate behavior generation module (535), an expected environment simulation module (540), a behavior selection module (550), an environment information analysis module (555), a post-alignment score calculation module (560), a belief state update module (565), an inappropriate decision judgment module (570), and an active information search module (580). The roles of each module and the relationships between modules will be described in detail below.
[0308] The user input acquisition module (510) can acquire input instructing a user to perform a task through a data input interface and store it in memory. For example, it acquires and stores user input such as "Bring me a banana" and transmits it to the belief management module (530). In one embodiment, the user input acquisition module (510) can acquire various forms of user input, such as text input, voice recognition, and image input, and can convert the acquired input into a natural language format and store it.
[0309] The environment information acquisition module (520) can receive environment information acquired from the surrounding environment through sensors equipped in the agent and store it in memory. For example, it receives and stores environment information acquired within the agent's observation range, such as "Cabinet 1 in the kitchen, refrigerator confirmed, shelf in the living room confirmed, interior of each storage space not observed."
[0310] The belief management module (530) receives user input from the user input acquisition module (510) and environmental information from the environmental information acquisition module (520), respectively, and can generate a hierarchical belief state using at least one artificial intelligence model. For example, it generates an initial hierarchical belief state such as "bananas will be in the cabinet or refrigerator" and transmits it to the candidate action generation module (535).
[0311] The candidate action generation module (535) receives a hierarchical belief state from the belief management module (530) and can generate at least one candidate action based on estimated information regarding the location or state of an object included in the hierarchical belief state. For example, it generates candidate actions such as "move to cabinet 1," "move to refrigerator," and "move to living room shelf 1" and transmits them to the expected environment simulation module (540). As illustrated in FIG. 13, when an updated hierarchical belief state is transmitted from the belief state update module (565), the candidate action generation module (535) can regenerate a new candidate action based on the updated belief state.
[0312] The expected environment simulation module (540) receives candidate actions from the candidate action generation module (535) and can perform a simulation to generate samples of expected environment information that are expected to be obtained when each candidate action is performed using at least one artificial intelligence model. For example, for a candidate action "move to living room shelf 1," multiple samples of expected environment information such as "there will be books," "there will be snacks," and "there will be a remote control" are generated and transmitted to the action selection module (550).
[0313] The action selection module (550) receives an expected environment information sample from the expected environment simulation module (540) and can select an optimal action through the information gain calculation module (551), the pre-alignment score calculation module (552), and the evaluation score calculation module (553). The information gain calculation module (551) calculates the information gain of each candidate action based on the expected environment information sample, the pre-alignment score calculation module (552) calculates a pre-alignment score based on the correlation between multiple expected environment information samples generated for each candidate action, and the evaluation score calculation module (553) can select an optimal action by calculating an evaluation score that combines the information gain and the pre-alignment score based on a predetermined weight.
[0314] The environment information analysis module (555) can analyze actual environment information obtained after the action selected by the action selection module (550) is executed. For example, when the action of moving to living room shelf 1 is executed, actual environment information such as "snack found" is obtained and analyzed and transmitted to the post-alignment score calculation module (560).
[0315] The post-alignment score calculation module (560) receives actual environmental information from the environmental information analysis module (555) and can calculate a post-alignment score based on the correlation between the expected environmental information sample and the actual environmental information sample. The calculated post-alignment score is transmitted to the belief state update module (565) and the inappropriate decision judgment module (570).
[0316] The belief state update module (565) receives posterior alignment scores and actual environment information from the posterior alignment score calculation module (560), determines update weights based on the posterior alignment scores, and can update the hierarchical belief state by applying the determined weights. The updated hierarchical belief state is transmitted to the candidate behavior generation module (535) and becomes the basis of the next cyclic loop.
[0317] The inappropriate decision-making judgment module (570) receives a post-alignment score from the post-alignment score calculation module (560), and if the post-alignment score is below a predefined threshold, it determines that the decision of the corresponding action is inappropriate and can activate the active information search module (580).
[0318] For example, if a completely different result is obtained from living room shelf 1 and the post-alignment score is below the threshold, the inappropriate decision-making judgment module (570) may determine that it is difficult to perform the task efficiently with only the current hierarchical belief state and that it is necessary to obtain additional information through active information search.
[0319] The active information search module (580) receives an activation signal from the inappropriate decision judgment module (570) and can actively acquire additional environmental information to resolve the uncertainty of the hierarchical belief state through the additional information inference module (581), the external information source search module (582), and the environment re-search control module (583).
[0320] The additional information inference module (581) can infer what additional information is needed to perform a task using at least one artificial intelligence model in an inappropriate decision-making situation.
[0321] For example, if a snack is found on living room shelf 1 and a low post-alignment score is calculated, the additional information inference module (581) performs the inference that "since food was found in a space outside the kitchen, additional public places such as countertops and dining tables must be searched," and in addition, the inference rule can be dynamically modified or reinforced to increase search efficiency in similar situations in the future.
[0322] The external information source search module (582) can generate a search query to obtain additional information inferred by the additional information inference module (581) and can obtain relevant knowledge from external information sources such as external knowledge bases, search engines, and APIs. For example, by generating a search query for "banana storage location" and searching an external knowledge base, it can obtain knowledge that "bananas are generally stored at room temperature and are often kept on kitchen countertops or in fruit baskets."
[0323] The environment re-exploration control module (583) can re-explore the surrounding environment by controlling a physical device based on the inference result of the additional information inference module (581), and can store the new environment information obtained through re-exploration in memory. For example, the presence of a banana can be confirmed by moving an agent toward the cooking counter and directly observing the location.
[0324] In this way, additional environmental information obtained by the active information search module (580) is transmitted to the belief state update module (565) and used to further update the hierarchical belief state, and a circular loop is repeated in which a new candidate action is generated by the candidate action generation module (535) based on the updated hierarchical belief state. Through this, the agent can efficiently perform tasks by independently obtaining additional information and updating the belief state even in inappropriate decision-making situations.
[0325]
[0326] The method and system for determining the behavior of an artificial intelligence agent based on hierarchical belief states in an environment where observation is limited, according to the embodiments of the present disclosure described above, can be applied to various industrial fields.
[0327] In one embodiment, the present disclosure may be applied to the fields of manufacturing and logistics.
[0328] In a smart factory environment, a robot agent operates in a partially observable environment where it is difficult to fully grasp the state of the entire process line due to limitations in the observation range of sensors, blockage of the line of sight by physical obstacles, etc. When applying the method according to the present disclosure, the robot agent generates a hierarchical belief state composed of a global belief (G), a region belief (R), and an object belief (S) from the environmental information of the currently observable process line, and can perform tasks such as part assembly, welding, and painting based thereon. For example, when the agent performs an assembly task without being able to identify the location of a specific part, a candidate action generation module (535) generates a candidate action based on the hierarchical belief state, and the expected environment simulation module (540) and the action selection module (550) select the search location with the highest information gain (IG), thereby efficiently verifying the location of the part. If, after the action is executed, the post-alignment score calculated by the post-alignment score calculation module (560) is below the threshold value and is judged as an inappropriate decision by the inappropriate decision judgment module (570), additional information regarding the location of the part is obtained through the additional information inference module (581), external information source search module (582), and environment re-search control module (583) of the active information search module (580), and the belief state update module (565) updates the hierarchical belief state so that the assembly task can be successfully completed.
[0329] In a logistics environment, the autonomous mobile robot can actively search for the location of cargo within the warehouse based on hierarchical belief states, generate an obstacle avoidance path in real time, and respond flexibly by performing external information source search and environment re-search through the active information search module (580) when an inappropriate decision is determined by the inappropriate decision judgment module (570). In addition, new inference rules obtained through the additional information inference module (581) can be dynamically modified and reinforced to continuously improve search efficiency in similar situations.
[0330] In other embodiments, the present disclosure may be applied to the medical and healthcare fields.
[0331] A surgical assistance robot operates in an environment where it is difficult to fully grasp the overall condition of the surgical site due to limitations in the camera field of view, obstruction of view by tissues and instruments, etc. When applying the method according to the present disclosure, the surgical assistance robot can perform precise surgical assistance movements by generating a hierarchical belief state from environmental information of the currently observable surgical site and by the action selection module (550) selecting an optimal surgical assistance action based on information gain and pre-alignment scores. For example, if the location of a surgical instrument is unclear, candidate actions can be generated based on object beliefs (S), and the search action with the highest information gain can be selected to determine the exact location of the instrument. If, after the execution of an action, it is determined to be an inappropriate decision by the inappropriate decision judgment module (570), necessary additional information can be inferred through the additional information inference module (581), and relevant information can be obtained from an external medical knowledge base through the external information source search module (582), thereby allowing the belief state update module (565) to update the hierarchical belief state to perform more accurate surgical assistance movements.
[0332] The hospital drug transport robot can perform the task of safely transporting drugs while actively responding to unpredictable environmental changes, such as obstacles in the hallway and the open / closed status of doors. In particular, it can re-explore the surrounding environment through the environment re-exploration control module (583) and determine the optimal movement path in real time based on the hierarchical belief state updated by the belief state update module (565).
[0333] In another embodiment, the present disclosure may be applied to the fields of autonomous driving and mobility.
[0334] Autonomous vehicles operate in various partially observable environments, such as road conditions outside the sensor range, limited visibility due to weather and lighting conditions, and obscuration by other vehicles and pedestrians. When applying the method according to the present disclosure, the autonomous driving agent generates a hierarchical belief state based on environmental information obtained from the current sensor, and the global belief (G) represents a high-level hypothesis regarding the overall road structure and traffic flow, the area belief (R) represents the risk and congestion levels for each road section, and the object belief (S) represents the probability regarding the location of individual vehicles and pedestrians. Based on this, the action selection module (550) can determine actions such as driving path planning, lane changes, and obstacle avoidance. If, after the execution of an action, the post-alignment score is below a threshold value and the decision is determined to be inappropriate, additional information regarding a detour route is obtained from an external information source, such as a traffic information service, through the external information source search module (582), and the belief state update module (565) updates the hierarchical belief state to perform safe detour driving. Furthermore, the inference rule updated through the additional information inference module (581) can continuously improve the efficiency of action decision in similar road conditions.
[0335] In drone delivery services, delivery tasks can be completed while actively adapting to unpredictable environmental changes, such as obstacles on the flight path and weather changes. In particular, new obstacles or weather changes on the flight path can be reflected in real time through the belief state update module (565), and if an inappropriate decision is determined by the inappropriate decision judgment module (570), an alternative flight path can be quickly determined through the active information search module (580).
[0336] In another embodiment, the present disclosure may be applied to the fields of agriculture and environmental monitoring.
[0337] An agricultural robot agent operates in an environment where the growth status of crops, the occurrence of pests and diseases, and soil moisture conditions can be observed only partially through sensors in a vast farmland. When applying the method according to the present disclosure, the agricultural agent generates a hierarchical belief state from the currently observable environmental information of the crops, and the action selection module (550) selects the search location with the highest information gain to efficiently perform tasks such as sowing, weeding, and harvesting. For example, when pests and diseases are detected in a specific area, the probability distribution of the object belief (S) changes significantly, and based on the judgment of the inappropriate decision judgment module (570), the active information search module (580) is activated to further search the state of surrounding areas and determine the optimal action to prevent the spread of pests and diseases, thereby performing pest control operations. The inference rules updated through the additional information inference module (581) can improve pest control efficiency in similar situations in the future.
[0338] In environmental monitoring drones, based on environmental information within a limited sensor range in tasks such as wildfire monitoring, marine pollution detection, and wildlife observation, the belief state update module (565) updates the hierarchical belief state, and the action selection module (550) can actively adjust the search range based on information gain. For example, a wildfire monitoring drone can update the area belief (R) centered on the point where smoke is detected and prioritize searching adjacent areas with high information gain, thereby quickly identifying the direction and scale of the wildfire's spread.
[0339] In another embodiment, the present disclosure may be applied to the field of construction and infrastructure inspection.
[0340] At a construction site, a robot agent operates in an environment where it is difficult to fully grasp the entire site due to complex structures, the movement of workers and equipment, etc. When applying the method according to the present disclosure, the construction robot agent generates a hierarchical belief state from currently observable environmental information of the construction site, maintains a probability distribution for the location of materials based on object beliefs (S), and the action selection module (550) selects the search action with the highest information gain, thereby enabling efficient performance of tasks such as material transport, welding, and bolt fastening.
[0341] Inspection drones for infrastructure facilities such as bridges, tunnels, and transmission towers operate in environments where it is difficult to assess the condition of the entire structure at once due to the complex shape of the structure and limitations in sensor range. When applying the method according to the present disclosure, the inspection drone probabilistically infers the probability of defect occurrence for each part of the structure based on a hierarchical belief state, and the action selection module (550) can effectively detect defects such as cracks and corrosion by prioritizing the inspection location with the highest information gain. If an inappropriate decision is determined by the inappropriate decision judgment module (570) during the inspection process, the design information and past inspection history of the relevant part are obtained from an external structure database through the external information source search module (582), and the belief state update module (565) updates the hierarchical belief state to perform more accurate defect detection.
[0342] In another embodiment, the present disclosure may be applied to the field of home service robots.
[0343] A home service robot operates in a dynamic, partially observable environment where environmental conditions change frequently due to the arrangement of various objects within the home, lighting conditions, and the movement of residents. When applying the method according to the present disclosure, the home service robot can understand user instructions in the form of natural language, generate a hierarchical belief state from currently observable home environment information, and perform tasks such as cleaning, organizing objects, and preparing food. For example, when a user instructs, "Bring me a banana," the belief management module (530) generates a hierarchical belief state for each location within the home, and the action selection module (550) can efficiently search for a banana by selecting the search location with the highest information gain. If, during the search process, the post-alignment score calculated by the post-alignment score calculation module (560) is below a threshold value and is judged as an inappropriate decision by the inappropriate decision judgment module (570), additional information regarding the location of the banana is inferred through the additional information inference module (581), and additional information is obtained through the external information source search module (582) or the environment re-search control module (583), thereby allowing the belief state update module (565) to update the hierarchical belief state, so that the task can be successfully completed. The inference rule updated through the additional information inference module (581) reflects the same user's object placement pattern in the future, enabling more efficient search to be performed from the beginning in similar search tasks.
[0344] In another embodiment, the present disclosure may be applied to the field of web-based agents and information retrieval.
[0345] A web-based agent operates in a partially observable environment where it must process user requests based on incomplete or unstructured information in a dynamically changing web environment. When applying the method according to the present disclosure, the web-based agent generates a hierarchical belief state from information on a currently accessible web page, and the global belief (G) may represent a high-level hypothesis regarding the overall structure and information arrangement of the website, the domain belief (R) may represent the information relevance of each page section, and the object belief (S) may represent the probability regarding the location of a specific information element. For example, if a user is instructed to search for specific product information, the action selection module (550) may select the web page search action with the highest information gain, and the belief state update module (565) may gradually access the target information while updating the hierarchical belief state based on the acquired information. If an inappropriate decision is determined by the inappropriate decision judgment module (570) during the search process, a necessary search query may be inferred through the additional information inference module (581), and additional information may be acquired through the external information source search module (582) to perform a more accurate information search.
[0346] As such, the method and system for determining the behavior of an artificial intelligence agent based on hierarchical belief states in an environment where observation is limited according to the present disclosure can be widely applied to all fields operating in partially observable environments, such as manufacturing, logistics, medical, autonomous driving, agriculture, construction, household services, and web-based information retrieval.
[0347] The present disclosure has a technical advantage over existing fixed policy-based agent systems in that it provides an adaptive mechanism that does not rely solely on pre-learned policies but dynamically updates hierarchical belief states at the time of inference, autonomously determines optimal behavior based on information gain and alignment scores, and improves itself through an active information search module (580) when inappropriate decisions occur.
[0348] Furthermore, the agent has a self-improvement characteristic in which its performance is continuously improved over time by dynamically modifying or reinforcing inference rules based on new information obtained through the additional information inference module (581) during the active information search process, which can contribute to continuously increasing the success rate and efficiency of task performance in various real-world environments.
[0349] Meanwhile, the embodiments according to the present disclosure described above may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the computer-readable recording medium may be those specifically designed and configured for the present disclosure or those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. Hardware devices may be modified into one or more software modules to perform processing according to the present disclosure, and vice versa.
[0350] The specific embodiments described in this disclosure are examples and do not limit the scope of this disclosure in any way. For the sake of brevity of the specification, descriptions of prior electronic configurations, control systems, software, and other functional aspects of said systems may be omitted. Additionally, the connections of lines or connecting members between components shown in the drawings are illustrative of functional connections and / or physical or circuit connections, and may be replaced or additionally represented as various functional connections, physical connections, or circuit connections in actual devices. Furthermore, unless specifically stated as “essential,” “importantly,” etc., a component may not be strictly necessary for the application of the embodiments of this disclosure.
[0351] Furthermore, although the detailed description of the present disclosure has been explained with reference to preferred embodiments of the present disclosure, those skilled in the art or those with ordinary knowledge in the art will understand that various modifications and changes can be made to the embodiments of the present disclosure without departing from the spirit and technical scope of the present disclosure as set forth in the claims below. Accordingly, the technical scope of the present disclosure should not be limited to the contents described in the detailed description of the specification but should be determined by the claims.
[0352] Various embodiments of the present disclosure have industrial applicability in that they enable efficient decision-making even under incomplete environmental information by maintaining hierarchical belief states in environments with limited observation, dynamically balancing search and exploit based on information gain and pre-aligned scores, and selecting optimal actions, and can be widely applied to improve the decision-making performance of artificial intelligence agents in various partially observable real-world environments such as robot manipulation, autonomous mobile robots, smart factories, and logistics automation.
Claims
1. As a method executed by a computer, A step of obtaining user input that directs a predetermined task through a data input interface; A step of obtaining information about the environment required for performing the above task; A step in which at least one processor uses at least one artificial intelligence model to generate a hierarchical belief state regarding an environment state for performing the task based on information about the environment; A step of generating at least one candidate behavior based on the above hierarchical belief state; and A method comprising: a step of selecting one of the at least one candidate action based on an evaluation that considers expected environmental information regarding the performance of the action for each of the at least one candidate action.
2. In Paragraph 1, A step of calculating action data corresponding to the above-mentioned selected action; and A method further comprising the step of inputting the above-mentioned action data into a subsequent processing component.
3. In Paragraph 2, A method in which the above-mentioned subsequent processing component is configured to execute the above-mentioned action based on the above-mentioned action data.
4. In Paragraph 1, A method in which the above hierarchical belief state includes estimated information regarding the possibility that an object exists at a specific location within an environment or the state of said object.
5. In Paragraph 1, The above hierarchical belief states consist of multiple layers of different levels, and The top layer includes comprehensive estimation information regarding the configuration of the entire environment and user behavior patterns, and The intermediate layer includes estimated information regarding the probability of object existence by region within the environment, and A method in which the lowest layer includes estimated information regarding the specific location or state of individual objects.
6. In Paragraph 1, The step of generating the above hierarchical belief state is, A method comprising the step of generating estimated information regarding the location or state of an object from information about the environment using at least one artificial intelligence model.
7. In Paragraph 6, The step of generating at least one candidate action above is, A step of determining a target candidate region associated with the task goal based on estimated information regarding the possibility of existence or location of an object included in the above hierarchical belief state; and A method comprising: a step of inferring the probability of an object existing within a target candidate region through at least one artificial intelligence model with respect to the determined target candidate region, and generating at least one candidate action corresponding thereto.
8. In Paragraph 1, The step of selecting one action among the above-mentioned at least one candidate action is, A simulation step of generating expected environment information samples expected to be obtained when performing each candidate action using the above-mentioned at least one artificial intelligence model; Based on the above-mentioned sample of expected environment information, a step of calculating an Information Gain representing the degree of uncertainty reduction of the hierarchical belief state for each of the at least one candidate behavior; and A method comprising the step of selecting one of the at least one candidate action based on the information gain above.
9. In Paragraph 8, The above simulation step is, The step of independently generating a plurality of expected environment information samples for each candidate action through the above-mentioned at least one artificial intelligence model; The step of selecting one action among the above-mentioned at least one candidate action is, For each of the above candidate actions, a step of calculating a pre-alignment score based on the degree of correlation between a plurality of generated expected environment information samples; and A method further comprising the step of selecting one of the at least one candidate action based on the information gain and the calculated pre-alignment score.
10. In Paragraph 9, The step of selecting one action among the above-mentioned at least one candidate action is, A step of calculating an evaluation score for each candidate action by applying weights to the information gain and the alignment score, respectively; and A method comprising the step of selecting one of the at least one candidate action based on the evaluation score.
11. In Paragraph 10, The step of calculating the above evaluation score is, A method comprising the step of: dynamically adjusting the weights according to an uncertainty index calculated based on the entropy or variance of the hierarchical belief states, wherein if the uncertainty index is above a threshold, the weight for the information gain is set relatively high, and if the uncertainty index is below the threshold, the weight for the prior alignment score is set relatively high to calculate the evaluation score.
12. In Paragraph 8, A step of obtaining actual environment information samples based on the execution of the above-mentioned selected action; A step of calculating a post-alignment score based on the correlation between the acquired actual environmental information sample and the predicted environmental information sample; and A method further comprising the step of determining an update weight to reflect the actual environment information sample into the hierarchical belief state based on the calculated posterior alignment score, and updating the hierarchical belief state by applying the determined weight.
13. In Paragraph 12, The step of updating the hierarchical belief state comprises: a step of prioritizing the updating of existence probability or attribute information for individual objects included in the lowest layer of the hierarchical belief state based on the actual environment information and the update weights; and A method comprising the step of propagating the updated information of the lowest layer to the upper layer to sequentially update the inference content of the upper layer, including the nature of the area where the object is located or the task-related context.
14. In Paragraph 12, A method comprising: a step of repeating from the step of generating candidate behaviors to the step of updating hierarchical belief states, wherein the step of terminating the repetition step is performed when the hierarchical belief state satisfies the goal condition of a predefined task, or when the state in which the posterior alignment score is above a threshold value persists for a predetermined number of times or more, and the amount of change in the hierarchical belief state converges to below the threshold value, thereby determining that the goal condition cannot be achieved or that the environment has been identified.
15. In Paragraph 12, A step of determining that the decision-making of the selected action is inappropriate if the calculated post-alignment score is below a predefined threshold; In the above-mentioned inappropriate decision-making situation, an active information search step for obtaining auxiliary information by performing an external knowledge source query or additional environmental interaction to resolve the uncertainty of the above-mentioned hierarchical belief state; and A method further comprising the step of dynamically modifying or augmenting the inference rules referenced by the artificial intelligence model when forming beliefs about the environment or generating the candidate behavior, based on the auxiliary information obtained above.
16. At least one memory; and At least one processor that reads at least one instruction stored in the above-mentioned at least one memory and performs a behavior decision method of an environment-adaptive hierarchical belief state-based artificial intelligence agent; comprising The above at least one instruction is, A step of obtaining user input that directs a predetermined task through a data input interface; A step of obtaining information about the environment required for performing the above task; A step in which at least one processor uses at least one artificial intelligence model to generate a hierarchical belief state regarding an environment state for performing the task based on information about the environment; A step of generating at least one candidate behavior based on the above hierarchical belief state; and A system comprising an instruction to perform the step of selecting one of the at least one candidate action based on an evaluation that considers expected environmental information regarding the performance of the action for each of the at least one candidate action.
17. In Paragraph 16, A plurality of neurons comprising an array including at least one register, at least one programmable logic, and at least one input interface; a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; and at least one routing network that controls the data flow between the plurality of neurons; comprising A system further comprising a Field Programmable Gate Array (FPGA) implementation for a predetermined artificial neural network, wherein each of the plurality of neurons is connected to at least one other neuron through the routing network to establish a transmission path for the weights.
18. In Paragraph 16, A plurality of neurons organized into an array comprising at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; comprising A system comprising an Application Specific Integrated Circuit (ASIC) for a predetermined artificial neural network, wherein each of the plurality of neurons is connected to at least one other neuron through any one of the plurality of synaptic circuits.