Multi-agent integrated robot using quantum computing and operating method therefor
The multi-agent integrated robot uses quantum computing to optimize agent unit behavior by evaluating environmental information and determining optimal policies, addressing collaborative learning challenges in decentralized environments.
Patent Information
- Application Number
- PCT/KR2025/010882
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-24
- Filing Date
- 2025-07-23
- Publication Date
- 2026-01-29
AI Technical Summary
Multi-agent reinforcement learning faces challenges in learning collaborative behavior due to the difficulty in optimizing actions among multiple agents, particularly in decentralized environments.
A multi-agent integrated robot utilizing quantum computing, where a quantum main unit collects and evaluates environmental information from multiple agent units, determining optimal behavioral policies through quantum computing to maximize the value function of each agent unit.
Enables optimized operation and control of a large number of agent units by leveraging quantum computing to determine ideal actions, even in complex environments with multiple agents.
Smart Images

Figure KR2025010882_29012026_PF_FP_ABST
Abstract
Description
Multi-agent integrated robot using quantum computing and its operating method
[0001] The present invention relates to the operation of a multi-agent, and more particularly, to a multi-agent integrated robot using quantum computing to optimize the behavior of a multi-agent.
[0002] Multi-Agent Reinforcement Learning (MARL) is a branch of reinforcement learning (RL). It deals with the process by which multiple agents, interacting with each other, learn to find optimal policies. The goal can be interpreted as finding the optimal actions for each agent or the optimal action strategy for the entire system.
[0003] MARL is important in diverse fields such as robotics, autonomous vehicles, and game theory, and involves unique complexities and challenges that do not arise in single-agent reinforcement learning.
[0004] Compared to single-agent reinforcement learning, the key feature of MARL is that multiple agents must interact and learn from each other's actions and outcomes, thereby finding the optimal action in a state where the actions of each agent are organically linked.
[0005] Multi-agent reinforcement learning (MARL) is a fully decentralized method in which each agent learns using only its own observation information and acts based on this, but it has the problem of difficulty in learning collaborative behavior.
[0006] To solve this problem, recent studies mainly utilize the Centralized Training and Decentralized Execution (CTDE) method, in which each agent trains a behavioral model by collecting all observation information from each agent, and then executes the behavioral model using only its own observation information.
[0007] Meanwhile, the rapid development of quantum computing is also bringing about significant changes in artificial intelligence learning. Quantum computing is known for its revolutionary data processing speeds, as it utilizes the phenomenon of quantum superposition to process qubits (quantum bits), the smallest unit of information processing, which are in a mixed state of 0 and 1. In neural network training, qubits serve as neurons, the basic units of neural networks. Quantum systems containing qubits can be designed to mimic conventional neural networks and function as neural networks.
[0008] Dong D. Chen C. Li H. Tarn TJ (2008). Quantum reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics. Part B, Cybernetics, 38(5), 1207-1220. 10.1109 / TSMCB. 2008.92574318784007.
[0009] Narayanan A. Menneer T. (2000). Quantum artificial neural network architectures and components. Information Sciences, 128(3-4), 231-255. 10.1016 / S0020-0255(00)00055-4.
[0010] Saggio, V. (2021). Experimental quantum speed-up in reinforcement learning agents. Nature 591, 229-233 (2021). https: / / arxiv.org / abs / 2103.06294
[0011] Chao Yu (2021), The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games, arXiv:2103.01955, https: / doi.org / 10.48550 / arXiv.2103.01955
[0012] Georgios Papoudakis (2019), Dealing with Non-Stationarity in Multi-Agent Deep Reinforcement Learning, arXiv:1906.04737
[0013] Yihan Wang (2020), Off-Policy Multi-Agent Decomposed Policy Gradients, arXiv:2007.12322, https: / doi.org / 10.48550 / arXiv.2007.12322
[0014] The purpose of the present invention to solve the above problems is to provide a multi-agent integrated robot using quantum computing and an operating method thereof.
[0015] One aspect of the present invention to achieve the above object provides a multi-agent integrated robot using quantum computing.
[0016] The multi-agent integrated robot using the above quantum computing includes a quantum main unit including a computational unit (CPU) that performs quantum computing; and a plurality of agent units whose actions are controlled by the quantum main unit.
[0017] Each of the above plurality of agent units collects environmental factors using at least one sensor, and operates by driving at least one power drive unit mounted inside according to the collected environmental factors.
[0018] Each of the plurality of agent units generates environmental information including the environmental elements and provides the information to the quantum main unit.
[0019] The quantum main unit is configured to collect all of the environmental information collected from each of the plurality of agent units, evaluate the contribution corresponding to each of the plurality of agent units through a quantum computing-based operation on the collected plurality of environmental information, determine a behavioral policy that maximizes the value function of the corresponding agent unit in an environment according to the environmental information of each of the plurality of agent units based on the contribution, and provide the determined behavioral policy to the corresponding agent unit.
[0020] The multi-agent integrated robot using the above quantum computing includes a group of agent units arranged in each of the four directions of up, down, left, and right based on the center where the quantum main unit is arranged.
[0021] The above group is composed of 25 agent units arranged in a grid pattern with 5 agent units arranged horizontally and 5 agent units arranged vertically, spaced apart from each other at a predetermined interval.
[0022] The at least one sensor is arranged at predetermined intervals along boundary lines corresponding to the upper, right, and lower directions centered on the group.
[0023] Each of the above agent units has a single motion axis configured to be movable in a direction perpendicular to the plane on which the agent unit is placed.
[0024] When using the multi-agent integrated robot and its operating method using quantum computing according to the present invention as described above, since quantum computing is used, even if a very large number of agent units are deployed, the operation of each agent unit can be optimized and controlled to enable the most ideal operation.
[0025] In particular, it is possible to propose a deployment structure that enables optimal operation in an environment where agent units are deployed in groups to have a single axis of movement and are controlled for movement.
[0026] Figure 1 is a conceptual diagram for explaining the operating environment of a multi-agent integrated robot using quantum computing according to one embodiment.
[0027] FIG. 2 is a diagram illustrating an example of a configuration of a multi-agent integrated robot using quantum computing according to one embodiment.
[0028] FIG. 3 is a two-dimensional planar drawing of a multi-agent integrated robot using quantum computing according to one embodiment.
[0029] FIG. 4 is a drawing showing a three-dimensional structure of a multi-agent integrated robot using quantum computing according to one embodiment.
[0030] FIG. 5 is a diagram showing an example of an expanded application of agent units of a multi-agent integrated robot using quantum computing according to FIG. 4.
[0031] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.
[0032] Terms such as first, second, A, and B may be used to describe various components, but these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component. The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0033] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0034] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0035] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0036] Hereinafter, a preferred embodiment according to the present invention will be described in detail with reference to the attached drawings.
[0037]
[0038] FIG. 1 is a conceptual diagram illustrating the operating environment of a multi-agent integrated robot using quantum computing according to one embodiment. FIG. 2 is a diagram exemplarily illustrating the configuration of a multi-agent integrated robot using quantum computing according to one embodiment.
[0039] Referring to FIG. 1, a multi-agent integrated robot (100) using quantum computing may include a quantum main unit (Quantum Computer, 101) including a computational unit (CPU) that performs quantum computing, and a plurality of agent units (Quantum Actor robots, 102, hereinafter also referred to as QAR) whose actions are controlled by the quantum main unit (101).
[0040] For example, the quantum main unit (101) may include a quantum computer having a computational unit (CPU) that performs quantum computing, and may further include a computing device having a computational unit (CPU) based on conventional electrical signals in addition to quantum computing.
[0041] In addition, the quantum main unit (101) can function as a cloud server that can be accessed from the outside via a wired or wireless network, or can be configured to communicate with a cloud server located in an external network, and can provide services such as Infrastructure-as-a-Service (IaaS), Platforms-as-a-Service (PaaS), Software-as-a-Service (SaaS), and Function-as-a-Service (FaaS) utilizing quantum computing resources to a user terminal that has accessed the cloud server by communicating with the cloud server.
[0042] In one embodiment, each of the plurality of agent units (102) may be a power device that includes at least one sensor, collects environmental information using the included at least one sensor, and drives at least one power drive unit (e.g., a motor) mounted therein to perform an action based on the collected environmental information.
[0043] In another embodiment, each of the plurality of agent units (102) may be a control unit that includes at least one sensor, collects environmental information using the included at least one sensor, and drives at least one unit control unit mounted therein to perform an action based on the collected environmental information. In one example, the control unit may be a power supply control device that controls the amount of driving of a motor.
[0044] In one embodiment, each of the plurality of agent units (102) has at least one motion axis and can perform movement in the direction of the motion axis as an action. For example, if the agent unit (102) has a single axis, the agent unit (102) may include a linear motor actuator having one motion axis as a power drive unit. As another example, if the agent unit (102) is a multi-axis robot having two or more motion axes, the agent unit (102) may be a robot arm in which each of the plurality of motors is arranged at a joint and each motor functions to rotate as a motion axis.
[0045] Here, the multi-agent integrated robot (100) can be defined as an object that includes a set of multiple unit robots (wherein the unit robots correspond to agent units (102)) and is configured to perform a specific task or purpose by controlling the mutual operation of each of the unit robots (wherein the device controlling the unit robots corresponds to the quantum main unit (101)).
[0046] For example, Extended PyMARL (EPyMARL), one of the Python-based codebases that provides reinforcement learning environments, provides several experimental environments using multiple agents, such as the objective (or mission) of loading all scattered food, the objective (or mission) of moving scattered shelves to a target point, or the objective (or mission) of winning a battle by attacking or healing depending on the unit type.
[0047] As a specific example, if an environment is assumed in which the purpose (or mission) of moving scattered shelves to a target point is performed using multiple loading and unloading robots, each of the multiple robots performing loading and unloading in the environment may correspond to an agent unit (102), and a robot or control unit that controls the mutually organic behavior of the agent units may be defined as a quantum main unit (101), and a concept that collectively refers to a set of agent units and the quantum main unit may be defined as a multi-agent integrated robot (100).
[0048] As another simple example, when a robot composed of multiple drive axes to grab an object to a specific location and move it is implemented as a multi-agent integrated robot, a control unit that controls the driving amount (or torque, power, etc.) of the corresponding drive axes by 1:1 arrangement corresponding to each of the multiple drive axes constituting the robot in order to achieve the purpose or mission of grabbing and moving the object with an appropriate force and direction can be interpreted as an agent unit (102), and the change in the position of the robot according to the driving amount of the corresponding agent unit and the repulsive force of the object transmitted to the corresponding drive axes can be defined as environmental information, and the main control unit that moves the object according to the desired purpose using the agent units corresponding to each of the drive axes can be defined as a quantum main unit (101).
[0049] In order to simulate the experimental environment related to this, various test environment setting methods and known algorithms for policy calculation according to the test environment, including https: / agents.inf.ed.ac.uk / blog / epymarl / , are publicly available, so ordinary technicians can use them.
[0050] Assuming an understanding of this multi-agent reinforcement learning environment, the environment is an external world in which agents can take action and observe the results thereof, and agent units (102) can take action in the environment in which they are each located and observe changes in the environment according to the action to obtain environmental elements.
[0051] For multi-agent reinforcement learning, it is common to use either a centralized training with decentralized execution (CTDE) framework or a decentralized training with decentralized execution (DTDE) framework.
[0052] In the case of the former (CTDE), the data for learning is a method in which each agent unit (102) learns by compiling all of the environmental information collected by each agent unit (102), and in the case of the latter (DTDE), the data for learning is a method in which the agent unit (102) learns only by using the environmental information collected by itself.
[0053] In any case, the operation process of the multi-agent after learning is the same in that each agent only uses the environmental information (or observation information) that it has collected.
[0054] In both frameworks, each agent unit (102) is equipped with a DQN (Deep Q-Network), a type of deep learning network, and learns independently. Then, the learned deep learning network, DQN, can be used to dynamically respond to changes in individual environmental factors.
[0055] Additionally, the agent unit (102) can generate environmental information including acquired environmental elements and provide the generated environmental information to the quantum main unit (101). For example, when acquiring environmental elements using an infrared sensor, the environmental elements may be light entering through the light-receiving portion of the infrared sensor. Generating environmental information from environmental elements may include processing the environmental elements into a form of a protocol capable of communicating with the quantum main unit (101).
[0056] The quantum main unit (101) can collect all environmental information collected from each of a plurality of agent units (102), evaluate the contribution of each of the plurality of agent units (102) to the purpose (or task) through quantum computing-based calculations on the collected environmental information, determine an action policy for maximizing the value function of each of the agent units (102) in the environment according to the currently collected environmental information, and provide the determined action policy to the corresponding agent unit (102).
[0057] Here, the contribution is a value individually calculated for each agent unit (102), and can be set so that the sum of the contributions calculated for each agent unit (102) is 1. For example, the initial contribution calculated for each agent unit (102) is standardized to have a value between 0 and 1, and the value obtained by dividing the standardized initial contribution by the sum of all agent units (102) is used as the contribution of the corresponding agent unit (102), so that the sum is 1. At this time, the method for determining the initial contribution can be calculated in an independent manner by a person skilled in the art depending on the purpose of the simulation environment, and is not limited to a specific formula.
[0058] Here, a behavior policy may consist of multiple behavioral guides that instruct the agent to perform a specific action based on specific environmental information. An agent unit (102) provided with a behavior policy determines the action to be taken based on the current environmental information based on the behavior policy and performs the determined action.
[0059] Here, the value function is a function that evaluates the state of the agent unit (102) according to the environmental information as a higher value the closer it is to achieving the goal (or mission) when a specific behavior policy is provided to the agent unit (102).
[0060] As a simple example, if the goal (or mission) is to reach a destination using a robot that can only move forward in one direction, the current position of the robot can be a state, and a value function can be created with a value that is inversely proportional to the distance between the current position and the destination position.
[0061] Considering a wide variety of simulation environments, the value function can be defined as shown in Equation 1 below.
[0062]
[0063] Referring to the above mathematical expression 1, the value function is a value (V) that represents the degree to which the provided agent unit (102) changes the state s closer to the goal when providing the action policy π in the state s according to the current environmental information. π(s)), where π(a|s) is the probability that the agent unit (102) takes action a in state s when π is provided as the action policy, A is a set of actions that the agent unit (102) can take, P(s', r|s, a) is the probability that the state changes to s' and reward r is received when action a is taken in state s, and γ is a weight representing the relative importance of the value of the future changed state s' compared to the current state s. At this time, the reward r is a value that represents the degree to which the agent unit (102) is advantageous in acting when the state s changes to state s' in a given environment (for example, in the case of an agent unit that moves an object to a destination point, since there is a reward that allows for less movement as the distance to the destination point is closer, it can be determined as a higher value as the distance to the destination point is closer), and may be a function that is individually defined by a typical technician according to each simulation environment.
[0064] The quantum main unit (101) can select one action policy among multiple predefined action policies based on a value function, which maximizes the result value of the value function, and provide the selected action policy to the agent unit (102). Expressed as a formula, the action policy π' selected in the current state s according to the environmental information of the agent unit (102) is as shown in the following mathematical formula 2.
[0065]
[0066] Here, the quantum main unit (101) calculates p, which is the number of changeable cases of the state confirmed through the environmental information of each agent unit (102), and calculates q, which is the number of cases of actions that the corresponding agent unit (102) can take for each estimated case, and then uses quantum computing to define in advance a number of action policies corresponding to pXq. As a simple example, if the environmental element observed through the environmental information is the 'distance' to the destination and the type of action that the agent unit (102) can take is motor angle adjustment, the action policies can be defined as action policies that compare the 'distance' with at least one threshold distance and adjust the angle between 0 and 180 degrees clockwise or counterclockwise depending on whether the result is greater than the threshold distance.
[0067] That is, the quantum main unit (101) can make full use of the advanced computational power of quantum computing to evaluate the value for reaching a goal (or task) in a specific action policy and state s using a value function for each of the plurality of agent units (102), select an action policy that can maximize the evaluated value, and continuously repeat the process of distributing it to each agent unit (102), and this process can be repeated until no change in the action policy distributed to each agent unit (102) occurs in any of the agent units (102).
[0068] In one embodiment, the quantum main unit (101) can multiply the aforementioned contribution, which indicates the degree of influence of each agent unit (102), on the resultant value of the value function according to the above mathematical expression 1, and use the resulting value as the resultant value of the value function. This has the advantage of being able to individually reflect the influence of each agent unit (102) in the environment in which it is located, thereby producing a value function.
[0069] By collecting all environmental information and determining an action policy that allows each agent unit (102) to determine the optimal action to achieve the purpose (or mission), the action policy is provided to the agent unit (102), and the agent unit (102) can perform the next action according to the provided collaboration policy to collect environmental information and repeatedly transmit it to the quantum main unit (101).
[0070]
[0071] FIG. 3 is a two-dimensional planar diagram of a multi-agent integrated robot using quantum computing according to one embodiment. FIG. 4 is a three-dimensional diagram of a multi-agent integrated robot using quantum computing according to one embodiment.
[0072] Referring to FIGS. 3 and 4, it can be seen that the multi-agent integrated robot (100) using quantum computing is represented in two-dimensional and three-dimensional structures according to the deployment environment.
[0073] First, referring to FIG. 3, the multi-agent integrated robot (100) may include groups of agent units (102) spaced apart in each of the four directions of up, down, left, and right based on the center where the quantum main unit (101) is placed. At this time, the groups of agent units (102) may be spaced apart from each other at predetermined intervals and placed in a grid shape. For example, the group of agent units (102) may be composed of 25 agent units (102) arranged in a grid shape with 5 agent units (102) horizontally and 5 agent units (102) vertically. In one embodiment, the groups may be placed on a square base substrate.
[0074] Here, a group of agent units (102) may be arranged with a predetermined number of sensors spaced apart from each other at predetermined intervals in each of the three directions except for the direction in which the quantum main unit (101) is located. For example, five sensors may be arranged at predetermined intervals along the boundary lines corresponding to the upper, right, and lower directions centered on the group of agent units (102). That is, when a total of 100 agent units (102) are installed, including 25 agent units (102) in each direction as shown in the drawing, a total of 60 sensors may be arranged, with 15 sensors in each direction. Here, each sensor may be shared by agent units (102) within a predetermined range centered on the area in which the sensor is arranged, and environmental factors collected from the sensor may be collected to generate environmental information. In one embodiment, the sensor is an infrared sensor, and the environmental factor may be light (wavelength of light, amount of light energy, etc.) emitted by the infrared sensor and reflected by an object and entering the light receiving portion of the infrared sensor.
[0075] Additionally, the agent unit (102) may have a single axis of motion, for example, each agent unit (102) may have said single axis of motion configured to be movable in a direction perpendicular to the plane on which the agent unit is placed (i.e., in the direction of entering or exiting the ground corresponding to the plane drawing with reference to FIG. 3).
[0076] Meanwhile, referring to FIG. 4, a group of agent units (102) may include a first group and a second group, each of which is arranged to face each other based on the plane of the base substrate on which the group is arranged.
[0077] At this time, each of the agent units (102) belonging to the first group may have a single axis of motion that can move in the first direction (Axis1) with respect to the plane of the base substrate.
[0078] Each of the agent units (102) belonging to the second group may have a single axis of motion that is movable in a second direction (Axis2) opposite to the first direction (Axis1) with respect to the plane of the base substrate.
[0079] In one example, a single motion axis may be arranged perpendicular to the plane of the base substrate, and the agent unit (102) may be formed in a rod shape with the direction of movement of the single motion axis as the length direction, and one of the two ends of the agent unit (102) facing the base substrate may be arranged so that the motion axis is exposed to the outside.
[0080]
[0081] FIG. 5 is a diagram showing an example of an expanded application of agent units of a multi-agent integrated robot using quantum computing according to FIG. 4.
[0082] Referring to FIG. 5, the multi-agent integrated robot (100) using quantum computing may further include a third group formed in a continuous grid shape adjacent to the first group and a fourth group formed in a continuous grid shape adjacent to the second group.
[0083] Additionally, the third group and the fourth group may be arranged to face each other based on the plane to which the base substrate belongs, the third group may be configured to have the same axis of motion as the first group, and the fourth group may be configured to have the same axis of motion as the second group.
[0084] That is, in the embodiment according to FIG. 5, the multi-agent integrated robot (100) using quantum computing can have the first to fourth groups arranged in each of four directions forming a 90-degree angle with respect to the center where the quantum main unit (101) is arranged.
[0085] Accordingly, the multi-agent integrated robot (100) using quantum computing in the embodiment according to FIG. 5 may include a total of 400 agent units (102) and a total of 200 sensors.
[0086] In the embodiment according to FIG. 5, since the sensors are not arranged between the third group and the first group, five sensors may be arranged along each of the upper and lower outer lines of the first group, and five sensors may be arranged along each of the upper, lower, and right outer lines of the third group, for a total of 25 sensors may be arranged along the outer lines of the first and third groups. Similarly, a total of 25 sensors may be arranged along the outer lines of the second and fourth groups.
[0087]
[0088] The methods according to the present invention may be implemented in the form of program instructions that can be executed by various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either singly or in combination. The program instructions recorded on the computer-readable medium may be those specifically designed and constructed for the present invention, or may be known and available to those skilled in the computer software art.
[0089] Examples of computer-readable media may include hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions may include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate with at least one software module to perform the operations of the present invention, and vice versa.
[0090] Additionally, the above-described method or device may be implemented by combining all or part of its configuration or function, or may be implemented separately.
[0091] Although the present invention has been described above with reference to preferred embodiments thereof, it will be understood by those skilled in the art that various modifications and changes may be made to the present invention without departing from the spirit and scope of the present invention as set forth in the claims below.
Claims
1. As a multi-agent integrated robot using quantum computing, A quantum main unit including a computational unit (CPU) that performs quantum computing; and A plurality of agent units whose actions are controlled by the quantum main unit; Each of the above plurality of agent units collects environmental factors using at least one sensor, and operates by driving at least one power drive unit mounted inside according to the collected environmental factors. Each of the plurality of agent units generates environmental information including the environmental elements and provides the environmental information to the quantum main unit, The quantum main unit is configured to collect all of the environmental information collected from each of the plurality of agent units, evaluate the contribution corresponding to each of the plurality of agent units through a quantum computing-based operation on the collected plurality of environmental information, determine an action policy that maximizes the value function of the corresponding agent unit in an environment according to the environmental information of each of the plurality of agent units based on the contribution, and provide the determined action policy to the corresponding agent unit. Multi-agent integrated robot using quantum computing.
2. In claim 1, Includes a group of agent units spaced apart in each of the four directions, up, down, left, and right, based on the center where the above quantum main unit is placed. The above group is arranged in a grid shape with a predetermined interval between each other, and is composed of 25 agent units arranged in a grid shape with 5 agent units horizontally and 5 agent units vertically. Multi-agent integrated robot using quantum computing.
3. In claim 1, At least one sensor above, They are arranged at predetermined intervals along the boundary lines corresponding to the upper, right, and lower directions centered on the above group, Each of the above agent units, having a single axis of motion configured to be movable in a direction perpendicular to the plane on which the agent unit is placed; Multi-agent integrated robot using quantum computing.
Citation Information
Patent Citations
Optimizing Policy Controllers for Robotic Agents Using Image Embedding
JP2020530602A
Methods and apparatus for pruning experience memories for deep neural network-based q-learning
KR1020180137562A
Patient monitoring and defibrillation system and method of driving the same
KR1020250176595A
Automobile seat with seat cover detachable device
KR102527484B1
Multi-agent integrated robot using quantum computing and its operation method
KR102762344B1