Techniques for controlling population of robots
By calculating vector field graphs and deflection fields to control the robot population, the problem of high-density robot population navigation and obstacle avoidance in complex areas is solved, and efficient and flexible control effects are achieved.
Patent Information
- Application Number
- CN202280101518.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art is difficult to effectively control high-density robot populations, especially in real-time communication environments that require fast response and high reliability. Conventional radio interface solutions cannot meet the needs of high population density and dynamic changes.
By calculating the vector field graph and the deflection field, radio access is provided to the robot population, and the vector field graph provides a global navigation route, while the deflection field is used for local deflection and obstacle avoidance, ensuring the collision-free movement of the group members.
It realizes efficient control of high-density robot populations, can navigate and avoid obstacles in complex areas, improves the flexibility and responsiveness of the system, and reduces the load of radio resources.
Smart Images

Figure CN120129882A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for controlling a robot swarm. In particular, but not limited thereto, a method and apparatus for controlling a robot swarm including a plurality of swarm members in an area in which a radio unit provides radio access to the robot swarm is provided. Background Art
[0002] The fifth generation of mobile communications (5G) provides flexibility, which is a key requirement for connected robots (e.g., cloud robots) and Industry 4.0. In addition, 5G radio access technology (RAT), such as the new radio (5G NR) specified by the 3rd Generation Partnership Project (3GPP), is a global communication standard with a growing ecosystem, which makes conventional radio interface issues no longer valid.
[0003] The common view is that 5G becomes an important part of the factory infrastructure of the future. The argument for 5G over other wireless technologies is the ability to support real-time communication with end-to-end latency as low as a few milliseconds at a high level of reliability. Some cloud robotics applications rely on real-time connectivity, for example, to enable instant movement of robots. Therefore, connectivity is of utmost importance.
[0004] A particular use case of interest is swarm control, which requires remote control of the velocity (i.e., speed and direction) of a swarm of robots. Controlling a swarm typically requires multiple radio network unicast transmissions to each swarm member, since different swarm members move at different speeds. Conventional multiple unicast transmissions result in high loads at the radio network and may lead to asynchronous behavior of swarm members, since the limited spectrum capacity of the radio network requires time-multiplexing of the unicast transmissions.
[0005] Existing technologies for drone swarms use offline plans that are transmitted to the drones, which execute the plans in parallel. Intensive radio communication between drones is required to maintain a safe area between each other. The ground station is responsible for selecting the various parts of the predetermined plan. For example, patent US9809306B2 relates to controlling unmanned aerial vehicles (UAVs) as a cluster to fly synchronously in an aerial show. Each UAV includes a processor that executes a local control module and a memory that can be accessed by the processor for use by the local control module. The system also includes a ground station system having a processor that executes a swarm manager module and having a memory that stores different flight plans for each of the UAVs. The flight plan is stored on the UAV. During flight operations, each local control module independently controls the corresponding UAV to execute its flight plan without continuous control from the swarm manager module. The swarm manager module is operable to initiate flight operations by triggering the initiation of flight plans by multiple UAVs at the same time. In addition, the local control module monitors the front-end and back-end communication channels, and when the channel is lost, the UAV is operated in a safe mode.
[0006] However, these conventional systems require a lot of computing and communication resources, so they are not efficient enough to respond to sudden changes or high population densities. In addition, they are only suitable for specific purposes. Summary of the invention
[0007] Therefore, a technique is needed that effectively controls an arbitrary number of group members and can therefore be extended to very high numbers or densities of group members without conflicts. In addition, there is a need to navigate group members with low communication bandwidth through complex areas that may change suddenly.
[0008] With respect to a first method aspect, a method for controlling a swarm of robots in an area is provided. The area includes a plurality of radio units for providing radio access to the swarm of robots. The swarm of robots includes a plurality of swarm members. The method includes or initiates the step of determining (e.g., calculating) a vector field map. The vector field map includes velocity vectors indicating speeds and directions for navigating swarm members through the area. The method also includes or initiates the step of determining (e.g., calculating) a deflection field. The deflection field indicates a deflection for deflecting swarm members relative to the vector field map. The method also includes or initiates the step of sending the vector field map and the deflection field to at least one swarm member via a radio unit to control the movement of the at least one swarm member in the area.
[0009] By sending (e.g., broadcasting) a vector field map, an embodiment can provide multiple global routes to all group members of the group in a radio resource efficient manner. For example, the vector field map can include one or more destinations, for example, where the direction of the vector field map converges and / or the speed of the vector field map decelerates. In addition, since the same vector field map is provided to all group members, an embodiment can ensure that the routes are disjoint and inherently conflict-free. For example, any smooth vector field map can define multiple collision-free routes. The deflection field enables an embodiment to effectively control local deflections (e.g., corrections) relative to the global route defined by the vector field map. Since the local group member group in the deflection zone receives (e.g., adjacent group members receive) the same deflection field, the deflection is applied consistently (e.g., simultaneously and uniformly), thereby inherently avoiding collisions. Therefore, embodiments of the technology can effectively use radio resources to control a robot group in a dynamic area.
[0010] In a first variation of any embodiment, the robot swarm may be controlled by broadcasting a deflection field (e.g., as a velocity vector). In a second variation of any embodiment, the deflection field may be determined using artificial intelligence (AI) (i.e., an AI agent trained by training data derived from the motions of swarm members).
[0011] The vector field map may associate a location in the region (e.g., each location in the region) with a velocity vector. The location (e.g., the regional resolution of the vector field map) may include each point or each portion of a grid in the region. The velocity and direction may be used by group members to navigate through the region.
[0012] The vector field map may be defined or may cover the entire region.Alternatively or additionally, the deflection field may be defined or may be non-zero in one or more islands (eg, compact regions) within the region.
[0013] The vector field map may be sent to each swarm member (e.g., by unicast or broadcast) for navigating each swarm member through the area. Alternatively or additionally, the deflection field may be transmitted to at least one swarm member (e.g., by unicast or multicast) for controlling movement of at least one swarm member in the area by deflecting (e.g., guiding) the at least one swarm member relative to (e.g., correcting) the vector field map.
[0014] The radio units may provide radio access to a radio access network (RAN) for the robot swarm. Here, the radio access may include sending data (e.g., vector field maps and deflection fields) from one or more radio units to swarm members in a downlink (DL), and optionally, receiving data (e.g., current position determined using a satellite-based radio navigation system (e.g., Global Navigation Satellite System GNSS)) from the swarm members at the radio units. Alternatively or additionally, the radio units may use (e.g., massive) multiple-input multiple-output (MIMO), e.g., for beamforming, to define a deflection zone in which the deflection field may be received.
[0015] The radio unit may comprise a radio base station (RBS) and / or a cell of a RAN. For example, the radio unit may provide centralized MIMO for beamforming transmission and / or beamforming reception. Alternatively or additionally, the radio unit may comprise a radio spot or a radio strip. For example, the radio unit may provide distributed MIMO and / or cell-free radio access, for example using distributed and phase-synchronized antennas.
[0016] The deflection field can be configured to deflect (e.g., reroute) group members in a deflection zone, for example, to reroute group members around an avoidance zone such as an obstacle (as an example of a deflection zone). The obstacle can be a stationary object or can move within the area. Alternatively or additionally, the deflection field can be configured to reroute group members to move along an alternate path, for example, temporarily changing to another course.
[0017] The deflection field may be or may include a deflection force field. The deflection force field may assign force vectors to different locations. The force vector may be the gradient of a velocity vector (eg, a combined deflection velocity field and a vector field map).
[0018] Alternatively or additionally, the deflection field may be or may include a deflection velocity field. The deflection velocity field may indicate a (e.g., local) correction to the vector field map. The deflection vector field may include a velocity vector (i.e., speed and direction) that is to be added to the velocity vector indicated by the vector field map for each group member.
[0019] The vector field map can be determined (e.g., calculated) by a distributed computing network (e.g., edge servers) or a centralized computing network (e.g., cloud servers). Servers in the distributed computing network can be spatially associated with each radio unit that transmits the deflection field and / or can be spatially associated with each deflection zone (e.g., where the deflection field is receivable and / or non-zero). Different servers in the distributed computing network can calculate vector field maps and / or deflection fields for different deflection zones. Alternatively or additionally, servers in the centralized computing network can determine vector field maps and / or deflection fields for multiple deflection zones.
[0020] The vector field graph may be a dynamic or static vector field graph.Alternatively or additionally, the vector field graph may include one or at least one convergence point, which may be referred to as a goal or objective.
[0021] One or more of the radio points (eg, as examples of radio units) may exclusively transmit deflection fields, while one or more base stations (eg, in addition to the radio points and / or as other examples of radio units) may transmit vector field maps.
[0022] The first method aspect may be performed by a group control entity. The group control entity and / or a centralized server or a distributed server network may calculate the vector field map. The group control entity may include a centralized server or a distributed server network.
[0023] The swarm control entity and / or one or more neural networks (also referred to as artificial intelligence agents or AI agents) may determine the deflection field and / or may determine (e.g., select) radio units (e.g., radio spots) in the area for transmitting the deflection field. The swarm control entity may include one or more neural networks (e.g., AI agents).
[0024] The group members may include at least one of a mobile robot, an automated guided vehicle (AGV), a drone, a bird or insect robot, a humanoid robot, an autonomous car, and a platoon truck. The area may include at least one of an indoor area (e.g., extending over multiple floors) and an outdoor area.
[0025] In any embodiments, the step of determining a deflection field may include selecting one or more radio units from a plurality of radio units, and rerouting group members around obstacles and / or into a deflection zone in the area by implementing at least one safety buoy on the selected one or more radio units, wherein the at least one safety buoy defines or acts as a source of the deflection field.
[0026] Embodiments of the technology may broadcast vector field maps and / or deflection fields using a multimedia broadcast and multicast service (MBMS) for crowd control. Alternatively or additionally, generating vector field maps and / or deflection fields may use a neural network trained by reinforcement learning (also known as AI-assisted).
[0027] A server (e.g., an edge cloud) may compute a vector field map required to navigate a group in an area. The vector field map may be updated by the server (e.g., in an edge cloud) (e.g., regularly or periodically or event-driven). For example, the vector field map differences may be updated accordingly. Embodiments of the techniques may stream the computed vector field map and / or differences (i.e., updates) in a radio-efficient manner using MBMS (e.g., evolved MBMS or eMBMS). The vector field map may include a (e.g., smoothed) velocity field.
[0028] The area may include (e.g., local or compact) deflection zones where a deflection field is applied or is non-zero. Examples of deflection zones are avoidance zones (e.g., safety zones or danger zones) that need to be avoided by the group and / or courses defined by the vector field map that need to be changed by at least some of the group members. For this or other purposes, local deflections (e.g., rerouting) may be added to the vector field map with the aid of the deflection field. At least one of the radio units (e.g., at least one radio point) may be operated as a virtual safety buoy, i.e., one or more of the radio units may broadcast a deflection velocity field (as an example of a deflection field) in a space-time manner. The radio points may be deployed over industrial areas.
[0029] A swarm control entity (e.g., an artificial intelligence agent, AI agent, also referred to as an AI strategy or simply an agent) can control (i.e., operate) at least one of the safety buoys by selecting the necessary radio units (e.g., radio points) to participate in the operation and determine the transmitted (e.g., broadcast) deflection velocity field.
[0030] The group members passing by one or more safety buoys add the vector field map and the received deflection velocity field together and perform corresponding deflection (eg, rerouting). In this article, the vector field map and the deflection velocity field may be collectively referred to as a velocity vector.
[0031] As an alternative or in addition to the deflection velocity field, a deflection force field can be used to achieve deflection (e.g., rerouting). In other words, the vector field map can be corrected by changing the velocity of the group members according to the deflection velocity field (e.g., locally and temporarily) and / or by changing the acceleration of the group members according to the deflection force field (e.g., locally and temporarily), thereby achieving deflection (e.g., rerouting). The deflection force field and the deflection velocity field are collectively referred to as the deflection field.
[0032] A safety buoy (short: buoy) acting as a source of a deflection field may mean that the safety buoy prepares or forms the deflection field. In a variant, the safety buoy may define the center of the deflection zone and / or the size of the deflection field may increase as the distance to the center of the deflection zone decreases. Alternatively or additionally, the deflection field may be uniform (e.g., within the deflection zone) or may explicitly include position information, or the signal-to-noise ratio (SNR) may be a scaling factor of the size of the deflection field (e.g., a scaling factor for the deflection strength).
[0033] The size of the deflection zone may be limited by radio reception (ie within which the transmitted deflection field is receivable).
[0034] In any embodiment, the step of determining the vector field map may be performed before the group members begin to move. Alternatively or additionally, the step of determining the deflection field may be performed while the group members are moving.
[0035] Determining the deflection field as the group members move can be accomplished by determining the deflection field after the group members begin to move and / or before the group members enter a deflection zone where the deflection field is applied. Alternatively or additionally, determining the deflection field dynamically or as the group members move can mean determining the deflection field in real time and / or in response to dynamically changing conditions in the area (e.g., moving obstacles).
[0036] Alternatively or additionally, the vector field map may encode the routes of the group members in a static environment. A static environment may encompass areas with no objects moving in the area (e.g., no obstacles), i.e., static parts of the environment, such as walls (e.g., in indoor areas) or roads and / or buildings (in outdoor areas).
[0037] In any embodiment, the deflection may be induced locally only in a deflection zone within the area and / or the deflection field may be transmitted only by a predetermined subset of radio units around the deflection zone. Alternatively or additionally, the vector field map may be transmitted independently of the deflection zone and / or may be transmitted throughout the area and / or may be transmitted by a base station covering the area.
[0038] The deflection zone (eg, obstacle) may be located in space and / or time. For example, the deflection zone may be centered around a moving obstacle and / or may exist temporarily.
[0039] The predetermined subset of the deflection fields sent can include only radio points or can be a single radio unit.
[0040] The vector field map can be sent from all radio units (e.g., base stations and radio points).
[0041] In any embodiment, at least one of the steps of determining the vector field map and determining the deflection field can be based on or can include the step of performing reinforcement learning (RL) to optimize the deflection of the swarm members. RL can output an optimization strategy used in determining the vector field map and / or determining the deflection field.
[0042] Reinforcement learning (RL) can be performed before the determining step and / or can be based on, for example, training data generated in a simulation of the swarm.
[0043] The strategy can be implemented by a neural network. The neural network can include an input layer, at least one intermediate layer, and an output layer. Each layer can include a plurality of neurons. The output of each neuron in one layer can be coupled to the input of one or more neurons in the next layer. Each coupling can be weighted according to a weight. All the weighted couplings at any one input can be summed at the input. The output of each neuron can be a non-linear (e.g., strictly monotonically increasing) function of the summed input. RL can optimize the weights of the neural network.
[0044] Alternatively or additionally, the strategy can be implemented by a Q-table including rows and columns for states and actions respectively. RL can optimize the Q-table, for example, according to the Bellman equation.
[0045] In any embodiment, the step of performing RL can include training the weights of the neural network. The neural network can be embodied in the strategy used in determining the vector field map and / or determining the deflection field. The neural network can be configured to sense and interpret the environment of the region. The weights can be trained by a positive reward according to the navigation of the vector field map and / or the desired result of the deflection according to the deflection field and / or a negative reward according to the navigation of the vector field map and / or the undesired result of the deflection according to the deflection field.
[0046] The input layer of the neural network can receive a state s and / or (e.g., long-term) reward R. The state s can include the position and / or velocity of the swarm members (e.g., as the first part of the state), and / or the vector field map and / or the deflection field (e.g., as the second part of the state), and / or the position and / or velocity of one or more deflection zones (e.g., obstacles) (e.g., as the third part of the state). The output layer of the neural network can provide an action a. The action can include, for example, at least one of the direction and velocity of the deflection field in the corresponding deflection zone. Alternatively or additionally, the output layer of the neural network can provide at least one of the center and diameter of the deflection zone.
[0047] The transition to a resulting next state based on a combination of state and action can be based on a simulation of the swarm (e.g., taking into account the propulsion forces and masses of the swarm members to calculate the changes in velocity and position under the influence of the deflection field) and / or can be determined online (e.g., where the swarm members provide feedback indicative of the acceleration and / or position measured by each swarm member).
[0048] A short-term reward r may be associated with each transition. The short-term reward r may include at least one of the following components, for example r = r 1 +r 2 +r 3 +…. The first reward component r 1 The second reward component r may depend on the distance D between adjacent group members or a change in the distance D. For example, a decrease in the distance D between adjacent group members or even a collision (i.e., D=0) may be associated with a decrease (or zero value) in the short-term reward. 2 can be associated with the change of each group member's trajectory due to the deflection field relative to the undeflected trajectory defined only by the vector field map. For example, the energy required for the deflection or change of the trajectory can correspond to a reduction in the short-term reward. In addition, the third reward component r 3 may be positive and associated with a corresponding one of the group members that arrives at a destination (eg, defined by a vector field graph).
[0049] For example, the distance D between adjacent group members can be 1 =+D / L 0 or -L 0 / D is associated (for example, where L 0 is a constant or scaling factor or grid length). Alternatively or additionally, the deflection that increases the path length L in direct space (or the trajectory in phase space) by a length ΔL can be expressed as the second component r 2 =-ΔL / L 0 Alternatively or additionally, with the third component r 3 =+L / L 0 The associated positive reward can be L 0 The path length L in direct space (or trajectory in phase space) is in units.
[0050] Alternatively or additionally, each group member may be associated with an expected trajectory resulting from integrating the current position and current velocity (as the first part of the current state) from the current vector field map and the current deflection field (as the second part of the current state). The integration may be performed until the destination is reached or until a maximum integration time T has been reached. The expected trajectory may be associated with a sum of short-term rewards r(t) associated with transitions along the trajectory The long-term reward R is associated with . Optionally, this sum is discounted by a discount factor 0<γ<1:
[0051]
[0052] The long-term reward associated with state s may correspond to the sum of the long-term rewards for each expected trajectory associated with each swarm member.
[0053] The weights of the neurons of the neural network can be randomly initialized. RL can be based on long-term rewards. For example, the agent or neural network can use the difference between the ground truth reward (e.g., the long-term reward determined based on simulations or measured transitions) and the expected reward (e.g., the next state produced by the neural network output or based on the action output by the neural network) as a loss function, and backpropagation through the loss function can be used to update the weights to improve the policy, i.e., maximize the expected long-term reward produced by the policy π(a,s).
[0054] In any embodiment, at least one of the plurality of swarm members may include a sensor for continuously capturing sensor data. The method may further include or initiate the step of receiving data based on the sensor data from the swarm member. The received data may be fed back to the RL for optionally optimizing the deflection of the swarm member as it moves (e.g., for optimizing a strategy for determining a deflection field).
[0055] The data may be received via a radio unit (eg, a radio point) and / or at a group control entity.
[0056] The sensors may include at least one of a position sensor for determining the position of the group members, a speed sensor for determining the speed (or at least the velocity) of the group members. Any sensor may include at least one of a LiDAR (Light Detection and Ranging) unit, a radar unit, a camera unit, and an ultrasonic transceiver.
[0057] RL may be performed continuously (e.g., periodically) in an operational (or field) deployment, i.e., based on data received during the movements of the swarm members.
[0058] In any embodiment, the vector field map and the deflection field may cause the group members to follow a trajectory. The steps of performing RL may include evaluating the long-term reward of the trajectories of the group members. Optionally, the steps of performing RL may include at least one of the following steps: controlling the trajectory for evaluation, the controlled trajectory including a random trajectory or a random destination according to the vector field map or a random deflection according to the deflection field; randomly selecting, according to the vector field map and the deflection field, a trajectory for evaluation from the trajectories executed by the group members; positively rewarding a trajectory having a relatively short or shortest length to a predefined destination; positively rewarding a trajectory having a relatively short or shortest time to a predefined destination; positively rewarding a trajectory having a relatively low or lowest energy consumption to a predefined destination; negatively rewarding or ignoring a trajectory deviating from a shortest trajectory predetermined value; negatively rewarding a trajectory causing a collision of the group members; and comparing with training data indicating the best trajectory.
[0059] The evaluated trajectory may be calculated according to the current policy or the current deflection field. The evaluated trajectory may be an expected trajectory, which may be different from the trajectory generated by the control of the group, for example because the expected trajectory is based on the current policy to be optimized by RL, such that the policy may change as the group members move along the trajectory.
[0060] In any embodiment, the step of determining the deflection field and / or the step of performing RL may further include at least one of the following: exploring the state including the vector field map and the deflection field by taking a random action to modify the deflection field; and exploiting past beneficial actions by taking an action based on at least one of a random sample of the policy after RL and past actions exceeding a minimum reward level.
[0061] In any embodiment, each trajectory in the region may be associated with a long-term reward. The long-term reward may indicate the negative cost generated by the group members reaching the destination in the region. RL may optimize the policy for determining the deflection field by modifying the velocity vectors of the group members to maximize the long-term reward.
[0062] The long-term reward may positively reward reaching the destination without a collision. The long-term reward (or negative total cost) may be generated by integrating the short-term reward (or cost) along the trajectory of each group member. The long-term reward may be implemented by a cost function (e.g., corresponding to the negative long-term reward). The cost function may also be referred to as a loss function.
[0063] Alternatively or additionally, RL may use a loss function, which is a smooth function of the weights of a neural network and represents (or approximates) the (negative) long-term reward. The policy may be optimized by modifying the weights of the neural network according to stochastic gradient descent to maximize the long-term reward or minimize the loss function.
[0064] In any embodiment, RL is executed in the context of a region (e.g., RL is executed based on feedback originating from the region). The context may include at least one production unit that operates in a simulator, or in real hardware that includes a swarm of robots moving in the region, or in hardware-in-the-loop (e.g., at least components of the swarm members, the positions and velocities of the simulated swarm members as hardware-in-the-loop), or in a digital twin of the region and the swarm members.
[0065] In any embodiment, the region may be divided into a plurality of sectors. The step of determining the vector field map may include defining start and destination points that can be connected by a plurality of trajectories in the region. Alternatively or additionally, the step of determining the vector field map may include generating a short-term reward field that indicates, for each sector, the value of the short-term reward for using the corresponding sector on the trajectory. Alternatively or additionally, the step of determining the vector field map may include generating an integral field that indicates, for each sector, the integral value of the (foregoing) long-term reward. The integral value may be integrated based on a plurality of values of the short-term reward for each sector along the trajectory from the start point to the destination point. Alternatively or additionally, the step of determining the vector field map may include generating the vector field map as a flow field by associating with each sector a velocity vector that indicates the direction to an adjacent sector and / or towards the destination based on the integral field.
[0066] The sectors may be implemented as grid squares or blocks.
[0067] Negative rewards may be referred to as costs. The value of the short-term reward may be realized by a negative cost value. The short-term reward field may be realized by a (negative) local cost field. The integral value of the long-term reward may be realized by integrating the cost values.
[0068] In any embodiment, the vector field map and / or the deflection field may further include a destination and / or at least one waypoint. The destination may be an attractor of a velocity vector that exerts an attraction on the swarm members. The waypoint may be associated with a deflection zone, where a deflection (e.g., a shift or a turn) in the same direction is applied to all swarm members within the corresponding one of the deflection zones.
[0069] The attractor may act contrary to the deflection field.
[0070] In any embodiment, the step of determining the vector field map may include updating the vector field map. The sending step may include sending the updated vector field map, or the difference between the updated vector field map and the previously sent vector field map. Optionally, the difference is encoded using a motion vector field based on video coding.
[0071] In any embodiment, to deflect the swarm members, the deflection field may be encoded with the position offsets of the individual swarm members. Optionally, the direction of the shift may be parallel throughout (or the foregoing) deflection region. Alternatively or additionally, the deflection field may be encoded with the change in the velocity of the individual swarm members. Optionally, the direction of the change may be parallel throughout (or the foregoing) deflection region. Alternatively or additionally, the deflection field may be encoded with the center of the deflection region, optionally the center of an obstacle. Alternatively or additionally, the deflection field may be encoded with a force that is parallel throughout the deflection region. Alternatively or additionally, the deflection field may be encoded with a repulsive force associated with the deflection region, optionally a radial force centered on the obstacle. Alternatively or additionally, the deflection field may be encoded with an attractive force associated with a path point, optionally a radial force centered on the path point.
[0072] In any embodiment, the radio unit includes at least one or more of the following: radio points, radio strips, radio units dedicated to controlling the swarm of machines, radio units dedicated to locally transmitting the deflection field and / or acting as safety buoys, at least one or each of the swarm members, a base station of a radio access network (RAN) providing radio access to the swarm of machines, and radio units deployed within another RAN.
[0073] In any embodiment, the sending step uses at least one of a multimedia broadcast and multicast service (MBMS) channel, point-to-point transmission, ultra-reliable low-latency communication (URLLC) according to fifth-generation (5G) mobile communication, massive machine type communication (mMTC) according to 5G mobile communication, non-cellular radio access technologies such as wireless fidelity (Wi-Fi), optical radio access technologies, optionally light fidelity (Li-Fi), unicast transmission, multicast transmission, and broadcast transmission.
[0074] In any embodiment, at least one radio unit among the plurality of radio units may perform unicast transmission to send a vector field map and / or a deflection field to different swarm members using time interleaving or time division multiplexing.
[0075] In any embodiment, the method may further include implementing an anti-collision system. Alternatively or additionally, the deflection field may include a uniform deflection field that is applied to or applicable to all group members in the deflection zone. Alternatively or additionally, the deflection field may be based on sensor data measured by and / or (e.g., at the group control entity or the aforementioned group control entity) received for determining the vector field map and / or for determining the deflection field. Alternatively or additionally, the deflection field may be based on collision events and / or RL may include tracking collision events. For example, RL may include reducing short-term rewards for collision events after updating the vector field map or the deflection field to suppress them. As an example, the method may include tracking collision events and reducing short-term rewards in response to collision events after updating the vector field map or the deflection field to suppress them.
[0076] The method may be implemented as a method for optimizing a strategy for determining a deflection field for group members within a control region. The determination of the vector field map and / or the determination of the deflection field may include providing an environment in which the group members will operate. The environment may include at least one region.
[0077] Alternatively or additionally, performing RL may include providing a set of training data (e.g., for the provided environment) that indicates actions in response to the appearance of obstacles. RL may be performed based on the training data. For example, the strategy generated by the training data may correspond to an initial strategy that will be used by the group control entity in another environment and / or will be optimized based on data received (e.g., in real time) from the group members.
[0078] In any embodiment, the determined deflection field may indicate a uniform velocity vector or a uniform force vector for one or each deflection zone within the region for the deflection of the group members relative to the vector field map. Alternatively or additionally, the deflection field (e.g., velocity vector or force vector) applied by each individual group member to control the movement of at least one group member in the region may further depend on the signal strength of the transmitted deflection field.
[0079] Regarding a second method aspect, a method of controlling a group member is provided. The group member includes at least one actuator configured to change a movement state of the group member that is part of a swarm of robots moving in an area. The method includes or initiates a step of receiving a vector field map. The vector field map includes velocity vectors indicating velocities and directions for navigating the group member through the area. The method further includes or initiates a step of receiving a deflection field. The deflection field indicates deflections for deflecting the group member relative to the vector field map. The method further includes or initiates a step of determining a position of the group member in the area. The method also includes or initiates a step of determining a change in the movement state based on the received vector field map and the received deflection field for the determined position. The method also includes or initiates a step of controlling at least one actuator to effect the changed movement state.
[0080] The second method aspect may be performed by the group member, such as by at least one or each of the swarm robot group members.
[0081] In any embodiment, the step of determining the change in the movement state may include combining the deflection field and the vector field map (e.g., adding their vectors). Alternatively or additionally, the step of determining the change in the movement state may include calculating a rotation vector from the gradient (e.g., curl vector operator) of the combined deflection field and vector field map. The rotation vector may be used to transform the current movement state into the changed movement state.
[0082] In any embodiment, the group member may act as a safety buoy upon deployment. This deployment may also be referred to as manual deployment.
[0083] In any embodiment, the received deflection field may indicate (e.g., for a deflection zone or each deflection zone within the area) a uniform velocity vector or a uniform force vector for deflecting the group member relative to the vector field map. Optionally, the step of determining the change in the movement state for the determined position in the deflection zone may include scaling the received uniform velocity vector or uniform force vector according to, for example, a signal strength of the deflection field received at the group member. The signal strength may be a reference signal received power (RSRP).
[0084] The second method aspect may also include any features and / or any steps disclosed in the context of the first method aspect, or corresponding features and / or steps thereto, e.g., a receiver counterpart of a transmitter feature or step. Vice versa, the first method aspect may also include any features and / or any steps disclosed in the context of the second method aspect, or corresponding features and / or steps thereto.
[0085] Any one of the group members may comprise or be implemented as a radio device, such as a user equipment (UE) according to 3GPP or a mobile station according to Wi-Fi. Alternatively or additionally, any one of the radio units may comprise or be implemented as a base station, such as an eNB or gNB according to 3GPP or an access point according to Wi-Fi.
[0086] Within any one of the one or more deflection zones, the deflection field may be locally transmitted (e.g., locally broadcast) by a local radio unit of the radio units. Alternatively or additionally, the deflection field may be transmitted (e.g., broadcast) by one of the group members (e.g., the first group member detecting an obstacle in the deflection zone, and / or e.g., a leading vehicle in a vehicle queue) to one or more adjacent group members (e.g., all group members in the same deflection zone and / or following vehicles following the transmitting group member). For example, one of the group members may forward the deflection field from a (e.g., fixed) radio unit to one or more adjacent group members.
[0087] The forwarding radio device may be a relay radio device. For example, the forwarding radio device may comprise a communication protocol stack configured to relay the deflection field from the radio unit to one or more adjacent group members, e.g., using the GPRS tunneling protocol (GTP), the User Datagram Protocol (UDP), or the Internet Protocol (IP).
[0088] In any embodiment, the deflection field may be sent (e.g., broadcast and / or forwarded) by one of the group members to one or more of its adjacent group members using a sidelink (SL) (i.e., wireless (e.g., radio or optical) device-to-device communication (e.g., Wi-Fi Direct or proximity-based service, ProSe, according to 3GPP TS23.303, version 17.0.0)). The SL transmission of the deflection field may be implemented according to 3GPP specifications, e.g., according to 3GPP document TS23.303 version 17.0.0 or its modified 3GPP LTE or 3GPP NR, or according to 3GPP document TS33.303 version 17.1.0 or its modified 3GPP NR.
[0089] The quality of service (QoS) required or configured for sending the deflection field, e.g., the maximum latency, may depend on at least one of the speed of the group members and the density of the group members. For example, the maximum latency may be inversely proportional to the product of the speed and the density.
[0090] Group members (e.g., radio devices) and radio units (e.g., nodes of a RAN) can be wirelessly connected in the uplink (UL), e.g., for feedback to the RL and / or downlink (DL) via the Uu interface. Alternatively or additionally, the SL can optionally implement direct radio communication between neighboring radio devices (e.g., group members and / or local radio units) using the PC5 interface.
[0091] Group members (e.g., radio devices such as UEs) and / or radio units (e.g., nodes of a radio access network RAN such as eNBs or gNBs) and / or the RAN can form or be part of a radio network, e.g., according to the 3rd Generation Partnership Project (3GPP) or according to the standard family IEEE802.11 (Wi-Fi). A first method aspect can be performed by one or more embodiments of a node of the RAN (e.g., a radio unit such as a base station) or a core network (CN) supporting the RAN. Alternatively or additionally, a second method aspect can be performed by one or more embodiments of a group member.
[0092] The RAN can include one or more base stations (e.g., performing the first method aspect). Whenever the RAN is mentioned, the RAN can be implemented by one or more base stations. Alternatively or additionally, the radio network can be a vehicular, ad-hoc, and / or mesh network including two or more radio devices (e.g., acting as group members and / or radio units).
[0093] Any group member can include and / or act as a radio device, e.g., a 3GPP user equipment (UE) or a Wi-Fi station (STA). The radio device can be a mobile device, a device for machine type communication (MTC), a device for narrowband Internet of Things (NB-IoT), or a combination thereof. Examples of UEs and mobile stations include mobile phones or tablet computers operable for navigation and autonomous vehicles. Examples of MTC devices or NB-IoT devices include robots, sensors, and / or actuators in manufacturing, automotive communication, and home automation, for example. MTC devices or NB-IoT devices can be implemented in manufacturing plants, household appliances, and consumer electronics.
[0094] A group member acting as a radio device can be wirelessly connected or connectable to another group member and / or a radio unit (e.g., at least one base station of the RAN) (e.g., according to the radio resource control RRC state or active mode).
[0095] A radio unit can be any station configured to provide radio access to any group member. Any radio unit can be implemented as a network node of a RAN (e.g., a base station or a radio access node), a cell of a RAN, a transmission and reception point (TRP) of a RAN, or an access point (AP). The radio units and / or radio devices within the group member can provide a data link to a host computer (e.g., a navigation server), which provides user data (e.g., a vector field map or information about a destination) to the group member and / or collects user data (e.g., a request to navigate to a destination) from the group member. Examples of base stations can include 3G base stations or Node B (NB), 4G base stations or eNodeB (eNB), 5G base stations or gNodeB (gNB), Wi-Fi APs, and network controllers (e.g., according to Bluetooth, ZigBee, or Z-Wave).
[0096] The RAN can be implemented according to the Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), 3GPP Long Term Evolution (LTE), and / or 3GPP New Radio (NR).
[0097] Any aspect of the technology can be implemented in the physical layer (PHY), media access control (MAC) layer, radio link control (RLC) layer, packet data convergence protocol (PDCP) layer, and / or radio resource control (RRC) layer of a protocol stack for radio communication, and / or in protocol data unit (PDU) layers such as the Internet Protocol (IP) layer and / or the application layer (e.g., for navigation). In this document, referring to the protocol of a layer can also refer to the corresponding layer in the protocol stack. Vice versa, referring to the layer of the protocol stack can also refer to the corresponding protocol of that layer. Any protocol can be implemented by a corresponding method.
[0098] In any aspect of the technology, a Multimedia Broadcast / Multicast Service (MBMS) bearer can be used, for example, in unicast or broadcast mode to send and receive a vector field map and / or a deflection field respectively. The MBMS bearer can be implemented according to 3GPP document TS23.246, version 17.0.0 regarding the MBMS architecture and function description, or 3GPP document TS26.346, version 17.1.0 regarding the MBMS protocol and codec.
[0099] On the other hand, a computer program product is provided. The computer program product includes a program code portion for performing any of the steps of the first and / or second method aspects disclosed herein when the computer program product is executed by one or more computing devices. The computer program product may be stored on a computer-readable recording medium. The computer program product may also be provided for download, for example, via a radio network, a RAN, the Internet, and / or a host computer. Alternatively or additionally, the method may be encoded in a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), or the functionality may be provided for download via a hardware description language.
[0100] Regarding a first device aspect, a device (e.g., a swarm control entity) for controlling a swarm of machines in an area is provided. The area includes a plurality of radio units for providing radio access to the swarm of machines. The swarm of machines includes a plurality of swarm members. The device includes a memory operable to store instructions and a processing circuit (e.g., at least one processor) operable to execute the instructions such that the device is operable to determine a vector field map that includes velocity vectors indicating velocities and directions for navigating the swarm members through the area. The device is also operable to determine a deflection field that indicates deflections for deflecting the swarm members relative to the vector field map. The device is further operable to transmit the vector field map and the deflection field to at least one swarm member via the radio units for controlling the movement of at least one swarm member in the area.
[0101] The device may further be operable to perform any of the steps of the first method aspect.
[0102] Regarding another first device aspect, a device (e.g., a swarm control entity) for controlling a swarm of machines in an area is provided. The area includes a plurality of radio units for providing radio access to the swarm of machines. The swarm of machines includes a plurality of swarm members. The device is configured to determine a vector field map. The vector field map includes velocity vectors indicating velocities and directions for navigating the swarm members through the area. The device is also configured to determine a deflection field. The deflection field indicates deflections for deflecting the swarm members relative to the vector field map. The device is further configured to transmit the vector field map and the deflection field to at least one swarm member via the radio units for controlling the movement of at least one swarm member in the area.
[0103] Alternatively or additionally, the device (e.g., a swarm control entity) includes: a vector field map determination module configured to determine a vector field map that includes velocity vectors indicating the velocity and direction for navigating swarm members through the area; a deflection field determination module configured to determine a deflection field that indicates a deflection for deflecting swarm members relative to the vector field map; and a transmission module configured to transmit the vector field map and the deflection field to at least one swarm member via a radio unit for controlling the movement of at least one swarm member in the area.
[0104] The device (e.g., a swarm control entity) may also be configured to perform any of the steps of the first method aspect.
[0105] Alternatively or additionally, the device (e.g., a swarm control entity) may include at least one radio unit (e.g., a base station) for transmission.
[0106] Regarding a second device aspect, the device (e.g., a swarm member) includes: at least one actuator configured to change the movement state of the swarm member that is part of a swarm of machines moving in an area; a memory operable to store instructions; and a processing circuit (e.g., at least one processor) operable to execute the instructions such that the device is operable to receive a vector field map. The vector field map includes velocity vectors indicating the velocity and direction for navigating swarm members through the area. The device is further operable to receive a deflection field. The deflection field indicates a deflection for deflecting swarm members relative to the vector field map. The device is also operable to determine the position of the swarm member in the area. The device is further operable to determine a change in the movement state based on the received vector field map and the received deflection field for the determined position. The device is also operable to control at least one actuator to achieve the changed movement state.
[0107] The device may further be operable to perform any of the steps of the second method aspect.
[0108] Regarding a further second device aspect, the device (e.g., a swarm member) includes at least one actuator configured to change the movement state of the swarm member that is part of a swarm of machines moving in an area. The device is configured to receive a vector field map. The vector field map includes velocity vectors indicating the velocity and direction for navigating swarm members through the area. The device is also configured to receive a deflection field. The deflection field indicates a deflection for deflecting swarm members relative to the vector field map. The device is also configured to determine the position of the swarm member in the area. The device is also configured to determine a change in the movement state based on the received vector field map and the received deflection field for the determined position. The device is also configured to control at least one actuator to achieve the changed movement state.
[0109] Alternatively or additionally, the device (e.g., a swarm member) includes: a vector field map receiving module configured to receive a vector field map including velocity vectors indicating velocity and direction for navigating the swarm member through the area; a deflection field receiving module configured to receive a deflection field indicating deflection for deflecting the swarm member relative to the vector field map; a position determination module configured to determine the position of the swarm member in the area; a movement state determination unit configured to determine a change in movement state based on the received vector field map and the received deflection field for the determined position; and an actuator control unit configured to control at least one actuator to achieve the changed movement state.
[0110] The device (e.g., a swarm member) may further be configured to perform any of the steps of the first method aspect.
[0111] Alternatively or additionally, the device (e.g., a swarm member) may include a radio device (e.g., a UE) for reception and / or position determination.
[0112] In yet another aspect, a communication system including a host computer is provided. The host computer includes processing circuitry configured to provide user data (e.g., a deflection field for deflection). The host computer further includes a communication interface configured to forward the user data to a cellular network (e.g., at least one of radio units, optionally to the RAN and / or a base station) for transmission to a UE implementing one of the swarm members. The processing circuitry of the cellular network is configured to perform any of the steps of the first and / or second method aspects. The UE includes a radio interface and processing circuitry configured to perform any of the steps of the first and / or second method aspects.
[0113] The communication system may further include a UE. Alternatively or additionally, the cellular network may further include one or more base stations configured to communicate with the UE wirelessly using the first and / or second method aspects and / or provide a data link between the UE and the host computer.
[0114] The processing circuitry of the host computer may be configured to execute a host application to provide user data and / or any of the host computer functions described herein. Alternatively or additionally, the processing circuitry of the UE may be configured to execute a client application associated with the host application.
[0115] Any one of the device, swarm member, radio device, UE, swarm control entity, base station, communication system, or any node or station for implementing the technology may further include any feature disclosed in the context of the method aspect, and vice versa. In particular, any of the units and modules disclosed herein may be configured to perform or initiate one or more steps of the method aspect. Description of the Drawings
[0116] Further details of embodiments of the technology are described with reference to the accompanying drawings, in which:
[0117] Figure 1 A schematic block diagram of an embodiment of a device for controlling a swarm of robots in an area having a plurality of radio units is shown;
[0118] Figure 2 A schematic block diagram of an embodiment of a device for controlling a swarm member having at least one actuator to change its movement state as part of the movement of a swarm of robots in an area is shown;
[0119] Figure 3 A flowchart of a method for controlling a swarm of robots in an area having a plurality of radio units for providing radio access to the swarm of robots is shown;
[0120] Figure 4 A flowchart of a method for controlling a swarm member having at least one actuator to change its movement state as part of the movement of a swarm of robots in an area is shown;
[0121] Figure 5 An overview of swarm control in an area having radio units according to an embodiment is schematically shown;
[0122] Figure 6 An embodiment for integrating artificial intelligence (AI) in the determination of a deflection field or a vector field map is schematically shown;
[0123] Figure 7 Shows the use of reinforcement learning as Figure 6 The detailed functional elements in an embodiment of an AI-assisted determination;
[0124] FIG. 8A to FIG. 8E An architecture of a system according to an embodiment is schematically shown;
[0125] Fig. 9 The calculation of the deflection caused by a deflection field is schematically shown;
[0126] Fig.10 A velocity vector field generated by combining a vector field map with a calculated deflection force field is schematically shown;
[0127] Fig.11 The function for combining a vector field map and a deflection field and its sequential application are schematically shown;
[0128] Fig.12 The calculation of a velocity command at a swarm member according to an embodiment is schematically shown;
[0129] Fig.13Schematically shows another application of the deflection force according to yet another embodiment;
[0130] Figures 14A-14C Schematically shows the effect of unicast transmission to group members;
[0131] Fig.15 Shows a schematic block diagram of an embodiment of a group control entity and group members.
[0132] Fig.15 Shows the implementation of Figure 1 A schematic block diagram of an embodiment of a group control entity of a device;
[0133] Fig.16 Shows the implementation of Figure 2 A schematic block diagram of a group member of a device;
[0134] Fig.17 Schematically shows an example telecommunications network connected to a host computer via an intermediate network;
[0135] Fig.18 Shows a generalized block diagram of a host computer communicating with user equipment via a partial wireless connection through a base station or radio device acting as a gateway; and
[0136] Fig.19 and 20 Shows a flowchart of a method implemented in a communication system that includes a host computer, a base station or radio device acting as a gateway, and a user equipment. Detailed Description
[0137] In the following description, for purposes of explanation and not limitation, specific details are set forth, such as a specific network environment, in order to provide a thorough understanding of the techniques disclosed herein. It will be apparent to those skilled in the art that the techniques may be practiced in other embodiments without these specific details. Additionally, while the following embodiments are primarily described with respect to New Radio (NR) or 5G implementations, it will be apparent that the techniques described herein may also be implemented for any other radio communication technology, including wireless local area network (WLAN) implementations according to the IEEE 802.11 standard family, 3GPP LTE (e.g., LTE-Advanced or related radio access technologies such as MulteFire), Bluetooth according to the Bluetooth Special Interest Group (SIG), particularly Bluetooth Low Energy, Bluetooth Mesh Network, and Bluetooth Broadcast, Z-Wave according to the Z-Wave Alliance, or ZigBee based on IEEE 802.15.4.
[0138] In addition, those skilled in the art will understand that the functions, steps, units, and modules explained herein can be implemented using software acting in conjunction with a programmed microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a general-purpose computer (e.g., including an advanced RISC machine (ARM)). It should also be understood that although the following embodiments are mainly described in the context of methods and devices, the present invention can also be embodied in a computer program product and in a system including at least one computer processor and a memory coupled to at least one processor, where the memory is encoded with one or more programs that can execute the functions and steps or implement the units and modules disclosed herein.
[0139] Figure 1 FIG. shows a schematic block diagram of an embodiment of a device 100 for controlling a swarm of machines, where the device 100 includes a vector field determination module 102, a deflection field determination module 104, and a transmission module 106. The swarm of machines can include a plurality of swarm members configured to move within an area having a plurality of radio units. According to an embodiment, the swarm of machines is controlled based on two fields: a vector field map and a deflection field.
[0140] The vector field determination module 102 is configured to determine a vector field map. The vector field map includes velocity vectors indicating the velocity and direction for navigating the swarm members through the area. According to an embodiment, the vector field map encodes the structure or geometry of the area to guide the swarm members through the area. For example, the vector field map indicates a road for a road vehicle that is a swarm member, or the vector field map indicates the terrain of an area for an aircraft that is a swarm member.
[0141] The structure or geometry of the area can include rigid objects or obstacles that cannot be ignored by the swarm members and affect their movement. Examples of objects or obstacles are: road boundaries or traffic lanes to be followed, buildings, traffic signs, trees, walls, doors, hills, mountains, lakes, hazardous ground structures (e.g., potholes, insufficient friction), or other static objects (non-dynamic) that can define the topology of the area.
[0142] The deflection field determination module 104 is configured to determine a deflection field. The deflection field can indicate a deflection for causing the swarm members to deflect relative to the vector field map. Thus, the deflection field allows for reacting to dynamically changing situations. For example, a moving object or an object that only occasionally appears can be a source of the deflection field so that the swarm members can avoid (e.g., pass by) such non-static obstacles, for example, by changing their course. In addition, the deflection field can also be used to guide (e.g., reroute) the swarm members in a specific direction (e.g., turn left or right) without having to be associated with an obstacle (e.g., a moving object).
[0143] Advantageously, the vector field map can represent the overall or global or static structure in a region. The deflection field can be used to respond to local deviations or dynamic changes (e.g., due to local perturbations or moving obstacles). Following standard notation, a field associates each point or region in a region with a physical quantity. In an embodiment, the vector field map associates each point or region in a region with a vector that can indicate the velocity (e.g., speed and direction of movement) that should be followed (i.e., applied) at that point or region of the region. The deflection field associates each point or region in a region with a vector that can indicate the desired detour around an obstacle or deflection in a certain direction (as indicated by the vector) at the corresponding point region in the region.
[0144] Thus, according to an embodiment, the vector field map can encode changes in the movement state of group members due to static conditions (static obstacles, roads, boundary conditions, etc.) or any changes in the movement state of group members that do not depend on time (e.g., static objects or obstacles), while the deflection field encodes desired changes in the movement state that are time-dependent (valid only for a certain time period) and may not be predictable in advance (or are unpredictable).
[0145] The movement or motion state of a group member can be defined to include information of the group member to describe its kinematic and / or dynamic physical state. This can include one or more of the following information: velocity, direction of movement, brake actuation, acceleration / deceleration, position (e.g., position relative to a reference point in the region (such as a radio unit or global)), altitude, orientation, etc. This information can refer to the current state or may include the upcoming state that may include position information.
[0146] The transmission module 106 is configured to send the vector field map and the deflection field, which are available to one or more group members to navigate in the region. The transmission module 106 can be implemented by or signal-connected to a radio unit to send the vector field map and the deflection field to the group members (e.g., using broadcast, unicast, multicast modes).
[0147] The vector field map and the deflection field can be sent separately, for example, at different time points and / or by different radio units. For example, if the vector field map is encoded with a static or baseline trajectory in the region, the vector field map can be transmitted less frequently than the deflection field. Alternatively or additionally, the vector field map can be transmitted periodically, and / or the deflection field can be transmitted in response to an event (i.e., event-driven, e.g., collision warning or observation of an obstacle).
[0148] According to an embodiment, the transmission module 106 can be configured to send the deflection field only locally (e.g., only through a subset of radio units), while the vector field map can be sent globally (e.g., throughout the region).
[0149] Device 100 can be implemented by a base station (e.g., eNB or gNB) and / or a group control entity. The group members and the group control entity 100 can communicate directly radio, e.g., at least for the transmission of the vector field map and the deflection field. The group members can be implemented by the device 200 mentioned below.
[0150] Figure 2 A schematic block diagram of an embodiment of a device 200 for controlling group members is shown. The device 200 includes a vector field map receiving module 202, a deflection field receiving module 204, a position determination module 206, a movement state determination module 208, and an actuator control module 210. According to an embodiment, the group member includes at least one actuator to be able to change the movement state.
[0151] The vector field map receiving module 202 can be configured to receive a vector field map, which can indicate the velocity vectors for navigating the group member through the area.
[0152] The deflection field receiving module 204 can be configured to receive a deflection field, which can indicate the deflection for deflecting the group member 200 relative to the vector field map.
[0153] The position determination module 206 can be configured to determine the position (e.g., position and / or orientation) of the group member. For this purpose, an embodiment can utilize one or more of the following: an available global positioning system or local positioning system; available sensors (e.g., cameras, radars); signals transmitted from a radio unit (e.g., for azimuth, triangulation, distance measurement based on signal power drop).
[0154] The movement state determination module 208 can be configured to update the movement state based on the received vector field map and the received deflection field. The update can include combining the vector field map with the deflection field to obtain the superposition of the two fields.
[0155] The actuator control module 210 can be configured to cause the group member to follow the updated movement state by correspondingly controlling the actuator. The actuator can change the movement state of the group member, which can for example include one or more of the following: steering, braking, accelerating, altitude adjustment, height adjustment (e.g., of the vehicle chassis), etc. For this purpose, the actuator can be coupled to and control the propulsion unit and / or the braking unit and / or the steering unit of the group member.
[0156] The device 200 can be implemented by a radio device (e.g., UE) and / or the corresponding one of the group members. The radio unit and the group member 200 can communicate directly radio, e.g., at least for the reception of the vector field map and the deflection field. The radio unit can be implemented by the device 100.
[0157] As a result, at least one group member can follow the superposition of two fields, namely the vector field map and the deflection field. As long as both fields are smooth enough (e.g., distinguishable), collisions can be inherently avoided, e.g., because adjacent trajectories resulting from the combination of the fields do not cross. Those skilled in the art know different conditions to ensure this. For example, the area can be divided into a grid or lattice structure.
[0158] For example, the area can be defined as an (n×m) grid, which means the area is constructed from m lines, each line including n square elements (n,m = 1,2,3,...). Then, for each grid square, two vectors can be defined, one indicating the vector field map of the corresponding grid square and one indicating the deflection acting on that grid square. At least some of the deflection vectors can be zero, which means the deflection field acts locally only in the regions where the corresponding deflection vectors are non-zero (so-called deflection zones). Now, when the adjacent vectors of the deflection field never cross, collisions can be avoided, which can in turn be ensured when the length difference of the adjacent vectors of the deflection field is shorter than the grid (lattice) spacing. When a greater deflection is required, the deflection field will be non-zero in a larger area to ensure the desired collision avoidance.
[0159] The area can be a two-dimensional surface or a three-dimensional space. It should be understood that it is not necessary to consider square or cubic lattices to divide the area - any known lattice or tessellation can be used to define or engrave a division structure on the area.
[0160] Figure 3 A flowchart of a method 300 for controlling a swarm of robots in an area is shown, the area having a plurality of radio units for providing radio access to the swarm of robots.
[0161] In step 302, a vector field map is determined, where the vector field map includes velocity vectors indicating the velocity and direction for navigating group members through the area.
[0162] In step 304, a deflection field is determined, where the deflection field indicates the deflection for deflecting group members relative to the vector field map.
[0163] In step 306, the vector field map and the deflection field are transmitted (e.g., in separate messages or in a single message or in a combined message) to at least one group member via the radio unit to control the movement of at least one group member in the area.
[0164] According to an embodiment, steps 302, 304, and 306 of method 300 can be performed by modules 102, 104, and 106, respectively, as Figure 1In particular, steps 302, 304 and 306 may be performed in a group control entity (or center) that controls some or all group members and monitors the operations of the group members.
[0165] Figure 4 A flow chart of a method 400 of controlling a swarm member is shown. A swarm member may include at least one actuator to change its movement state. For example, a swarm member may act as part of a robotic swarm moving in an area.
[0166] In step 402, a vector field map is received, wherein the vector field map indicates velocity vectors for navigating group members through the area.
[0167] In step 404, a deflection field is received indicating a deflection for deflecting group members relative to the vector field map.
[0168] In step 406, the location (eg, position, orientation, speed and / or direction or movement) of the group members in the area is determined.
[0169] In step 408, an update (ie, change) to the movement state is determined based on the received vector field map and the received deflection field. The update may be determined by determining local field values of the vector field map and the deflection field (if any) at the determined location.
[0170] In step 410, at least one actuator is controlled to achieve the updated movement state.
[0171] According to an embodiment, steps 402, 404, 406, 408, and 410 of method 400 may control Figure 2 In particular, steps 402, 404, 406, 408, 410 may be performed by at least one or each swarm member receiving the vector field map and / or the deflection field from the exemplary swarm control entity 100.
[0172] Figure 5An embodiment of a system 500 for controlling a swarm of machines is schematically shown, which includes an embodiment of swarm members 200 in an area 502 having a plurality of radio units 504 and 506. In area 502, the swarm members 200 move along a trajectory 514 derived from a combination of a vector field map 510 and a deflection field 512. The combining step can be performed by each swarm member 200 to derive a velocity vector 201 at any given position along its trajectory. Each swarm member 200 can use the velocity vector to correspondingly control at least one of its actuators to follow the determined velocity vector 201. That is, by adjusting the swarm member velocity vectors to follow the determined velocity vector 201, the individual swarm members 200 follow the trajectory defined by the vector field map 510 and the deflection field 512. In other words, the integral of all the velocity vectors 201 corresponds to the trajectory 514, as Figure 5 schematically shown for three trajectories.
[0173] The plurality of radio units 504 and 506 provide radio access (i.e., radio coverage) in area 502.
[0174] A subset 506 of the radio units (e.g., radio points) is connected to the swarm control entity 100 and provides radio access (i.e., radio coverage) at least in the deflection zone (i.e., where the deflection field 512 is non-zero). Area 502 can be a radio cell associated with a base station 504 connected to a server network 101 (e.g., cloud or edge computing).
[0175] The radio points 506 for transmitting the 306 deflection field 512 can be arranged at or near the obstacle 508. Although Figure 5 an exemplary radio point 506 is shown, the radio units 506 for transmitting the 306 deflection field 512 can also be formed as radio strips, e.g., along a baseline trajectory defined by the vector field map 510.
[0176] Area 502 includes an exemplary obstacle 508, and the swarm control entity 100 is configured to determine the 304 deflection field 512 to deflect the swarming of the swarm members 200 around the obstacle 508. For example, when entering area 502, the swarm members 200 initially follow the direction indicated by the vector field map 510. The swarm members 200 can receive the deflection field 512 from one or more radio points 506, and the deflection field 512 can be non-zero only in the deflection zone around the obstacle 508, e.g., the deflection zone where the periphery of the obstacle 508 intersects the baseline trajectory defined by the vector field map 510.
[0177] Before reaching the obstacle 508, the deflection field 512 can first cause a deflection to the right (observed in the direction of movement of the group member 200), and then turn left to bypass the obstacle 508. Finally, when the obstacle 508 has been passed, the deflection field 512 can cause a right turn again (observed in the direction of movement) to align with the direction indicated by the vector field map 510 again and leave the area 502.
[0178] The area 512 can cover a specific area of interest (e.g., a factory hall), but can also be combined with other areas to cover a larger area (e.g., the road system to a destination). In the latter case, each area 502 can be associated with one or more base stations 504. The area 502 can also represent an indoor area with multiple walls (e.g., as static boundary conditions) and doors through which the group member 200 will move. The vector field map 510 can take into account rigid walls and other rigid obstacles to guide the group member from the starting point to the destination. The deflection field 512 can take into account all dynamic or non-permanent obstacles 508 (e.g., other moving objects).
[0179] According to an embodiment, all objects that may affect the movement of the group member 200 can be assigned to the vector field map 510 or the deflection field 512 to ensure that the group member 200 does not collide with these objects. However, the assignment can be freely chosen. Advantageously, in order to utilize computer resources (e.g., computing power, bandwidth of the network interface) in the most efficient way, all objects that do not change their movement state during the movement of the group member 200 can be encoded in the vector field map 510, and all objects that can change their movement state can be encoded by the deflection field 512.
[0180] Although according to an embodiment the vector field map 510 and the deflection field 512 can be determined as a continuous function of all points in the area 502, computer resources are utilized more efficiently if the area is divided by a grid or a mesh or a cell structure. According to an embodiment, the grid spacing can be constant, but can also depend on the position or time in the area 502. For example, in an area where an obstacle 508 is to be expected (e.g., at a traffic intersection), the spacing can be narrower than in other areas where the collision risk is relatively low.
[0181] Figure 6 Flowcharts of methods 300 and 400 are shown, which have further details for controlling the group member 200 to achieve a desired goal (e.g., reaching a destination without collision).
[0182] In step 602, the process begins. In step 302, the system calculates the dynamic vector field map 510. In this calculation, as described above, the system can consider the (rigid) topology of region 502 to find a route from a starting point to a target (e.g., a destination), which can be optimized based on criteria elaborated later.
[0183] In steps 306, 402, the vector field map 510 is streamed or broadcast into region 502. Optionally and locally, in steps 306, 404, the deflection field 512 is streamed or broadcast. For at least one of these transmissions, evolved Multimedia Broadcast and Multicast Service (eMBMS) can be used.
[0184] As long as no deflection is required or no obstacle occurs, method 400 can continue with step 409 as a sub-step of step 408, where the group member 200 determines a rotation vector (e.g., by calculating the curl or gradient of 409) to deviate from the current motion state to an altered motion state to follow the baseline trajectory defined by the vector field map 510.
[0185] In the presence of a deflection zone or obstacle 508, methods 300 and 400 can proceed to step 408, where the group member 200 determines 409 a rotation vector (e.g., by calculating the curl or gradient) based on the combination 407 of the received 402 vector field map 510 and the received 404 deflection field map 512 to deviate from the current motion state to an altered motion state, thereby following the trajectory according to group control, which includes correction of the deflection field map 512 relative to the baseline trajectory defined by the vector field map 510.
[0186] Alternatively or additionally, when an obstacle 508 or another emergency event occurs in step 604, methods 300 and 400 can include evaluating how to handle the event such that the group member 200 can still proceed safely to the desired target. To this end, in step 606, one or more virtual safety buoys are deployed, e.g., as a sub-step of step 304. In step 608, e.g., as another sub-step of step 304, optionally a virtual safety buoy control agent is started for each deployed virtual safety buoy. In step 304, a deflection force field 512 is determined for the event detected at step 604. According to steps 306, 402, the determined deflection field 512 (e.g., deflection force field) is streamed spatially (e.g., locally within the deflection zone 508) into region 502.
[0187] In step 407, which is a sub-step of step 408, the group member 200 may combine the vector field map 510 and the received deflection field 512, where the combination may be the sum (e.g., vector sum) of both the received fields 510 and 512. However, if the deflection field 512 is determined or received as a force field, a transformation may be used to obtain the corresponding velocity vector (e.g., using the fact that the local velocity is the integral of the acceleration or force).
[0188] In step 409, which is a sub-step of step 408, a rotational vector that also indicates the deviation caused by the event detected in step 604 is calculated, which may be used to control the actuator to perform deflection.
[0189] The sending 306 and receiving 404 of the local deflection field may be implemented using the location services specified by 3GPP, such as according to 3GPP document TS22.261, version 19.0.0; or TS22.071, version 17.0.0; or TS23.273, version 17.6.0.
[0190] In step 612, the cumulative trajectory error is calculated. This trajectory error may be calculated based on the estimated optimal route (e.g., by comparing it with the actual current route). This error may occur if for some reason the group member 200 cannot perform the desired change in its motion state (e.g., insufficient actuator power, wind, slope, etc.).
[0191] According to an embodiment, the trajectory error may be calculated using the location determination module 206 of the group member 200. For example, the group member 200 may be configured to send the calculated trajectory error back to the group control entity 100. Alternatively or additionally, the group control entity 100 may be configured to determine the trajectory error by determining the subsequent location of the group member 200. To this end, the group control entity 100 may utilize the radio units 504, 506 to locate the group member, such as using the tracking of the mobile device according to the 3GPP specification, optionally using 5G positioning according to 3GPP document TS38.455 version 17.2.0.
[0192] In step 610, the control agent policy (e.g., the agent initiated as the virtual safety buoy control agent in step 608) can be retrained based on the calculated error. This step can be performed by the swarm control entity 100. In this retraining 610, the agent will modify the deflection field 512 with the aim of reducing the trajectory error (e.g., in step 608). Thereafter, steps 304, 306 (optionally 402), 404, 406, and 408 are repeated. If the error is still higher than a predetermined threshold in step 612, the loop of steps 610, 608, 304, 306 (optionally 402), 404, 406, and 408 is repeated again to further improve the result. This repetition can continue until the error is acceptably small (e.g., below a predetermined threshold).
[0193] Alternatively or additionally, the initial training of the control agent initiated in step 608 can be performed. This initial training can be performed in simulation or using a digital twin, and can be based on training data to ensure that the agent calculates the deflection field 512 with acceptable accuracy in the field. In this initial training, steps 608, 610 are repeatedly executed to train the agent to generate the best deflection field 512, which allows the swarm members 200 to effectively bypass the obstacle 508 or handle an emergency. As will be elaborated in more detail below, this initial training can be based on artificial intelligence (AI), such as a reinforcement learning (RL) process, to maximize a reward or minimize a loss function (e.g., the time required for all swarm members to bypass an exemplary obstacle), optionally based on a random selection of paths that have previously successfully optimized the reward or loss function.
[0194] Finally, if the deflection field 512 is determined 304 by the swarm control entity 100 and applied by the swarm members (e.g., according to steps 404, 408, 410), which results in the movement of the swarm members 200 corresponding to the best route, method 300 and / or 400 can stop in step 614, or method 300 and / or 400 can continue to control the swarm according to a baseline trajectory (e.g., steps 302, 306, 402, 406, 408, 410) and maintain the deflection (e.g., steps 304, 404).
[0195] The streaming of the vector field map 510 and / or the deflection field 512 in steps 306, 402, 404 can be achieved by leveraging evolved multimedia broadcast and multicast service (eMBMS). This is a mobile radio technology specified by 3GPP, which enables the transmission of multimedia content broadcasts, for example, through the licensed spectrum of 4G Long Term Evolution (4G LTE). Over the years, the consumption of multimedia content has increased. It started from analog broadcasts and then evolved towards digital broadcasts, video on demand, podcasts, and live video streaming. Since the beginning of cellular technology, attempts have been made to provide a seamless user experience for multimedia content.
[0196] eMBMS can transmit data in both unicast mode (i.e., using a dedicated channel between one transmitter and one receiver) and multicast mode (i.e., one transmitter to multiple receivers). This gives operators the option to achieve fast scalability and significant network efficiency gains (when transmitted in multicast mode) while delivering high-quality voice and data services (when transmitted in unicast mode).
[0197] Devices for Internet of Things (IoT) or Machine-to-Machine (M2M) communication are seamlessly connected to a central server or a network of distributed servers (cloud), such as to implement the group control entity 100. Although most IoT devices send and receive very few data packets, the sheer number of IoT devices will exceed the currently available network capacity. eMBMS can enable the efficient transmission of common configurations, commands, software updates to multiple devices. eMBMS provides mechanisms or configuration options by which IoT devices (e.g., implementing group members) can be independently addressed by location. Compared with existing unicast control mechanisms, use cases such as turning street lights on and off require minimal signaling.
[0198] This makes eMBMS a favorable technology for transmitting 306 information to group members 200 in embodiments. The resulting group control can be used as an industrial application in an existing factory or mine, etc., where it is necessary to control a group of robots through a building or another area. Embodiments achieve this by utilizing a vector field map 510 indicating static paths or global navigation in the area and the transmission of a deflection field that provides dynamic (e.g., emergency) control or static avoidance zones for group members.
[0199] Therefore, a prerequisite for the operation of embodiments is that group control is based on the vector field map 510. Group members 200 receive updates to the vector field map 510 (e.g., via eMBMS as described above).
[0200] The physical effects of this technology can include selecting radio units (e.g., radio points 506, see Figure 5 ) participating in the operation, and speed vectors (as part of the vector field map 510) that are calculated and broadcast for a certain mode of operation of group members to be achieved.
[0201] According to embodiments, possible timings can be as follows:
[0202] 1. The group control entity 100 (e.g., an agent) selects radio units 504, 506 (e.g., radio points) that need to participate in the operation, for example as a sub-step of step 304.
[0203] 2. The swarm control entity 100 (e.g., an agent) determines 302 a vector field map 510 to control the points or swarm members 200, thereby implementing a specific trajectory 514 on the swarm members 200.
[0204] 3. The swarm control entity 100 (e.g., an agent) manipulates the swarm members 200 by determining 302 a vector field map 510 and transmitting 306 it by means of the selected radio units 504, 506 (e.g., radio points) (e.g., radio units participating in the eMBMS channel). The selected radio units 504, 506 may use the eMBMS channel to transmit a spatially efficient velocity vector, e.g., a vector field map that does not intersect static obstacles and / or satisfies boundary conditions (e.g., of a road).
[0205] 4. The swarm control entity 100 (e.g., an agent) determines 304 a deflection field 512 (e.g., a deflection velocity vector or a deflection force field).
[0206] 5. The swarm control entity 100 instructs (e.g., explicitly or implicitly by transmitting a deflection force field) the swarm members 200 to combine 407 the vector field map 510 and the deflection field 512 to perform a deflection (e.g., rerouting).
[0207] Method 400 receives and performs a deflection (e.g., rerouting), and may be performed by the corresponding swarm member 200. Method 400 may include the step of generating dynamic information (e.g., velocity commands for the swarm members) from static information (e.g., the vector field map 510 and avoidance zones) in real time.
[0208] Deflection (e.g., rerouting) is a technical effect achieved by embodiments of the technique, where the transmitted information (e.g., the deflection field 512) may depend on at least one obstacle 508.
[0209] According to an embodiment, the swarm members 200 may be of different or the same type. The swarm members 200 may include, for example, one or more of the following various robots: small to large mobile robots, automated guided vehicles (AGVs), unmanned aerial vehicles, bio-inspired bird or insect-like robots, and humanoid robots, autonomous vehicles, or platooning trucks. The technique may be embodied for indoor or outdoor environments.
[0210] The deployment of the radio units 504, 506 (e.g., radio points or radio strips) may determine the resolution of the achievable deflection (e.g., alternative trajectories). Alternatively or additionally, the deflection field 512 (e.g., a deflection force field or a deflection velocity field) may be a spatial field, i.e., the deflection field 512 may be a function of the position of the corresponding swarm member 200 in the region 502.
[0211] Alternatively or additionally, the deflection field 512 can also depend on time, i.e., the deflection field 512 can change over time (e.g., if the obstacle 508 leaves only for a given period of time).
[0212] Thus, the embodiments provide maximum flexibility in adjusting the baseline trajectory to many possible scenarios. Additionally, radio resources are used in an efficient manner, and there is no high demand at the group member 200 for operating in the area 502 controlled by the group control entity 100.
[0213] According to an embodiment, reinforcement learning (RL) is used to generate the vector field map 510 and / or the deflection field 512 assisted by AI. RL is a machine learning training method based on rewarding desired behaviors and / or punishing undesired behaviors. For RL, an agent can perceive and interpret its environment, take actions, and learn through trial and error. RL can be applied to configure the group control entity 100 (e.g., training the agent of the group control entity) to learn a strategy that maximizes the (expected) cumulative reward, e.g., maximizing the number of group members 200 reaching a goal (e.g., a destination). The destination can be a convergence point in the vector field map 510.
[0214] Here, the strategy can be a function (or mapping).
[0215] π:A×S→[0,1]
[0216] π(a,s)=Prob(a t =a∣s t =s)
[0217] For a state s∈S and an action a∈A, it defines the probability of taking action a for state s. The agent will learn a strategy (or tactic) that maximizes the reward. To this end, various reward functions can be defined as measures of desired outcomes, such as the minimum required time, minimum distance, minimum distance, minimum consumption of resources, maximum safety, collision-free, maintaining a maximum safe distance from walls or other members, or a combination thereof. The reward function can be considered the opposite of the loss function (e.g., reversing the sign).
[0218] The agent (e.g., method 300) can interact with the environment (e.g., objects in the area including the group member 200) at discrete time steps t. At each time t, the agent can receive the current state s t and the reward r t-1 associated with the latest state transition (s t-1 ,a t ,s t-1 ,a t-1 ,s t ) caused by the latest action a t .
[0219] Long-term goals help prevent agents from stopping at fewer goals. Over time, agents are trained to avoid negatives and seek positives. This RL is adopted in artificial intelligence (AI) as a way to guide unsupervised machine learning through rewards and punishments.
[0220] An example of RL uses a deep Q-network. In addition to reinforcement learning (RL) techniques, this example also utilizes neural networks. This example exploits the self-directed environmental exploration of RL. Future actions are based on a random sample of past beneficial actions learned by the neural network.
[0221] This technology can be implemented using Ray, such as Ray 2.0.0, which is an open-source project developed by the RISE Lab at the University of California, Berkeley. As a general and universal distributed computing framework, Ray allows for the flexible running of any computationally intensive Python workload, including distributed training or hyperparameter tuning, as well as deep RL and production model serving.
[0222] To implement RL, RLlib within Ray can be used. RLlib is an open-source library for reinforcement learning (RL) that provides support for production-grade, highly distributed RL workloads while maintaining a unified and simple API for various industry applications. RLlib supports training agents in a multi-agent setting, either purely from offline (e.g., historical) datasets or using an externally connected simulator.
[0223] A policy optimizer or policy gradient optimizer can use proximal policy optimization (PPO). PPO is a policy gradient method for reinforcement learning, motivated by algorithms with data efficiency and reliable performance, having the benefits of trust region policy optimization (TRPO) while only using first-order optimization (see, for example, as described in J. Schulman et al., “Proximal Policy Optimization Algorithms”, https: / / arxiv.org / abs / 1707.06347). The clipped objective of PPO supports multiple stochastic gradient descent (SGD) passes through the same batch of experience. RLlib's multi-GPU optimizer fixes the data in GPU memory to avoid unnecessary transfers from host memory, thus significantly improving performance compared to naive implementations.
[0224] Figure 7 An embodiment showing step 302 of determining a vector field map 510 generated with the aid of artificial intelligence (AI) using reinforcement learning (RL) is presented. The same technique can also be used to determine a deflection field 512, for example, as Figure 6 described.
[0225] Reinforcement learning is a form of machine learning based on intelligent agents that should take actions in an environment. Actions are driven by maximizing rewards 724. To this end, the state 722 of the agent (e.g., the movement state of the swarm member 200) and its environment 712 (e.g., the area 512) are defined. Next, a set of actions 730 is defined, such as moving in a specific direction at a given speed. Based on the interaction with the defined environment 712, rewards 724 (or penalties) are given to achieve better results in the next round. Rewards 724 can be given based on criteria such as avoiding an avoidance area (e.g., a dangerous area) such as the deflection area 508, time, energy consumed, and the travel distance required to achieve a goal (e.g., traveling from a starting point to a destination without colliding with other participants or leaving the environment). In particular, during reinforcement learning, specific actions can be randomly selected based on previous successful actions (e.g., all actions that achieve at least a minimum level of reward).
[0226] The result of reinforcement learning is a policy π (indicated at reference numeral 718), which defines appropriate actions for the state 722 of the agent within the environment 722. The action can be a vector in the vector field map 510 and / or the deflection field 512 at a given position or area within the area 512. In other words, the set of all actions can give the vector field map 510 and / or the deflection field 512 determined in step 302.
[0227] After this general process, Figure 7 a process 710 for determining the policy π is shown, where the process 710 includes a loop of steps 712, 714, 716, 718.
[0228] In step 712, the environment is determined, for example, by feedback provided by the swarm member 200 or corresponding data is input. Next, the state s of the swarm member 200 is determined, and based on this state, possible rewards can be assigned (e.g., whether a collision occurs in the determined state, or whether the minimum distance between the swarm member or the swarm member and an obstacle is satisfied). The environment 712 can include at least one serving radio cell 740, for example, operating in a simulator, in real hardware (HW), in HW in the loop, or in a digital twin scenario. In the observation space, the state s can include the current position of the swarm member and (e.g., the global) vector field map.
[0229] In the observation space 720, the reward r is a function that forces the RL to optimize the policy (i.e., the agent). In this use case, avoiding the avoidance area (e.g., the dangerous area) is important. Therefore, a reward (e.g., positive or increasing) is associated with the swarm member 200 avoiding the avoidance area 508, and a reward (i.e., penalty) (e.g., zero or decreasing) is associated with one or more swarm members 200 that do not avoid the avoidance area 508.
[0230] Alternatively or additionally, the reward can be reduced based on a normalized value calculated, for example, as a deviation from an original trajectory (e.g., a baseline trajectory defined only based on the vector field map 510), so as to train the agent to determine 304 a deflection field 512 that keeps the swarm members 200 close to the original behavior.
[0231] A swarm member 200 (or each swarm member) that reaches its destination can be associated with the highest reward (e.g., within a range from 1 to 10 or any other number, optionally many times greater than the absolute value of the negative reward for entering the avoidance zone). If the swarm member 200 (or the corresponding one swarm member 200) does not reach at all (e.g., does not reach the destination defined by the vector field map), the reward can be zero.
[0232] According to an embodiment, the reward can be selected, for example, according to preferences or what is still acceptable and what is not. For example, any situation that may cause human safety problems is unacceptable under any circumstances and may be severely punished. In other cases, a collision that does not cause too much damage is still acceptable.
[0233] In step 714, preprocessing is performed, where based on the current state, the next position to move or the possible direction can be determined.
[0234] In step 716, filtering can be implemented to exclude specific directions that are less favorable (involving too high a cost, e.g., higher than a threshold).
[0235] In step 718, a policy is determined, which allows determining 302 the vector field map 510 and / or the deflection field 512. Since the loop of steps 712, 714, 716, 718 can be executed multiple times for a given environment and / or state 722 of the swarm members 200, the validity of the previously determined vectors (of the vector field map 510 / deflection field 512) can end, or the vector can be replaced by a new (further optimized) vector. In the next round, the action space A (including all actions a) indicated at the reference marker 730 can be used to determine 304 (e.g., modify) the deflection field 512 (e.g., a deflection velocity field or a deflection force field).
[0236] According to an embodiment, the policy evaluation 710 can be pre-executed during the training process based on the trainer class 702. Alternatively or additionally, the policy evaluation 710 can be performed on-site, for example, to further improve the performance of the swarm control entity 100.
[0237] Therefore, according to an embodiment, reinforcement learning (RL) can be used to evaluate the deflection field 512 (e.g., a deflection force field) through machine learning (also known as artificial intelligence or AI).
[0238] Alternatively or additionally, during the evaluation, the deflection field 512 (e.g., a deflection force field) can be manually turned on and off. For example, an operator or user of the swarm members 200 in the monitoring area 502 can evaluate the effect of the deflection field 512 by triggering the generation of the deflection field 512 within an area that the operator / user can indicate. For this purpose, a user interface can be provided in the swarm control entity 100. The user interface can be configured to indicate the area and the type of interference (e.g., the type of obstacle). The user interface can also be used on-site when the user / operator becomes aware of an upcoming obstacle or perturbation to the movement of at least one swarm member 200.
[0239] Alternatively or additionally, for example, a spline function (e.g., approximated by a polynomial) can be used to approximate or estimate the trajectory 514 provided by the deflection field 512 (e.g., a deflection force field) (see, for example, Figure 5 ). The trajectory 514 can be a bypass trajectory around an exemplary obstacle 508 (see, for example, Figure 5 ), and can be locally defined only in the vicinity of the obstacle 508 or the deflection area 508. This offers the advantage that the evaluation of the effectiveness of the deflection field 512 can be checked in a fast and efficient manner, since the spline function is quickly computed on the swarm members 200 and / or clearly depicted for the evaluation. Thus, the determined deflection field 512 can be approved or discarded. Only few resources are required for this purpose.
[0240] In the case of a large number of radio points or radio strips 506, the estimation (or evaluation) of the forces generated by the deflection field 512 can become complex, especially when multiple obstacles 508 are close to each other, such that the deflection fields 512 generated from different obstacles 508 will penetrate each other. As a result, the individual swarm members 200 will combine multiple forces acting from different directions. This situation can be analogized to the gravitational fields generated by the constellations of various celestial bodies in space. However, also here, the field vectors will add up, and the system will evaluate the result based on the evaluated trajectory 514 (e.g., caused by the superposition of the deflection fields 512).
[0241] Although estimating the trajectory 514 of the swarm members 200 may become increasingly complex as their expected travel distance becomes farther and the number of sources of the deflection field 512 increases, embodiments using reinforcement learning provide sufficient resources to account for these effects. In addition to (e.g., static) navigation according to a vector field map, due to the dynamic nature of objects (e.g., obstacles, other swarm members) moving in space, embodiments are also able to incorporate other components of complexity (e.g., the deflection field 512). Similarly, unexpectedly fast objects may appear and temporarily affect the fields 510 and / or 512, and embodiments are also able to handle them.
[0242] The same complexity can occur in a factory environment (e.g., in a hall or a multi-room building). For a static setup, according to an embodiment, the vector field map 510 can be pre-computed. However, a dynamically changing environment complicates problems that can be easily handled by machine learning according to the embodiment.
[0243] During the training process, the training strategy determines 302, 304 an optimized trajectory by broadcasting appropriate velocity vectors (i.e., the deflection field 512, e.g., the deflection velocity field) by the radio points 506. The trajectory can be optimal in several ways. In an embodiment, due to the reward function, the swarm control entity 100 (e.g., the agent of the swarm control entity 100 generated by RL using the reward function) can cause a trajectory that is closest to the original (e.g., baseline) trajectory. Here, the original trajectory can be a trajectory generated by integrating the velocity field map 510 (i.e., without the deflection field).
[0244] In the same or a further embodiment, it can be the case that a complete rerouting of the swarm members 200 is determined by the system (e.g., such that no swarm member 200 enters the avoidance zone). Alternatively or additionally, the swarm control entity 100 can optimize on the shortest paths of the swarm members 200. This is beneficial in terms of energy consumption.
[0245] According to a further embodiment, the system can define short-term rewards (e.g., for successful or failed avoidance and / or minimum deviation) and long-term rewards (e.g., the discounted sum of the short-term rewards and the reward for reaching the destination) for RL.
[0246] An exemplary architecture of the system 500 can include three components of the server network 101 (e.g., the edge cloud), the swarm members 200, and the radio unit 506 for the deflection field 512 (e.g., a safety buoy).
[0247] FIG. 8A to FIG. 8E This architecture with exemplary three components responsible for different aspects is schematically shown.
[0248] The first component 810 can be implemented in a server (e.g., the server network 101, such as the edge cloud) and / or the swarm control entity 100, where the vector field map 510 is determined 302. This calculation can be based on the reachability map 815 (e.g., generated by scanning the environment with LiDAR), which indicates at least one accessible area 817 and / or at least one inaccessible area 819 (e.g., a building, a wall, or other restrictions collectively referred to as boundaries and boundary conditions) or the avoidance zone. The area 502 can be a combination of the reachability map 815 and the avoidance zone 508.
[0249] Fig. 8A An example of the reachability map 815 is shown, where an example of the component 810 of the vector field map 510 is in Figure 8B is shown in
[0250] The vector field graph 510 determined 302 does not cross the inaccessible area 819 (e.g., by aligning the velocity vectors parallel to the boundary) and will point to the target 850 (i.e., the destination). The target 850 can be the final or intermediate destination (also called waypoint) of the group member 200. Thus, the target 850 represents the sink point of the vector field graph 850. This determination 302 can be pre-executed before the group member 200 starts moving through the area 502.
[0251] The second component 820 (an example of which is shown in Figure 8C ) can be implemented in the group control entity 100 or the group member 200 (e.g., if the corresponding group member is an obstacle to another group member), and processes the obstacle 508 or any other deflection area 508. The obstacle 508 can be regarded as the source of the deflection field 512, that is, the vectors of the deflection field 512 point away from the obstacle 508. The deflection field 512 is calculated by the group member entity 100. The deflection field 512 represents the force that routes all group members 200 around the obstacle 508. There are various possible deflection fields 512, and the implemented RL determination 304 optimizes the deflection field 512, which ensures that no group member 200 collides with the obstacle 508 or with each other, and at the same time reaches the target 850 at the minimum cost. RL can achieve this optimization.
[0252] The calculated vector field graph 510 (in the first component 810) and the deflection field 512 (in the first component 820) can be sent, for example, using broadcast or multicast or unicast such as eMBMS.
[0253] The third component 830 can be implemented in (each) group member 200, and combines the two fields, the vector field graph 510 and the deflection field 512, to determine the unique (e.g., velocity) vector to follow at the position of each group member 200, thereby generating a trajectory 514 (e.g., see Figure 5 ). Fig.8D An example of the combined field is shown. When following the depicted vectors, the group member 200 will reach the target 850 without colliding with the obstacle 508 (or avoidance area 508) or the inaccessible area 819.
[0254] Fig. 8EAn exemplary situation in part 840 (on the lower right hand side below) is schematically shown, where an obstacle 508 in the accessible area 817 is another group member (e.g., stationary or moving at a lower speed), optionally another embodiment of the device 200. The system 500, such as the swarm control entity 100 or a swarm member 200, notices that there is a risk that a swarm member 200 may collide with the obstacle 508. To avoid this, according to an embodiment, the system determines a deflection field 512 (not shown), which when combined with a vector field map 510 (not shown) encoding at least one accessible area 817 and / or at least one inaccessible area 819, results in a detour represented by a vector 514. Thus, the swarm member 200 will safely bypass the obstacle 508 without a collision.
[0255] Any embodiment may use reinforcement learning (RL) as an example of a machine learning process. As described in detail with respect to Figure 7 the output of RL is an optimized policy (i.e., the agent of the swarm control entity 100), which is configured to control the correct movement of the swarm members 200 following the velocity vectors broadcast by the radio points 506 (e.g., according to the action space definition 730).
[0256] Furthermore, any embodiment of this technique may use RL, for example, for at least one of the following: selecting the radio unit 506 (e.g., radio point); determining the vector field map 510; and determining the deflection field 512. RL may be implemented by approximate dynamic programming or neural dynamic programming.
[0257] Further details of the first component 810 (e.g., implementation in an edge cloud) may be summarized as follows.
[0258] According to an embodiment, crowd pathfinding and steering using flow field tiles is used as a technique to solve the computational problem of moving, for example, hundreds to thousands of individual agents on a large scale map, as described, for example, in E. Emerson, “Crowd Pathfinding and Steering Using Flow Field Tiles”, 2019. By using dynamic flow field tiles, embodiments implement a steering pipeline with features such as obstacle avoidance, flocking, dynamic formation, crowd behavior, and support for arbitrary physical forces, all without the burden of a heavy central processing unit (CPU) that repeatedly reconstructs separate paths for each swarm member 200. Furthermore, the swarm members 200 can move or react immediately regardless of path complexity, giving immediate feedback to the swarm control entity 100 (e.g., the agent).
[0259] To this end, region 502 can be divided into an n×m grid (grille structure), and for each n×m grid sector, there can be three different n×m 2D arrays or data fields used by pathfinding and steering techniques. These three field types include
[0260] (i) a cost field,
[0261] (ii) an integration field, and
[0262] (iii) a flow field.
[0263] For the cost field (i) that can be achieved by negative short-term rewards, the cost field stores a predetermined "path cost" value for each grid sector (e.g., a square). These values are used as inputs when constructing the integration field. For region 502, the cost can encode one or more of the following: terrain, ground conditions, areas indicating high population density (e.g., having many potential obstacles or many people), and other conditions that affect the overall performance of system 500.
[0264] For the integration field (ii) that can be achieved by negative (partial) long-term rewards, the integration field stores the integrated "cost to target" value for each grid sector and is used as an input when constructing the flow field.
[0265] For the flow field (iii), the flow field includes the path target direction. In other words, each grid sector can be associated with at least one vector indicating the direction to be used on the way to target 850, e.g., implemented as a discretization of vector field diagram 510.
[0266] In cases where the above grid or any other tessellation is used as an input, various algorithms ultimately result in a dynamic vector field 510 that can be used for swarm control according to an embodiment. The dynamic vector field 510 can be extended with multiple functions including general control mechanisms such as obstacle avoidance, swarm behavior, and dynamic formation already mentioned. All of these routings are valid for each agent without recalculation. Thus, agents according to an embodiment can respond quickly to changes regardless of the complexity of the journey.
[0267] Periodically update swarm members 200 on map updates using unicast transmissions for those members that do not support multicast or using MBMS.
[0268] Any embodiment in any aspect can implement MBMS according to 3GPP using at least one of the following features.
[0269] In a first variant, the group control entity 100 (e.g., a server) or method 300 communicates with the UEs 200 containing group members on the downlink (DL) using a unicast bearer at the start of a group communication session. When the group control server 100 triggers the use of an MBMS bearer in the Evolved Packet System (EPS) for DL Vertical Application Layer (VAL) service communication, the Network Resource Model (NRM) server (also referred to as the Management Information Model MIM server) decides to establish an MBMS bearer in the EPS using the procedures defined in 3GPP document TS23.468 version 17.0.0. For example, a vehicle communication application using vehicle-to-any communication (V2X application) is an example of a VAL service. The NRM server provides the UEs 200 with MBMS service description information associated with one or more MBMS bearers obtained from the BM-SC. The UE 200 starts using one or more MBMS bearers to receive the DL VAL service and stops using the unicast bearer for DL group control server communication, e.g., according to 3GPP document TS23.434 regarding "Service Enabler Architecture Layer for Verticals (SEAL)", version 18.2.0, clause 14.3.4.3 regarding "Use of dynamic MBMS bearer establishment".
[0270] In a second variant that can be combined with the first variant, for the radio resource efficient transmission 302 of the vector field diagram 510 and / or the transmission 304 of the deflection field 512, e.g., the transmission 302 of the change (i.e., update) of the vector field diagram 510 and / or the transmission 304 of the change of the deflection field 512, advanced video coding can be applied, e.g., using the codec H.264 as described in "Advanced video coding for generic audiovisual services", https: / / www.itu.int / rec / T-REC-H.264.
[0271] Aspects of the embodiments implemented in the group members 200 can be summarized as follows:
[0272] Group member 200 receives the dynamic vector field 510 via unicast or multicast transmission. After successfully positioning itself on the vector field map 510, each group member 200 calculates its group member velocity vector based on the velocity vector received from the group control entity 100 (e.g., an agent). When added to the current position (or orientation), the resulting vector moves the group member from its previous orientation to the desired orientation (i.e., the orientation indicated by the vectors of the vector field). The vector field map 510(x, y) can provide a common velocity vector for the positions (x, y) of the respective group members 200. The corrected or updated velocity vector of the group member 200 at the position (x, y) can be calculated as follows:
[0273] Group member velocity vector(t) =
[0274] Vector field map(x(t), y(t)) + Deflection velocity field(x(t), y(t))(1)
[0275] Group member position(t + 1) =
[0276] Group member position(t) + Group member velocity vector(t)·Δt(2)
[0277] where Δt is the time increment used by the system, which can be set to 1. The group member velocity vector 201 is as Figure 5 shown.
[0278] The motion control system of a corresponding one of the group members 200 (e.g., on the group member) moves the corresponding one of the group members 200 in the direction indicated by the corresponding group member velocity vector 201, for example, according to equation (1). Alternatively or additionally, the motion control system moves each group member 200 to the calculated group member position, for example, according to equation (2).
[0279] In a variant of any one of the embodiments, deflection (e.g., rerouting) is caused at the force level (or acceleration, i.e., the rate of change of velocity) according to the deflection force field, for example, as opposed to the correction of the vector field map at the velocity level (e.g., by superimposing the deflection velocity field). For this purpose, equations 2(1) and (2) above are modified to read
[0280] Group member velocity vector(t + Δt) = (1’)
[0281] Deflection force field(x(t), y(t))·Δt
[0282] + Fundamental steering force according to the gradient of the velocity field map·Δt
[0283] + Group member velocity vector(t)
[0284] Group member position(t + Δt) = (2’)
[0285] The position of the group member (t) + the velocity vector of the group member (t + Δt)·Δt.
[0286] Where Δt is still the time increment used by the system, and "(..)" refers to the argument of a function or field.
[0287] In any embodiment, at least one of the following features or steps can be used to implement any feature or step related to implementing the deflection field for a safety buoy (e.g., determination 304 and transmission 306).
[0288] The safety buoy is used in an embodiment to achieve collision avoidance with respect to the obstacle 508. Such a collision avoidance mechanism can determine the (repulsive) force to maintain a clearance (i.e., minimum distance) from the obstacle 508. Thus, the RL or group control entity 100 (i.e., also referred to as an agent) controls the group members 200 (i.e., which can also include agents), for example, when the obstacle appears to block its path (e.g., the spatial part of the trajectory in the phase space), forcing the group members 200 to avoid the obstacle 508.
[0289] Since the agent or group member 200 only follows vectors, it is not necessary to install, for example, collision sensors, or the group member 200 does not need to be able to generate appropriate parameters for avoidance maneuvers in real time. Continuous querying of the centrally maintained vector field (e.g., according to the combined fields 510 and 512 of 407) and execution of the correction procedure 409 can be sufficient to achieve collision avoidance. To avoid collisions, embodiments treat all obstacles 508 (or deflection zones 508) as simple geometric shapes. A common solution is to use a spherical shape (a two-dimensional circle or a three-dimensional sphere).
[0290] Fig. 9 An embodiment for determining and performing the deflection 409 caused by the deflection field 512 around the obstacle 508 with the obstacle center 508a is schematically shown.
[0291] The original velocity vector of the vector field 510 (simply referred to as the original velocity vector 510) is extended to the extended vector 510a to check its proximity to the obstacle 508. By using this extension, the group member 200 can start an avoidance maneuver in a timely manner. If the trajectory of its route initially crosses the obstacle 508 or the deflection zone, the original vector 510 should be rotated 409. To avoid collisions, the original velocity vector 510 is rotated in the direction of the deflection force 512 so that the group member avoids the obstacle 508 along its gradient vector 514 (e.g., see Fig. 9 ).
[0292] The deflection force acting on the group member can be a scaled maximum deflection force That is, it can be restricted by defining a normalized deflection force (e.g., a value between 0 and 1) multiplied by the maximum deflection force. The maximum deflection force can be selected based on specific circumstances (e.g., the capabilities of the group members, the density of the group members, the mass of the group members, and other factors).
[0293] For example, the deflection force can be calculated as follows:
[0294]
[0295] The normalization can be linear in the distance d, e.g., (1 - d / D), or inversely proportional to the distance d, e.g., D / d.
[0296] As a specific example, the deflection force can be calculated as follows:
[0297]
[0298] where d is the (future) minimum distance between the group member and the center of the obstacle if the trajectory of the group member is not deflected, e.g.,
[0299]
[0300] The safety buoy can be configured to define or determine a deflection (e.g., force) vector field 512. The force field can be calculated based on the repulsive force of the obstacle 508. The calculation of the force can be implemented according to J. Barraquand, B. Langlois, and J.-C. Latombe, "Numerical potential field techniques for robot path planning" in IEEE Transactions on Systems, Man, and Cybernetics, Vol. 8, pp. 224 - 241, March - April 1992, doi:10.1109 / 21.148426.
[0301] Fig.10 An embodiment of the velocity vector field 1000 of the group generated by step 407 of the combined deflection field 512 and the vector field map 510 is schematically shown.
[0302] In the case of an emergency route change, the safety buoy 506 is deployed. The safety buoy uses a broadcast communication method (e.g., using point - to - point 5G, Wi - Fi, or Li - Fi) to transmit the deflection force field 512 spatially around its vicinity.
[0303] The received deflection force field 512 is added 407 to the vector field map 510 (e.g., a dynamic vector field), e.g., by each group member 200 at their respective positions, or for all points of the map grid, as Fig.11It is schematically shown. Note that the repulsiveGrad component corresponds to the deflection field 512 (i.e., originating from the safety buoy 506), while the attractiveGrad component corresponds to the vector field map 510 (i.e., representing the destination 850, and optionally other waypoints).
[0304] In this embodiment, the attractive gradient can be the vector field map 510, which can in this case be derived as the gradient of a scalar potential.
[0305] Fig.11 Functionally and in the order of their application, an exemplary implementation of step 408 for determining a change in the movement state is schematically shown. The vector field map 510 is determined as the gradient from a first scalar potential (referred to as the attractive gradient). The deflection field 512 is determined as the gradient from a second scalar potential (referred to as the repulsive gradient).
[0306] As Fig.11 schematically shown, the gradient vector of the movement state change is calculated based on the attractive gradient and the repulsive gradient.
[0307] In step 410, the swarm member 200 calculates and applies a speed command (e.g., a speed value) for a certain actuator, e.g., the speed command needs to be applied to the rotor or wheels of the swarm member 200 to move in the desired direction.
[0308] Fig.12 An exemplary implementation of step 410 for determining the speed command is schematically shown. When step 410 is executed, the corresponding unit 210 receives the gradient vector field. By default, the starting point of the vector is in the coordinate system of the robot itself, so it shows the direction relative to the robot. In order to make the end point of the vector a navigation point, it must be transformed into the coordinate system of the map.
[0309] Fig.13 Another application of the deflection force in the deflection zone 508 is schematically shown. As Fig.13 schematically shown, the deflection field 512 can be uniform (i.e., parallel within the deflection zone 508), and does not necessarily decrease with the distance to the center 508a of the deflection zone 508.
[0310] According to the first embodiment, the emitter of the deflection field 512 can be at least one of the following: already deployed and existing radio units 506, such as radio points or radio strips in a factory cell; a dedicated device deployed at the center 508a of the deflection zone 508 as a safety buoy; and the swarm members 200 of the swarm of machines.
[0311] According to the second embodiment, the transmission 306 may use at least one of the following: an existing MBMS channel, optionally, where the operation of certain radio points broadcasts a locally effective velocity vector on the MBMS channel, for example, according to the MBMS SEAL process; and optionally local broadcast or peer-to-peer transmission over 5G, Wi-Fi, Li-Fi, etc.
[0312] According to the third embodiment, the obstacle center 805a is necessary information for calculating the deflection force 512. In the absence thereof, the swarm members 200 may apply the received velocity vector 1000 to their current velocity, but it is important to note that in this case, the velocity vector representing the entire coverage area of the deflection zone 508 (e.g., a cell) will have very much the same velocity vector as that broadcast by the deflection field 512. The natural effect (which decreases as a function of distance) in the case of an electric field and a gravitational field is absent.
[0313] Fig.13 The default case of the deflection force field 512 is schematically shown. The velocity vector may be broadcast at various time values, thus producing some pulsation effect. There may be positive and negative directions. All swarm members 200 receive the same field 512. If it is necessary to consider them based on the distance from the buoy, for example, to make the swarm turn smaller and the cell edge smaller, then, for example, the signal-to-noise ratio (SNR) may be a weighting factor in the deflection force calculation 408 or 410. Alternatively or additionally, the transmitted signal strength may affect the size of the deflection zone 508.
[0314] Figures 14A-14C The effect of unicast transmission to the swarm members is schematically shown. In the case of unicast transmission 306 (and the corresponding reception 404), the number of connected swarm members 200 may affect the alternate routing on the swarm members 200.
[0315] For example, as Fig.14A shown, a cell may transmit a left vector 512 every 10 ms. If 2 swarm members are connected, then as Fig. 14B shown, the cell transmits to each swarm member for 20 ms at full load. Therefore, these swarm members 200 receive fewer left-turning vectors than Fig.14A in the first case. Fig. 14B And Fig. 14C show the effect of unicast transmission.
[0316] In Fig. 14CIn it, the horizontal chain of the vector does not indicate the direction of the vector, but instead indicates time division multiplexing of unicast transmissions according to the pattern to group members 1, 2, and 3 respectively. Unicast conversion may require attaching the corresponding group member 200 to the radio point 506. The transmission capacity is distributed among the group members 200. The pulse rate of the deflection field 512 decreases as the number of connected group members 200 increases. Therefore, in some cases, broadcasting and / or multicasting is preferred.
[0317] In any embodiment, buoys 506 (e.g., manually deployed) can encode their direction (e.g., obtained from their compass), e.g., encoded in the access point (AP) name or service set identifier (SSID).
[0318] In a fourth implementation, a single central agent (e.g., at the group control entity 100) or a collaborative multi-agent deployment (e.g., including agents at each group member 200) can be implemented. Training can be performed in a real-time deployment, e.g., if the group members 200 have computational resources for training and the sensors resume during the time the policy is being trained to find the best deflection (e.g., rerouting) of the group members 200.
[0319] In a fifth embodiment, the center of the alternate route or the curvature of the deflection does not necessarily have to be the geometric center of the radio point. A position offset can also be encoded into the velocity vector.
[0320] Any embodiment can implement collision avoidance. There are two main cases regarding collision avoidance: 1) groups and 2) AGVs and UAVs with more advanced sensors and processing capabilities.
[0321] A swarm with minimal sensor information can be achieved according to S. Mayya, P. Pierpaoli, G. Nair, and M. Egerstedt, “Localization in Densely Packed Swarms Using Interrobot Collisions as a Sensing Modality,” in IEEE Transactions on Robotics, vol. 35, no. 1, pp. 21–34, Feb. 2019, doi: 10.1109 / TRO.2018.2872285. Their proposal is that a less conservative coordination control strategy can be adopted to achieve collision avoidance in the swarm, where collisions are not only tolerated but can potentially be used as a source of information. In the paper, they take collisions as a sensing modality for providing information about the robot's surrounding environment as the research direction. They envision a collection of mobile robots with no sensors other than binary tactile sensors that can determine whether a collision has occurred, and let the robots use this information to determine their positions. They apply a probability localization technique based on mean-field approximation, which allows each robot to maintain and update the probability distribution over all possible positions. Simulated and real multi-robot experiments illustrate the feasibility of the proposed method.
[0322] Alternatively or additionally, collision avoidance can be achieved according to Seyed Zahir Qazavi, Samaneh Hossesini Semnani, “Distributed Swarm Collision Avoidance Based on Angular Calculations,” https: / / arxiv.org / abs / 2108.12934, which proposes Angular Swarm Collision Avoidance (ASCA) as an algorithm for motion planning for large teams of agents. ASCA is distributed, real-time, low-cost, and based on full robots in two- and three-dimensional spaces. In this algorithm, each agent calculates its direction of movement at each time step based on its own sensing (knowing the relative positions of other agents / obstacles), i.e., each agent does not need to know the states of neighboring agents. The proposed method calculates the possible interval of its movement at each step and then quantifies its speed magnitude and direction based on this. ASCA is parameter-free and only requires robot and environmental constraints, such as the maximum allowable speed of each agent and the minimum possible separation distance between agents. It is shown that ASCA is faster in simulation compared to the prior art algorithms ORCA and FMP.
[0323] In these cases, local decisions on the robot to avoid collisions do not require communication with external sources.
[0324] AGVs or UAVs with sensors can be implemented using existing products for collision avoidance with specific task sensors, such as according to https: / / www.sick.com / au / en / end-of-line-packaging / automated-guided-vehicle-agv / collision-avoidance-on-an-automated-guided-vehicle-agv / c / p514346.
[0325] Alternatively or additionally, (1) three-dimensional collision avoidance can be achieved by controlling the height of the swarm members 200 and by swarm rules in two or three dimensions, implicitly implemented by, for example, the locally broadcast deflection field 512.
[0326] Since the deflection field 512 can be transmitted locally and since the combination of fields 510 and 512 can be implicitly collision-free, the radio resources for controlling the swarm are used efficiently, which also means energy savings.
[0327] Preferably, in any embodiment, the radio unit 506 that is the source of the deflection field 512 is static or stationary.
[0328] The position (or orientation) of the swarm members for feedback in training can be based on simulation or digital twin, for example, if no position is received from real-world swarm members. Alternatively or additionally, unicast feedback or camera measurements of the positioning (or location) can be implemented.
[0329] Optionally, a training strategy to determine 304 a deflection field and apply the strategy to each radio unit 506 (e.g., radio point).
[0330] This technology can be applied to uplink (UL), downlink (DL), or direct communication between radio devices, such as device-to-device (D2D) communication or sidelink (SL) communication.
[0331] Each of the transmitting station 100 and the receiving station 200 can be a radio device or a base station. Here, any radio device can be a mobile or portable station and / or any radio device that can be wirelessly connected to a base station, a RAN, or another radio device. For example, the radio device can be a user equipment (UE), a device for machine type communication (MTC), or a device for the Internet of Things (IoT) (e.g., narrowband). Two or more radio devices can be configured to be wirelessly connected to each other, for example, in an ad-hoc radio network or via a 3GPP SL connection. In addition, any base station can be a station that provides radio access, can be part of a radio access network (RAN), and / or can be a node connected to the RAN for controlling radio access. For example, the base station can be an access point, such as a Wi-Fi access point.
[0332] In this document, whenever noise or signal-to-noise ratio (SNR) is mentioned, corresponding steps, features, or effects for noise and / or interference or signal-to-interference-and-noise ratio (SINR) are also disclosed.
[0333] Fig.15 A schematic block diagram of an embodiment of the device 100 is shown. The device 100 includes processing circuitry, for example, one or more processors 1504 for executing the method 300 and a memory 1506 coupled to the processor 1504. For example, the memory 1506 can be encoded with instructions for implementing at least one of the modules 102, 104, and 106.
[0334] One or more processors 1504 can be a microprocessor, a controller, a microcontroller, a central processing unit, a digital signal processor, an application specific integrated circuit, a field programmable gate array, or a combination of one or more of any other suitable computing devices, resources, or a combination of hardware, microcode, and / or encoded logic, which can operate to provide the transmitter function and / or the function of the group control entity 100, either alone or in combination with other components of the device 100, such as the memory 1506. For example, one or more processors 1504 can execute instructions stored in the memory 1506. Such functions can include providing the various features and steps discussed herein, including any benefits disclosed herein. The expression "the device is operable to perform an action" means that the device 100 is configured to perform the action.
[0335] As Fig.15 Schematically shown, the device 100 can be embodied by a group control entity 1500, such as a transmitting base station or a transmitting UE. The transmitting station 1500 includes an interface 1502 (e.g., radio) coupled to the device 100 for radio communication with one or more stations, such as a radio unit 504, 506, or a UE implementing the group member 200.
[0336] Fig.16 FIG. 1 is a schematic block diagram showing an embodiment of device 200. Device 200 includes processing circuitry, e.g., one or more processors 1604 for executing method 400 and a memory 1606 coupled to processor 1604. For example, memory 1606 may be encoded with instructions for implementing at least one of modules 202, 204, 206, 208, and 210.
[0337] One or more processors 1604 may be a microprocessor, controller, microcontroller, central processing unit, digital signal processor, application specific integrated circuit, field programmable gate array, or any other suitable combination of one or more of computing devices, resources, or a combination of hardware, microcode, and / or encoded logic that can operate alone or in combination with other components of device 200, such as memory 1606, to provide UE functionality or the functionality of group member 200. For example, one or more processors 1604 may execute instructions stored in memory 1606. Such functionality may include providing the various features and steps discussed herein, including any benefits disclosed herein. The statement that "the device is operable to perform an action" means that device 200 is configured to perform the action.
[0338] As Fig.16 shown, device 200 may be implemented as group member 1600, e.g., as a receiving UE. Group member 1600 includes a radio interface 1602 coupled to device 200 for radio communication with one or more transmitting stations, e.g., acting as a transmitting base station or a transmitting UE.
[0339] Referring Fig.17 , according to an embodiment, communication system 1700 includes a telecommunications network 1710, such as a 3GPP-type cellular network, which includes an access network 1711 (such as a radio access network) and a core network 1714. Access network 1711 includes a plurality of base stations 1712a, 1712b, 1712c, such as NB, eNB, gNB, or other types of wireless access points, each base station defining a corresponding coverage area 1713a, 1713b, 1713c. Each base station 1712a, 1712b, 1712c may be connected to core network 1714 via a wired or wireless connection 1715. A first user equipment (UE) 1791 located in coverage area 1713c is configured to wirelessly connect to or be paged by corresponding base station 1712c. A second UE 1792 in coverage area 1713a may wirelessly connect to corresponding base station 1712a. Although multiple UEs 1791, 1792 are shown in this example, the disclosed embodiments are equally applicable to the case where a single UE is in the coverage area or a single UE is connected to corresponding base station 1712.
[0340] Any base station 1712 may embody at least one of radio units 504 and 506 and / or group control entity 100. Any UE 1791, 1792 may embody group member 200.
[0341] Telecommunication network 1710 is itself connected to host computer 1730, which may be implemented in hardware and / or software of an independent server, a cloud-implemented server, a distributed server, or as processing resources in a server farm. Host computer 1730 may be under the ownership or control of a service provider, or may be operated by or on behalf of a service provider. The connections 1721, 1722 between telecommunication network 1710 and host computer 1730 may extend directly from core network 1714 to host computer 1730, or may go via an optional intermediate network 1720. Intermediate network 1720 may be one or a combination of more than one of a public, private, or hosted network; intermediate network 1720 (if any) may be a backbone network or the Internet; in particular, intermediate network 1720 may include two or more sub-networks (not shown).
[0342] Fig.17 Communication system 1700 as a whole enables a connection between one of the connected UEs 1791, 1792 and host computer 1730. The connection may be described as an over-the-top (OTT) connection 1750. Host computer 1730 and the connected UEs 1791, 1792 are configured to transmit data and / or signaling via OTT connection 1750, using access network 1711, core network 1714, any intermediate network 1720, and possibly other infrastructure (not shown) as mediators. In the sense that the participating communication devices through which OTT connection 1750 passes are not aware of the routing of the uplink and downlink communications, OTT connection 1750 may be transparent. For example, it is not necessary to inform base station 1712 of the past routing of incoming downlink communications having data originating from host computer 1730 to be forwarded (e.g., switched) to the connected UE 1791. Similarly, base station 1712 does not need to know the future routing of outgoing uplink communications from UE 1791 to host computer 1730.
[0343] By means of method 200 performed by any one of UEs 1791 or 1792 and / or any one of base stations 1712, the performance or range of OTT connection 1750 may be improved, for example, in terms of increased throughput and / or reduced latency. More specifically, host computer 1730 may indicate to system 500 (such as group control entity 100 or group member 200 (e.g., at the application layer)) at least one of vector field map 510 and deflection field 512. For example, host computer may determine 302 vector field map 510 so as to deliver packets to an address, for example, according to an online order.
[0344] Now, a reference will be made to Fig.18 describe an example implementation of embodiments of the UE, base station, and host computer discussed in the previous paragraphs. In communication system 1800, host computer 1810 includes hardware 1815, and hardware 1815 includes a communication interface 1816 configured to establish and maintain a wired or wireless connection with interfaces of different communication devices of communication system 1800. Host computer 1810 further includes processing circuitry 1818, which may have storage and / or processing capabilities. In particular, processing circuitry 1818 may include one or more programmable processors suitable for executing instructions, application specific integrated circuits, field programmable gate arrays, or a combination thereof (not shown). Host computer 1810 also includes software 1811 stored in or accessible by host computer 1810 and executable by processing circuitry 1818. Software 1811 includes host application 1812. Host application 1812 may be operable to provide services to remote users, such as UE 1830 connected via an OTT connection 1850 terminated at UE 1830 and host computer 1810. When providing services to remote users, host application 1812 may provide user data transmitted using OTT connection 1850. The user data may depend on the location of UE 1830. The user data may include supplementary information or targeted advertisements (also: ads) delivered to UE 1830. The location may be reported to host computer by UE 1830, for example, using OTT connection 1850 and / or by base station 1820, for example, using connection 1860.
[0345] Communication system 1800 further includes base station 1820, which is disposed in a telecommunication system and includes hardware 1825 enabling it to communicate with host computer 1810 and UE 1830. Hardware 1825 may include a communication interface 1826 for establishing and maintaining a wired or wireless connection with interfaces of different communication devices of communication system 1800, and a radio interface 1827 for establishing and maintaining at least a wireless connection 1870 with UE 1830 located in a coverage area ( Fig.18 not shown) served by base station 1820. Communication interface 1826 may be configured to facilitate connection 1860 to host computer 1810. Connection 1860 may be direct, or it may pass through a core network of the telecommunication system ( Fig.18(not shown) and / or via one or more intermediate networks external to the telecommunications system. In the illustrated embodiment, the hardware 1825 of the base station 1820 further includes processing circuitry 1828, which may include one or more programmable processors suitable for executing instructions, application specific integrated circuits, field programmable gate arrays, or combinations thereof (not shown). The base station 1820 also has software 1821 stored internally or accessible via an external connection.
[0346] The communication system 1800 also includes the UE 1830 already mentioned. Its hardware 1835 may include a radio interface 1837, which is configured to establish and maintain a wireless connection 1870 with a base station serving the coverage area where the UE 1830 is currently located. The hardware 1835 of the UE 1830 also includes processing circuitry 1838, which may include one or more programmable processors suitable for executing instructions, application specific integrated circuits, field programmable gate arrays, or combinations thereof (not shown). The UE 1830 also includes software 1831, which is stored in the UE 1830 or accessible by the UE 1830 and executable by the processing circuitry 1838. The software 1831 includes a client application 1832. The client application 1832 may be operable to provide services to a human or non-human user via the UE 1830 with the support of the host computer 1810. In the host computer 1810, the execution of the host application 1812 may communicate with the execution of the client application 1832 via an OTT connection 1850 terminated at the UE 1830 and the host computer 1810. When providing services to the user, the client application 1832 may receive request data from the host application 1812 and provide user data in response to the request data. The OTT connection 1850 may convey both the request data and the user data. The client application 1832 may interact with the user to generate the user data it provides.
[0347] Note that Fig.18 the host computer 1810, base station 1820, and UE 1830 described in Fig.17 may be the same as one of the host computers 1730, base stations 1712a, 1712b, 1712c, and one of the UEs 1791, 1792, respectively. That is, the internal workings of these entities may be as Fig.18 shown, and independently, the surrounding network topology may be Fig.17 the network topology of
[0348] In Fig.18In [the figure], the OTT connection 1850 is abstractly depicted to show the communication between the host computer 1810 and the UE 1830 via the base station 1820, without explicitly referring to any intermediate devices and the exact routing of messages through these devices. The network infrastructure can determine the routing, which can be configured to hide it from the UE 1830 or the service provider operating the host computer 1810 or both. When the OTT connection 1850 is active, the network infrastructure can further make a decision by which it dynamically changes the routing (e.g., based on load balancing considerations or reconfiguration of the network).
[0349] The wireless connection 1870 between the UE 1830 and the base station 1820 is in accordance with the teachings of the embodiments described throughout this disclosure. One or more of the various embodiments improve the performance of the OTT service provided to the UE 1830 using the OTT connection 1850, where the wireless connection 1870 forms the last hop. More precisely, the teachings of these embodiments can reduce latency and increase data rate, thus providing benefits such as better responsiveness and improved QoS.
[0350] A measurement process can be provided for the purpose of monitoring data rate, latency, QoS, and other factors improved by one or more embodiments. There can also be an optional network function for reconfiguring the OTT connection 1850 between the host computer 1810 and the UE 1830 in response to changes in the measurement results. The measurement process and / or the network function for reconfiguring the OTT connection 1850 can be implemented in the software 1811 of the host computer 1810 or the software 1831 of the UE 1830 or both. In an embodiment, sensors (not shown) can be deployed in or associated with the communication devices through which the OTT connection 1850 passes; the sensors can participate in the measurement process by providing values of the monitored quantities illustrated above or other physical quantities from which the software 1811, 1831 can calculate or estimate the monitored quantities. The reconfiguration of the OTT connection 1850 can include message format, retransmission settings, preferred routing, etc.; the reconfiguration need not affect the base station 1820, and the base station 1820 may not know or be aware of it. Such processes and functions can be known and practiced in the art. In certain embodiments, the measurement can involve proprietary UE signaling that facilitates the measurement of throughput, propagation time, latency, etc. by the host computer 1810. The measurement can be implemented because the software 1811, 1831 causes messages, especially empty or "virtual" messages, to be sent using the OTT connection 1850 while it monitors propagation time, errors, etc.
[0351] Fig.19 is a flowchart showing a method implemented in a communication system according to one embodiment. The communication system includes a host computer, a base station, and a UE, which may be reference Fig.17 and18 Those described. For simplicity of the present disclosure, only references to Fig.19 will be included in this paragraph. In a first step 1910 of the method, the host computer provides user data. In an optional sub-step 1911 of the first step 1910, the host computer provides user data by executing a host application. In a second step 1920, the host computer initiates a transmission carrying the user data to the UE. In an optional third step 1930, in accordance with the teachings of the embodiments described throughout the present disclosure, the base station sends the user data carried in the transmission initiated by the host computer to the UE. In an optional fourth step 1940, the UE executes a client application associated with the host application executed by the host computer.
[0352] Fig. 20 is a flowchart showing a method implemented in a communication system according to an embodiment. The communication system includes a host computer, a base station, and a UE, which may be those referred to in Fig.17 and 18 described. For simplicity of the present disclosure, only the reference numerals of Fig. 20 will be included in this paragraph. In a first step 2010 of the method, the host computer provides user data. In an optional sub-step (not shown), the host computer provides user data by executing a host application. In a second step 2020, the host computer initiates a transmission carrying the user data to the UE. In accordance with the teachings of the embodiments described throughout the present disclosure, the transmission may be relayed via the base station. In an optional third step 2030, the UE receives the user data carried in the transmission.
[0353] It is apparent from the above description that at least some embodiments of the technology allow for an improved selection of relay radio devices and / or an improved selection of SL connection establishment. The same or additional embodiments may ensure that the traffic relayed by the relay radio device is given appropriate QoS handling.
[0354] Many advantages of the present invention will be fully understood from the foregoing description, and it is apparent that various changes may be made in the form, construction, and arrangement of the units and devices without departing from the scope of the present invention and / or sacrificing all of its advantages. Since the present invention may be varied in many ways, it will be recognized that the present invention should be limited only by the scope of the appended claims.
Claims
1. A method (300) for controlling a swarm of robots in a control area (502), the area (502) including a plurality of radio units (504, 506) for providing radio access to the swarm of robots, the swarm of robots including a plurality of swarm members (200; 1600; 1791; 1792; 1830), the method (300) including or initiating: Determining (302) a vector field map (510), the vector field map (510) including velocity vectors indicating the velocity and direction for navigating the swarm members (200; 1600; 1791; 1792; 1830) through the area (502); Determining (304) a deflection field (512), the deflection field (512) indicating a deflection for deflecting the swarm members (200; 1600; 1791; 1792; 1830) relative to the vector field map (510); and Transmitting (306) the vector field map (510) and the deflection field (512) to at least one swarm member (200; 1600; 1791; 1792; 1830) via the radio units (504, 506) to control the movement of the at least one swarm member (200; 1600; 1791; 1792; 1830) in the area (502).
2. The method (300) according to claim 1, wherein the step of determining (304) the deflection field (512) includes: Selecting one or more radio units (504, 506) from the plurality of radio units (504, 506); and rerouting the swarm members (200; 1600; 1791; 1792; 1830) around an obstacle (508) in the area (502) and / or into a deflection area (508) by implementing at least one safety buoy on the selected one or more radio units (504, 506), wherein the at least one safety buoy defines or acts as a source of the deflection field (512).
3. The method (300) according to claim 1 or 2, wherein, the deflection is locally induced only in a deflection area (508) within the area (502), and / or the deflection field (512) is transmitted (306) only by a predetermined subset (506) of the radio units (504, 506) around the deflection area (508), and / or wherein the vector field map (510) is transmitted independently of the deflection area (508) and / or is transmitted throughout the area (502) and / or is transmitted by a base station (504) covering the area (502).
4. The method (300) according to any one of claims 1 to 3, wherein, at least one of the steps of determining (302) the vector field map (510) and determining (304) the deflection field (512) is based on or includes: Execute reinforcement learning RL to optimize the deflection of swarm members (200; 1600; 1791; 1792; 1830), where the RL outputs an optimization strategy used in determining (302) the vector field map (510) and / or determining (304) the deflection field (512). Optionally, wherein the step of executing RL includes training the weights of a neural network, the neural network being embodied in the strategy used in determining (302) the vector field map (510) and / or determining (304) the deflection field (512), the neural network being configured to sense and interpret the environment (712) of the region (502), and training the weights by positive rewards for the desired results of navigation according to the vector field map (510) and / or deflection according to the deflection field (512) and / or negative rewards for the undesired results of navigation according to the vector field map (510) and / or deflection according to the deflection field (512).
5. The method (300) according to claim 4, wherein at least one of the plurality of swarm members (200; 1600; 1791; 1792; 1830) includes a sensor for continuously capturing sensor data, and the method (300) further comprises or initiates: Receiving data based on the sensor data from the swarm member (200; 1600; 1791; 1792; 1830). Wherein, The received data is fed back to the RL for optionally optimizing the deflection of the swarm member (200; 1600; 1791; 1792; 1830) when the swarm member (200; 1600; 1791; 1792; 1830) moves.
6. The method (300) according to claim 4 or 5, Wherein, Each trajectory in the region (502) is associated with a long-term reward, the long-term reward indicating the negative cost incurred by the swarm member (200; 1600; 1791; 1792; 1830) reaching the destination (850) in the region (502), and wherein the RL optimizes the strategy for determining (304) the deflection field (512) by modifying the velocity vector of the swarm member (200; 1600; 1791; 1792; 1830) to maximize the long-term reward.
7. The method (300) according to any one of claims 4 to 6, wherein the RL is executed in the environment (712, 740) of the region (502), the environment (712, 740) including at least one production unit operating in at least one of: a simulator, real hardware including the swarm of machines moving in the region (502), hardware in a loop including at least components of the swarm members, and a digital twin of the region (502) and the swarm members (200; 1600; 1791; 1792; 1830).
8. The method (300) according to any one of claims 1 to 7, Wherein, The vector field map (510) and / or the deflection field (512) further includes a destination (850) and / or at least one waypoint. The destination (850) is an attractor of the velocity vectors that exert an attraction force on the group members (200; 1600; 1791; 1792; 1830). The path points are associated with a deflection area in which a shift or turn is applied to all group members (200; 1600; 1791; 1792; 1830) within a respective one of the deflection areas in the same direction.
9. The method (300) according to any one of claims 1 to 8, wherein, the step (302) of determining the vector field map (510) includes updating the vector field map (510), and the step of transmitting (306) includes transmitting the updated vector field map (510) or transmitting the difference between the updated vector field map (510) and the previously transmitted vector field map. Optionally, wherein the difference is encoded using a motion vector field based on video coding.
10. The method (300) according to any one of claims 1 to 9, wherein in order to deflect the group members (200; 1600; 1791; 1792; 1830), the deflection field (512) is encoded to have at least one of the following: a shift in the position of each group member (200; 1600; 1791; 1792; 1830), optionally, wherein the direction of the shift is parallel throughout the deflection area (508); a change in the velocity of each group member (200; 1600; 1791; 1792; 1830), optionally, wherein the direction of the change is parallel throughout the deflection area (508); the center (508a) of the deflection area (508), optionally, the center (508a) of the obstacle (508); forces parallel throughout the deflection area (508); a repulsive force associated with the deflection area, optionally, a radial force centered on the obstacle (508) at the center (508a); and an attractive force associated with the path point, optionally, a radial force centered on the path point.
11. The method (300) according to any one of claims 1 to 10, wherein, the radio units (504, 506) include at least one or more of the following: radio points; radio strips; radio units dedicated to controlling the swarm of machines; radio units dedicated to locally transmitting the deflection field and / or acting as safety buoys; at least one or each of the group members (200; 1600; 1791; 1792; 1830); base stations of a radio access network RAN providing radio access to the swarm of machines; and radio units deployed within another RAN.
12. The method (300) according to any one of claims 1 to 11, wherein, the transmitting step (306) uses at least one of the following: a multimedia broadcast and multicast service MBMS channel; point-to-point transmission; ultra-reliable low-latency communication URLLC according to the fifth-generation 5G mobile communication; massive machine type communication mMTC according to 5G mobile communication; a non-cellular radio access technology, optionally a Wi-Fi unit. Optical radio access technology, optionally a Li-Fi unit; Unicast transmission; Multicast transmission; and Broadcast transmission.
13. The method (300) according to any one of claims 1 to 12, wherein, at least one radio unit (504, 506) of the plurality of radio units (504, 506) performs unicast transmission to send (306) the vector field map (510) and / or the deflection field (512) to different group members (200; 1600; 1791; 1792; 1830) using time interleaving or time division multiplexing.
14. The method (300) according to any one of claims 1 to 13, wherein, the determined (304) deflection field (512) indicates a uniform velocity vector or a uniform force vector for each or one of the deflection regions (508) within the region (502) for the deflection of the group members (200; 1600; 1791; 1792; 1830) relative to the vector field map (510), optionally, wherein the velocity vector or the force vector applied by the group members (200; 1600; 1791; 1792; 1830) for controlling the movement of at least one of the group members (200; 1600; 1791; 1792; 1830) within the region (502) also depends on the signal strength of the transmitted deflection field (512).
15. A method (400) for controlling group members, the group members (200; 1600; 1791; 1792; 1830) including at least one actuator configured to change the movement state of the group members (200; 1600; 1791; 1792; 1830) that are part of a swarm of machines moving within a region (502), the method (400) including or initiating: receiving (402) a vector field map (510) that includes velocity vectors indicating velocities and directions for navigating the group members (200; 1600; 1791; 1792; 1830) through the region (502); receiving (404) a deflection field (512) that indicates deflections for deflecting the group members (200; 1600; 1791; 1792; 1830) relative to the vector field map (510); determining (406) the position of the group members (200; 1600; 1791; 1792; 1830) within the region (502); determining (408) a change in the movement state based on the received vector field map (510) and the received deflection field (512) for the determined position; and controlling (410) the at least one actuator to achieve the changed movement state.
16. The method (400) according to claim 15, wherein, the step of determining (408) a change in the movement state includes at least one of the following: combining (407) the deflection field (512) and the vector field map (510); and Calculate a rotation vector from the gradient of the combined deflection field (512) and vector field map (510), where the rotation vector transforms the current movement state into a changed movement state.
17. The method (400) according to any one of claims 15 or 16, wherein, the received (404) deflection field (512) indicates a uniform velocity vector or a uniform force vector of a deflection region (508) within the region (502) for the deflection of the group members (200; 1600; 1791; 1792; 1830) relative to the vector field map (510), optionally, wherein the step (408) of determining a change in the movement state for a determined position in the deflection region includes: scaling the received uniform velocity vector or uniform force vector according to the signal strength of the deflection field (512) received (404) at the group member (200; 1600; 1791; 1792; 1830).
18. The method (400) according to any one of claims 15 to 17, further comprising any feature or step as claimed in any one of claims 1 to 14, or any corresponding feature or step thereto.
19. A computer program product comprising program code portions for performing the steps according to any one of claims 1 to 14 or 15 to 18 when the computer program product is executed on one or more computing devices (1504; 1604), the program code portions optionally stored on a computer-readable recording medium (1506; 1606).
20. A swarm control entity (100; 1500; 1712; 1820) for controlling a swarm of robotic devices in a region (502), the region (502) including a plurality of radio units (504, 506) for providing radio access to the swarm of robotic devices, the swarm of robotic devices including a plurality of group members (200; 1600; 1791; 1792; 1830), the swarm control entity (100; 1500; 1712; 1820) including a memory (1506) operable to store instructions and a processing circuit (1504) operable to execute the instructions such that the swarm control entity (100; 1500; 1712; 1820) is operable to: determine a vector field map (510) that includes velocity vectors indicating velocities and directions for navigating group members (200; 1600; 1791; 1792; 1830) through the region (502); determine a deflection field (512) that indicates a deflection for deflecting group members (200; 1600; 1791; 1792; 1830) relative to the vector field map (510); and The vector field map (510) and the deflection field (512) are sent by the radio unit (504, 506) to at least one group member (200; 1600; 1791; 1792; 1830) to control the movement of the at least one group member (200; 1600; 1791; 1792; 1830) in the area (502).
21. The radio device (100; 1500; 1712; 1820) according to claim 20 is further operable to perform the steps of any one of claims 2 to 14.
22. A group control entity (100; 1500; 1712; 1820) for controlling a swarm of robots in an area (502), the area (502) including a plurality of radio units (504, 506) for providing radio access to the swarm of robots, the swarm of robots including a plurality of group members (200; 1600; 1791; 1792; 1830), the group control entity (100; 1500; 1712; 1820) comprises: a vector field map determination module (102) configured to determine a vector field map (510), the vector field map (510) including velocity vectors indicating the velocity and direction for navigating group members (200; 1600; 1791; 1792; 1830) through the area (502); a deflection field determination module (104) configured to determine a deflection field (512), the deflection field (512) indicating a deflection for deflecting group members (200; 1600; 1791; 1792; 1830) relative to the vector field map (510); and a transmission module (106) configured to send the vector field map (510) and the deflection field (512) to at least one group member (200; 1600; 1791; 1792; 1830) by the radio unit (504, 506) to control the movement of the at least one group member (200; 1600; 1791; 1792; 1830) in the area (502).
23. The radio device (100; 1500; 1712; 1820) according to claim 22 is further configured to perform the steps of any one of claims 2 to 14.
24. A group member (200; 1600; 1791; 1792; 1830) includes at least one actuator configured to change the movement state of the group member (200; 1600; 1791; 1792; 1830) which is part of a swarm of robots moving in an area (502), the group member (200; 1600; 1791; 1792; 1830) includes a memory (1506) operable to store instructions and a processing circuit (1504) operable to execute the instructions such that the group member (200; 1600; 1791; 1792; 1830) is operable to: Receive a vector field map (510), the vector field map (510) including velocity vectors indicating the velocity and direction for navigating a group member (200; 1600; 1791; 1792; 1830) through the region (502); Receive a deflection field (512), the deflection field (512) indicating a deflection for deflecting a group member (200; 1600; 1791; 1792; 1830) relative to the vector field map (510); Determine the position of the group member (200; 1600; 1791; 1792; 1830) in the region (502); Based on the received vector field map (510) and the received deflection field (512), determine a change in the movement state for the determined position; And Control the at least one actuator to effect the changed movement state.
25. The group member (200; 1600; 1791; 1792; 1830) according to claim 24, further operable to perform the steps of any one of claims 16 to 18.
26. A group member (200; 1600; 1791; 1792; 1830) includes at least one actuator and is configured to change the movement state of the group member (200; 1600; 1791; 1792; 1830) which is part of a swarm of machines moving in a region (502), the group member (200; 1600; 1791; 1792; 1830) further Comprises: A vector field map receiving module (202) configured to receive a vector field map (510), the vector field map (510) including velocity vectors indicating the velocity and direction for navigating a group member (200; 1600; 1791; 1792; 1830) through a region (502); A deflection field receiving module (204) configured to receive a deflection field (512), the deflection field (512) indicating a deflection for deflecting a group member (200; 1600; 1791; 1792; 1830) relative to the vector field map (510); A position determining module (206) configured to determine the position of the group member (200; 1600; 1791; 1792; 1830) in the region (502); A movement state determining unit (208) configured to determine a change in the movement state based on the received vector field map (510) and the received deflection field (512) for the determined position; And An actuator control unit (210) configured to control the at least one actuator to effect the changed movement state.
27. The group member (200; 1600; 1791; 1792; 1830) according to claim 26, further configured to perform the steps of any one of claims 16 to 18.
28. A communication system (1700; 1800) including a host computer (1730; 1810), Comprises: A processing circuit (1818) configured to provide user data; And A communication interface (1816), configured to forward user data to a cellular or ad-hoc radio network (1710) for transmission to a user equipment UE (200; 1600; 1791; 1792; 1830), wherein the UE (200; 1600; 1791; 1792; 1830) includes a radio interface (1602; 1837) and a processing circuit (1604 1838), and the processing circuit (1604; 1838) of the UE (200; 1600; 1791; 1792; 1830) is configured to perform the steps according to any one of claims 15 to 18.
Citation Information
Patent Citations
Controlling unmanned aerial vehicles as a flock to synchronize flight in aerial displays
US9809306B2