METHOD FOR CONTROLLING AT LEAST ONE DRONE AND CONTROL THEREOF

DE602023011282T2Active Publication Date: 2026-01-28THALES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE602023011282
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-12-19
Filing Date
2023-12-18
Publication Date
2026-01-28
Estimated Expiration
2043-12-18

AI Technical Summary

Technical Problem

Coordinating multiple drones for efficient observation of large, predetermined geographical areas is complex and time-consuming, especially when areas to be observed change during a mission, and existing methods do not adequately account for drone capabilities and constraints.

Method used

A control method using an artificial neural network to process observation maps, drone positions, and convoy information, enabling optimized drone movement and image sensor positioning to efficiently cover all or selected areas, incorporating learning phases to adapt to changing conditions.

Benefits of technology

The method significantly reduces the time required to observe geographical areas by optimizing drone travel and image sensor positioning, ensuring comprehensive coverage with minimal redundancy.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a method for controlling at least one drone.

[0002] The invention also relates to a controller for at least one drone.

[0003] Drones, or UAVs (Unmanned Aerial Vehicles), include, for example, one or more image sensors to observe geographical areas. Specifically, the geographical areas to be observed are known in advance, or predetermined, and a drone moves between these areas to observe them successively.

[0004] Observing large, predetermined geographical areas often involves a very large number of drone shots. This is often time-consuming. By using several drones in parallel, the time required to observe each area within the predetermined zones is reduced.

[0005] However, coordinating multiple drones is often complex and difficult for an operator to implement.

[0006] US20060085106A1 describes an example of a method for observing an area using one or more drones, the method consisting of determining movement commands from an observation map defining areas to be monitored, a lighting map indicating their observation status, and information relating to the position and orientation of the drones.

[0007] Therefore, it is desirable to further reduce the travel time of the drone(s), while observing either all of the predetermined geographical areas, or at least a minimum number of these predetermined geographical areas.

[0008] One problem concerning the optimization of travel distances is the traveling salesman problem. This problem determines the minimum distance between a plurality of cities, while visiting each city only once. When several agents travel between the cities, the problem is called the Multiple Traveling Salesman Problem (MTSP).

[0009] However, such approaches do not take into account certain constraints and / or capabilities of drones observing predetermined geographical areas.

[0010] Also, in some cases, the areas to be observed are likely to change during a mission, thus modifying the problem to be solved.

[0011] One aim of the present invention is therefore to obtain a control method exhibiting optimized control of at least one drone to observe predetermined geographical areas.

[0012] To this end, the invention relates to a control method according to claim 1

[0013] Indeed, the control process allows for highly efficient observation of the areas to be monitored, as it takes into account several maps as well as the position of the drone and / or the image sensor. Thus, the control process enables observation of the areas to be monitored in a very short timeframe.

[0014] According to other advantageous aspects of the invention, the control method comprises one or more features according to claims 2 to 9, taken individually or in all technically possible combinations.

[0015] The invention also relates to a controller for at least one drone according to claim 10.

[0016] These features and advantages of the invention will become apparent upon reading the following description, given solely by way of non-limiting example and with reference to the accompanying drawings, in which: [ Fig 1 ] there figure 1 is a schematic view of a plurality of drones and a geographical area with a convoy of vehicles; [ Fig 2 ] there figure 2 is a schematic view of a controller configured to control one or more of the drones of the figure 1 ; Fig 3 ] there figure 3 is a schematic view of an example of an observation map comprising several areas to be observed by drones of the figure 1 ; Fig 4 ] there figure 4 is a schematic view of a plurality of maps including another example of the observation map of the figure 3 , a lighting map, a convoy map and a drone information map; Fig 5 ] there figure 5 is a schematic view of a neural network of the controller of the figure 2 , And [ Fig 6 ] there figure 6 is a flowchart of a control process for controlling one or more of the drones of the figure 1 .

[0017] With reference to the figure 1 , a set 10 includes several drones 12, also called UAVs (from the English "Unmanned Aerial Vehicle" for "unmanned aircraft"), and a convoy of vehicles 14 in a geographical region 16, also called region 16. Alternatively, the set 10 includes a single drone 12.

[0018] Each drone 12 includes, for example, at least one propulsion device 18 configured to move the drone 12, such as at least one rotor equipped with a motor, a controller 20 configured to control the drone 12 and at least one image sensor 22 configured to observe an observation pattern 24 in the geographical region 16.

[0019] Each drone 12 also includes, for example, a localization device, for example of the GNSS type (Global Navigation Satellite System), configured to determine a position of the drone 12. In one example, the localization device is integrated into the controller 20. In another example, the localization device is on board the drone 12, and separate from the controller 20.

[0020] As an example, each drone 12 also includes at least one communication device 26 configured to communicate with the other drones 12 and / or with a control station, not shown. The communication device 26 is, for example, configured to implement communication according to one or more different protocols, such as µTMA, 5G, and Wi-Fi.

[0021] Each drone 12 includes, for example, an obstacle detection and avoidance system configured to detect and avoid detected obstacles, an anti-collision system configured to avoid collisions with other drones 12 or other aircraft, and / or an altitude control system relative to the ground, configured to also enable terrain-following flight of the drone 12. The system or each of the obstacle detection and avoidance system, the anti-collision system and / or the control system is, for example, integrated into the controller 20 or, alternatively, forms an on-board system separate from the controller 20.

[0022] Each drone 12 includes, for example, a housing 28, specifically configured to allow the fixing or attachment of the propulsion device 18, the image sensor 22 and the communication device 26, and to receive the controller 20 in an internal space.

[0023] When the set 10 comprises several drones 12, each drone 12 is preferably at least partially identical to the others. In particular, each drone 12 comprises the same or similar components. Alternatively, the drones 12 comprise different components.

[0024] In the example of the figure 1 Each controller 20 is mounted on its respective drone 12. Alternatively, not shown, the controller(s) 20 are mounted on the remote control station of the drone 12. In this case, the control station, including the controller(s) 20, is, for example, located on the ground. In another example, or as a complement, the control station is mounted on a land vehicle, a ship, or an aircraft.

[0025] With reference to the figure 2 , the controller 20 includes for example a receiving module 30 configured to receive several cards, drone information and / or convoy information, a processing module 32 configured to determine at least one command for the drone 12 and a control module 36 configured to control the drone 12.

[0026] The controller 20 includes, for example, in addition, not shown, interface devices configured to exchange data, signals and / or commands with other elements of the drone 12, for example with the propulsion device 18, with the image sensor 22 and with the communication device 26.

[0027] The processing module 32 of the controller 20 includes at least one artificial neural network 34 configured to receive as input state variables comprising at least some of the maps, drone information, and optionally convoy information. The artificial neural network 34 is further configured to provide as output the command to move the drone 12 and / or to position the image sensor 22. The artificial neural network 34 is referred to hereafter as the neural network.

[0028] The receiving module 30, the processing module 32 and the control module 36 are each, for example, integrated into at least one computer 38.

[0029] In this case, each of the modules among the receiving module 30, the processing module 32 and the control module 36 is at least partially in the form of software executable by a processor 40 and stored in a memory 42 of the computer 38.

[0030] Alternatively or in addition, each of the modules among the receiving module 30, the processing module 32 and the control module 36 is integrated, at least partially, into a physical device, such as for example a programmable logic circuit, such as an FPGA (from the English "Field Programmable Gate Array"), or in the form of a dedicated integrated circuit, such as an ASIC (from the English "Application Specific Integrated Circuit").

[0031] The image sensor 22 is preferably capable of modifying the observation pattern 24 by rotating the image sensor 22 relative to a housing 28 of the drone 12. In particular, as illustrated for example on the figure 1 The image sensor 22 is equipped with a pivot motor 44 configured to change the orientation of the image sensor 22 relative to the housing 28. The image sensor 22 is thus particularly capable of taking images of different areas from the same position of the drone 12, and in particular capable of changing the observation pattern 24, by pivoting the image sensor 22 relative to the housing 28.

[0032] As an alternative or supplement not shown, the image sensor 22 is equipped with a digital control capable of modifying the observation pattern 24 of the image sensor 22, in particular independently of an orientation of the image sensor 22 itself.

[0033] According to one example, the image sensor 22 is capable of modifying the observation pattern 24 by enlarging, also called zooming, a part of the region covered by the observation pattern 24 and / or by reducing, also called zooming out, the part of the region covered by the observation pattern 24.

[0034] The image sensor 22 includes, for example, a camera, in particular an optical camera, and / or an infrared sensor. As an optional addition or alternative, the image sensor 22 includes a radar or lidar. The image sensor 22 is, for example, equipped with an image processing device capable of classifying detected objects, for example, by processing using machine learning. For example, the processing device is configured to detect vehicles or the presence of people in each image taken by the image sensor 22. In another example, such a processing device is integrated into the controller 20. In yet another example, the image processing device is carried on board the drone 12 and separate from the controller 20, or it is carried on board the control station (not shown), located remotely from the drone 12.

[0035] With reference to the figure 1 The convoy of vehicles 14 includes, for example, one or at least two vehicles, such as land vehicles. Alternatively, or in addition, the convoy of vehicles 14 includes at least one vehicle as well as pedestrians. In particular, the drone(s) 12 are configured to monitor the geographical area 16 in which the convoy of vehicles 14 is likely to operate.

[0036] A 100% control method for the drone(s) 12 will now be described, with reference to the figure 6 This process is implemented by the drone(s) 12, and in particular by the controller(s) 20.

[0037] The control process 100 includes a learning phase 102 and an operating phase 104.

[0038] The exploitation phase 104 includes, for example, a step of obtaining 110 an observation card 200, a first reception step 112, a second reception step 114, a third reception step 116, a determination step 118 and an ordering step 120.

[0039] During acquisition step 110, controller 20 obtains observation card 200, an example of which is illustrated on the figure 3 .

[0040] Observation map 200 defines a portion or fraction of geographic areas within geographic region 16 as areas to be observed by at least one of the drones, also referred to as observation cells. Specifically, geographic region 16 comprises several geographic areas, and observation map 200 defines a selection of some of these areas as the observation zones.

[0041] For example, during retrieval step 110, controller 20 obtains observation map 200 based on a predetermined, unrepresented terrain map comprising the plurality of geographic areas of geographic region 16, and based on a predetermined list of elements to be observed.

[0042] The predetermined terrain map includes, in particular, the set of geographical areas, for example arranged according to a regular grid.

[0043] The predetermined terrain map is, for example, of type 2.5D. As an example, a predetermined 2.5D terrain map is a non-flat 2D (two-dimensional) map that follows terrain deformations due to ground elevation. Specifically, the predetermined terrain map defines, for each geographic area, a width and length with associated geographic positions, as well as a single elevation per geographic area. The elevation could be, for example, the ground elevation, the maximum height of a building in the geographic area, or the maximum height of a tree in the geographic area. In one example, each geographic area has a length and a width between 5m and 10m.

[0044] The predetermined list of elements to be observed includes, for example, several elements of interest for observation. For example, the predetermined list of elements includes, with reference to the figure 3 , the convoy of vehicles 14, a river 203, a forest 204, a road 206 with a bridge 208.

[0045] During the acquisition step 110, the controller 20 selects in particular certain geographical areas, and defines them as the areas to be observed 201. As an illustration, the controller 20 selects the areas on the edge of the forest 204, at least certain areas of the road 206, around the bridge 208 and around the convoy of vehicles 14 as the areas to be observed 201.

[0046] According to a variant of retrieval step 110, the observation card 200 is stored in memory. Specifically, the observation card 200 is predetermined, for example, by an operator, or by a device or computer not shown. In this case, retrieval step 110 of the observation card 200 includes transmitting the observation card 200 to the computer 38, specifically to the receiving module 30.

[0047] Another example of the observation map 200 is illustrated on the figure 4 For example, each area to be observed is assigned an observation importance value, for example between 0 and 1. In the representation of the figure 4 The observation area 201 is clearer if the value is higher. When the observation importance value is equal to 1, the observation area 201, designated as important area 201a, is very important, for example, for a mission. When the value is close to 0, for example, equal to 0.1, the observation area 201, designated as less important area 201b, is less important for the mission. For example, the observation importance value depends on the position of the vehicle convoy 14.

[0048] During the first reception stage 112, the controller 20 receives a lighting card 210. The lighting card 210 is for example stored in a part of the memory 42 of the computer 38, and transmitted to the receiving module 30 during the first reception stage 112.

[0049] The lighting map 210 includes, for each area to be observed 201, an observation status by the drone 12.

[0050] The observation state is, for example, a value, for each observation zone 201, bounded between 0 and 1. When an observation zone 201 has been completely observed by drone 12, the value is, for example, equal to 1. When an observation zone 201 has not been observed at all by drone 12, the value is equal to 0.

[0051] The observation state of an area to be observed 201 indicates in particular whether the observation pattern 24 of the image sensor 22 covers, or has covered in a previous iteration, the area to be observed 201 partially or completely.

[0052] The observation pattern 24 corresponds, for example, to an image taken by the image sensor 22 at a given instant, preferably comprising areas within the image that are associated with observation intensity values. For example, a central area of ​​the image, such as a circle of predetermined diameter, is associated with a high observation intensity value. For example, peripheral areas of the image are associated with observation intensity values ​​lower than the high observation intensity value. The peripheral areas form, for example, peripheral circles around the central area. The central area includes, for example, a focal point of the image sensor 22.

[0053] According to one example, the observation pattern 24 is limited by a maximum distance determined from the image sensor 22 or around the position of the drone 12.

[0054] In one example, controller 20 applies the observation state value to a time decay law. For instance, the time decay law reduces the value after a predetermined period, or sets the value to 0 after the predetermined period. The predetermined period corresponds, in particular, to the obsolescence of information or a situation in the observation zone 201. In another example, or as an optional complement, controller 20 receives a command to set the value of a specific observation zone 201 to 0. This allows, in particular, forcing the drone 12 to observe the observation zone 201 again.

[0055] During the second reception stage 114, the controller 20 receives the drone information. The drone information includes, for example, the position of the drone 12, the orientation of the image sensor 22, and / or the observation pattern 24. In one example, the drone information also includes the destination of the drone 12.

[0056] For example, with reference to the figure 4 , drone information is presented in the form of a drone information card 212. In this case, the drone information card 212 includes, for example, the position 214, the observation pattern 24 and the destination 216 of the drone 12. In an alternative not shown, drone information is presented in the form of scalar values.

[0057] During the third reception stage 116, the controller 20 receives convoy information including for example a position and / or destination of a convoy of vehicles 14.

[0058] For example, with reference to the figure 4 The convoy information is presented in the form of a convoy map 220, showing, for example, the same geographical region 16 as the predetermined terrain map. The convoy map 220 includes, in particular, the position 222 and the destination 224 of the convoy of vehicles 14. In an alternative not shown, the convoy information is presented in the form of scalar values.

[0059] During determination step 118, the controller 20 determines the command for the drone 12 by the neural network 34.

[0060] The neural network 34 receives state variables as input. The state variables include at least the observation map 200, the lighting map 210, and drone information.

[0061] The neural network 34 provides as output the command to move the drone 12 and / or to position the image sensor 22.

[0062] The drone 12 movement control includes, in particular, a translation and / or rotation control of the drone 12 within a geographical reference point.

[0063] The positioning control of the image sensor 22 includes, in particular, a control for rotating the image sensor 22 relative to the housing 28 of the drone 12. Alternatively or in addition, the positioning control includes a control for modifying the observation pattern by enlarging and / or reducing an area observed by the image sensor 22.

[0064] As an example, neural network 34 also receives as input a state variable containing convoy information. This allows, in particular, for increased observation around the convoy of vehicles 14.

[0065] In one example, the neural network 34 also receives as input a state variable comprising a map of points of interest. In another example, the map of points of interest includes drone information, specifically the values ​​for the drone's position 12, its destination 12, and the position of the observation pattern 24. Alternatively or in addition, the map of points of interest includes at least some elements from the predetermined list of elements to be observed, such as the river 203, the forest 204, the road 206, and the bridge 208.

[0066] According to one example, neural network 34 also receives as input a state variable comprising a lighting map per terrain type, obtained by multiplying lighting map 210 with a binary map representing a terrain type.

[0067] The binary map indicates the areas that need to be observed and those that do not need to be observed for a given terrain.

[0068] According to one example, the neural network 34 also receives as input a state variable comprising a map of remaining-to-observe including all the areas to be observed 201 whose state of observation corresponds to a state of partial observation by the drone 12 or a state of non-observation by the drone 12.

[0069] With reference to the figure 5 , the neural network 34 includes for example a perception part 300 and a decision part 302 following the perception part 300.

[0070] During the determination step 118, the neural network 34 receives at least some of the state variables described above. The state variables are received by the neural network 34, for example, as global maps 304 covering the entire geographical region 16, as partial maps 306 forming part of a global map 304, and / or as scalars 308.

[0071] The perception component 300 comprises several network blocks 310, each including, for example, a convolutional neural network (CNN), an activation node (e.g., a ReLU), and a pooling node. The perception component 300 also includes, for example, at least one multilayer perceptron (MLP) block 312.

[0072] Each network block 310 and the MLP neural network block 312 processes the input variables and transmits the processed data to the decision part 302.

[0073] The decision section 302 comprises an actor module 314 and a critical module 316 implementing reinforcement learning. The decision section 302 generates outputs 318, namely the command to move the drone 12 and / or to position the image sensor 22.

[0074] During command step 120, the controller 20 commands the movement of the drone 12 and / or the positioning of the image sensor 22 according to the command. Thus, the drone 12 moves specifically to the position of the command and / or the image sensor 22 rotates according to the command and / or changes the observation pattern 24.

[0075] With reference to the figure 6 The process also includes, for example, an update step 122.

[0076] Update step 122 is preferably implemented as a follow-up to the implementation of command step 120.

[0077] During the update step 122, the controller 20 updates at least the lighting map 210 and the drone information according to the movement of the drone 12 and / or the positioning of the image sensor 22 implemented during the command step 120. The controller 20 thus obtains an updated lighting map and updated drone information including in particular the current drone position 12, in particular after the movement of the drone 12 according to the command, and the current orientation of the image sensor 22 and / or the current observation pattern 24, in particular after the rotation of the image sensor 22 according to the command.

[0078] Depending on one variant or as an optional addition, the lighting map 210 is updated based on the presence of other drones. Such a presence is transmitted, for example, via inter-drone communication or relayed by the vehicle convoy 14. In this case, for example, the observation statuses are updated for the observation areas 201 that are covered, at least partially, by the observation pattern 24 of the other drone(s).

[0079] According to a variant or as an optional supplement, the observation state of certain areas to be observed 201 is set to zero, or put into an unlit state, when a revisit command is given to this area to be observed 201 by the operator, for example present on the convoy of vehicles 14 (via communication with the convoy).

[0080] The updated lighting map is identical to the lighting map 210 received during the first reception stage 112, except for the observation state(s) of each area to be observed 201 which is covered, during the command stage 120, at least partially by the observation pattern 24 of the drone 12.

[0081] For example, the controller 20 modifies the observation state of each area to be observed 201 according to the coverage of the observation pattern 24 of the area to be observed 201 during the command step 120. In particular, when the observation pattern 24 covers, during the command step 120, this area partially, the controller 20 increases the observation state by a first value, and when the observation pattern 24 covers this area entirely, the controller 20 increases the observation state by a second value greater than the first value.

[0082] When the process includes the update step 122, the process preferably further includes at least one repetition of the determination step 118, in which the neural network 34 receives as input at least the observation map 200, the updated lighting map and the updated drone information.

[0083] Preferably, the process includes at least N repetitions, N being greater than 2, preferably greater than 100.

[0084] In one example, each iteration includes the implementation, in this order, of the retrieval step 110, the first reception step 112, the second reception step 114, the third reception step 116, the determination step 118, and the ordering step 120, and possibly the update step 122, particularly before a subsequent iteration. In another example, only the first implementation includes the retrieval step 110, and in subsequent iterations, the observation map 200 remains the same.

[0085] The learning phase 102 is preferably implemented at least once before an implementation of the exploitation phase 104.

[0086] During learning phase 102, the neural network 34 is trained, notably by reinforcement learning.

[0087] In particular, a 400 software agent, illustrated especially on the figure 2 , trains the neural network 34 through successive iterations to obtain, after one or more iterations, the trained neural network 34 ready for controlling drone 12. The software agent 400 is illustrated on the figure 2 as part of the controller 20, it is preferably only present during the learning phase 102. In particular, it is only embedded in the controller 20 or in another computer separate from the controller 20 during the learning phase 102.

[0088] For example, the learning phase 102 is implemented at least partially, preferably entirely, by a computer separate from the controller 20. The software agent 400 is presented for example in the form of a software block, executable by the processor 40 and stored in the memory 42 of the computer 38 during the learning phase 102.

[0089] The 400 software agent does not, ideally, understand a predetermined model, but implements actions based on a strategy whereby the 400 software agent learns a function associating a current state with an action to be performed. Such an approach is also called "model-free" in the context of reinforcement learning.

[0090] By "current state", it is understood in this description in particular a current state, during a current iteration, of the state variables received as input to the neural network 34 during the exploitation phase 104. In particular, the current state includes at least the observation map 200, the lighting map 210, the drone information, and as an optional complement the convoy information.

[0091] The action implemented by the software agent 400 includes, for example, the command to move the drone 12 and / or the command to position the image sensor 22.

[0092] Preferably, the software agent 400 implements the learning of the neural network 34 by calculating and optimizing a convergent gain G, in particular maximizing the expected value of the gain G, to train the neural network 34. The gain G is notably determined according to the following equation: G t = ∑ k = 0 ∞ γ k r t + k + 1 Or : rt is the reward received at time t γ ∈ [0,1] is a reduction factor penalizing future rewards.

[0093] According to one example, software agent 400 implements a Markov decision process, also called MDP (from the English "Markov Decision Process"), or a partially observable Markov decision process, also called POMDP (from the English "Partially Observable Markov Decision Process").

[0094] According to one example, learning sentence 102 includes an algorithm from the class of algorithms of "Temporal Difference (TD) learning" (name literally meaning Temporal Difference Learning).

[0095] For example, the software agent 400 implements for learning one or more of the algorithms among the "DQN", or "Deep Q Network" algorithm and an algorithm according to an "actor-critical" approach including an actor modeling a political function and a so-called critical parameter estimating a relevance of modifications and indicating a direction of optimization of the actor.

[0096] When the 400 software agent implements an algorithm according to the so-called "actor-critic" approach, it implements, for example, the "PPO" algorithm (for "Proximal Policy Optimization"), the "DDPG" algorithm (for "Deep Deterministic Policy Gradient"), in particular "TD3" (for "Twin Delayed DDPG"), or the "SAC" algorithm (for "Soft Actor Critic"), the latter including in particular stochastic modeling, and including modeling of entropy in a calculated objective.

[0097] According to one example, the learning phase 102 further includes training with several software agents 400 in parallel to obtain parameters enabling one of the aforementioned algorithms to achieve convergence.

[0098] As an example, learning phase 102 includes the implementation of a PBT (Population Based Training) approach to obtain hyperparameters.

[0099] For example, and with reference to the figure 6 , the learning phase 102 includes a determination step 402 of an action by the software agent 400, an application step 404 of the action, an elaboration step 406, and a transmission step 408.

[0100] During determination step 402, software agent 400 determines the action. For example, software agent 400 determines commands or instructions. As an example, determination step 402 includes the calculation of auxiliary indicators and / or the normalization of variables for application to the neural network 34. For example, the action includes the command to move drone 12, the command to position image sensor 22, and / or the movement of the convoy.

[0101] During application step 404, the software agent 400 applies the action, specifically the command. For example, the drone information is modified as a result of the drone 12's movement. For instance, the lighting map 210 is modified, specifically the observation state of each area to be observed 201 that is at least partially covered by the observation pattern 24 of the image sensor 22 when the command is applied.

[0102] During the elaboration step 406, for example an interpreter, not shown, determines a reward for the software agent 400 and a new state including the state variables modified during the application step 404. The interpreter is presented for example in the form of a software block, executable by the processor 40 and stored in the memory 42 of the computer 38 during the learning phase 102.

[0103] For example, the interpreter implements, during the elaboration step 406, a first sub-step 410, a second sub-step 412 and a third sub-step 414 to obtain the reward.

[0104] During the first substep 410, the interpreter determines a lighting progress map corresponding to the difference between the lighting map 210 of the current iteration and the lighting map 210 of a previous iteration, that is, specifically, an iteration directly preceding the current iteration. The lighting progress map includes, in particular, the differences in the observation states of the areas to be observed 201.

[0105] During the second substep 412, the interpreter multiplies the observation card 200 with the lighting progress card to obtain a reward card. Specifically, the interpreter multiplies, zone by zone, the status of each zone on the lighting progress card with the observation card 200.

[0106] During the third substep 414, the interpreter obtains the reward, for example a scalar value, by applying a predetermined function to the reward card.

[0107] The interpreter applies the predetermined function to all areas of the reward map, specifically to an observation state value associated with each area. The predetermined function is, for example, a sum or average of the values ​​in the reward map areas.

[0108] According to one example, the interpreter obtains the reward further depending on the state variables received by the neural network 34 as input during the exploitation phase 104, including for example convoy information and / or the map of points of interest.

[0109] During transmission step 408, the interpreter transmits the reward and the new state to the software agent 400.

[0110] The determination step 402, the application step 404, the development step 406, and the transmission step 408 are preferably implemented in that order.

[0111] The determination step 402, the application step 404, the elaboration step 406, and the transmission step 408 are preferably implemented several times, in order to obtain the trained neural network 34.

[0112] In particular, with reference to the figure 6 , the learning phase 102 further includes a modification step 416, implemented for example as a consequence of an implementation of the transmission step 408 and before an implementation of the determination step 402 of a subsequent iteration.

[0113] The modification step 416 modifies the neural network 34 based on collected interaction data, such as state variables, the action(s), and / or the reward. Specifically, the reward is optimized using a loss or cost function from a reinforcement learning algorithm, such as the DQN, PPO, DDPG, and / or SAC algorithms. This allows for the iterative adoption of better actions.

[0114] The control process 100 is for example implemented for a single drone 12 and / or by a single controller 20.

[0115] Alternatively, the control method 100 includes the control of several drones 12, each comprising at least one image sensor 22 configured to observe a respective observation pattern 24.

[0116] In one example, controller 20 is a centralized controller controlling several drones 12. Alternatively, each drone 12 includes its own controller 20.

[0117] Preferably, the control method 100 includes the control of several drones 12, each comprising at least one image sensor 22 configured to observe a respective observation pattern 24.

[0118] As an example, controller 20 is trained using the CTDE principle (Centralized Training for Decentralized Execution). For instance, controller 20 is trained by modeling several drones 12, but during the operational phase 104, each drone 20 is controlled independently, while also receiving states from the other drones 12.

Claims

1. A method for controlling (100) at least one drone (12) comprising at least one image sensor (22) configured to observe an observation pattern (24), the control method (100) comprising an operating phase (104), comprising the steps of: - acquiring (110) an observation map (200) for a region (16) comprising a plurality of geographical zones, the observation map (200) defining part of the geographical zones of the region (16) as zones to observe (201); - reception (112) of a lighting map (210), comprising, for each zone to observe (201), a state of observation by the drone (12); - reception (114) of drone information relating to a position of the drone (12) and / or an orientation of the image sensor (22); - determination (118) of at least one control for the drone (12) by at least one neural network (34), the neural network (34) receiving, as input, state variables comprising at least the observation map (200), the lighting map (210) and said drone information, the neural network (34) providing at the output at least one command for the movement of the drone (12) and / or of the positioning of the image sensor (22), the control method further comprising a learning phase (102) of the neural network (34), during which the neural network (34) is trained by reinforcement learning, wherein the learning phase (102) comprises a step of developing (406) a reward for a software agent (400) based on a reward map, the development step (406) comprising: - a first sub-step (410) of determination of a lighting progress map corresponding to a difference between the lighting map (210) of a current iteration and the lighting map (210) of a previous iteration; - a second sub-step (412) of determination of the reward map, by multiplying the observation map (200) with the lighting progress map; - a third sub-step (414) of acquiring the reward by application of a predetermined function to the reward map.

2. The control method (100) according to claim 1, wherein the image sensor (22) is apt to modify the observation pattern (24) by rotating the image sensor (22) with respect to a housing (28) of the drone (12), by enlarging a portion of the region (16) covered by the observation pattern (24) and / or reducing the portion of the region (16) covered by the observation pattern (24).

3. The control method (100) according to claim 1 or claim 2, further comprising a command step (120) comprising a movement of the drone (12) and / or a positioning of the image sensor (22) according to the command.

4. The control method (100) according to claim 3, further comprising a step of updating (122) the lighting map (210) and the drone information according to the movement of the drone (12) and / or the positioning of the image sensor (22), in order to obtain an updated lighting map and an updated drone information, the control method (100) further comprising a repetition of the determination step (118), during which the neural network (34) receives, as input, at least the observation map (200), the updated lighting map and said updated drone information.

5. The control method (100) according to claim 4, wherein during the update step (122), the state of observation of each observation zone (201) is modified when the observation pattern (24) covers the zone to observe (201) at least partially, preferentially entirely, during the command step (120).

6. The control method (100) according to any of the preceding claims, wherein the acquiring step (110) comprises acquiring the observation map (200) based on a predetermined terrain map comprising the plurality of geographic zones of the region (16) and based on a predetermined list of elements to observe.

7. The control method (100) according to any of the preceding claims, wherein the neural network (34) receives as input, the state variables further comprising at least a map among: - a map of points of interest comprising values of a position of the drone (12), of a destination of the drone (12), and a position of the observation pattern (24); - one lighting map per type of terrain, acquired by multiplying the lighting map (210) and a binary map representing a type of terrain; - a remains-to-observe map comprising all the zones to observe (201), the state of observation of which corresponds to a partial state of observation by the or each drone (12) or a non-state of observation by the or each drone (12).

8. The control method (100) according to any of the preceding claims, wherein the neural network (34) receives, as input, the state variables further comprising convoy information comprising a position and / or a destination of a convoy of vehicles (14).

9. The control method (100) according to any of the preceding claims, wherein, during the development step (406), the reward is further developed according to at least a part of the state variables received by the neural network (34) at the input during the operating phase (104).

10. A controller (20) of at least one drone (12) comprising at least one image sensor (22) configured to observe an observation pattern (24), the controller (20) comprising a reception module (30) configured to receive an observation map (200) for a region (16) comprising a plurality of geographic zones, the observation map (200) defining a portion of the geographical zones of the region (16) as zones to observe (201), the reception module (30) being further configured to receive a lighting map (210), comprising, for each zone to observe (201), a state of observation by the drone (12), the reception module (30) being further configured to receive drone information relating to a position of the drone (12) and / or an orientation of the image sensor (22), the controller (20) further comprising a processing module (32) configured to determine at least one command for the drone (12), wherein the processing module (32) comprises at least one neural network (34) configured to receive, as input, state variables comprising at least the observation map (200), the lighting map (210) and said drone information, and configured to output at least one command for the movement of the drone (12) and / or of the positioning of the image sensor (22), the neural network having been trained during a learning phase, during which the neural network is trained by reinforcement learning, the learning phase (102) comprising a step of developing (406) a reward for a software agent (400) according to a reward map, the development step (406) comprising: - a first sub-step (410) of determination of a lighting progress map corresponding to a difference between the lighting map (210) of a current iteration and the lighting map (210) of a previous iteration; - a second sub-step (412) of determination of the reward map, by multiplying the observation map (200) with the lighting progress map; - a third sub-step (414) of acquiring the reward by application of a predetermined function to the reward map.