A cooperative control system for unmanned aerial vehicles (UAVs) under complex interference environments
An environmental situation map is generated by the data acquisition, fusion, and decision-making modules. The PER-D3QN is used to determine the control strategy. The collaborative control module enables the UAV to perform efficient and intelligent search in complex environments, solving the problem of the impact of dynamic changes on the search and achieving high-precision and high-efficiency target search.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINESE PEOPLES ARMED POLICE FORCE GUANGXI ZHUANG AUTONOMOUS REGION CORPS HOSPITAL
- Filing Date
- 2026-04-15
- Publication Date
- 2026-06-02
Smart Images

Figure CN122131804A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent control of unmanned aerial vehicles (UAVs), specifically to, but not limited to, a collaborative control system for UAVs in complex interference environments. Background Technology
[0002] When drones search for targets in complex and dynamic environments, factors such as adverse weather, complex terrain, and diverse obstacles can interfere with the search process. To reduce the interference of these factors and improve the accuracy of target searching, related technologies provide pre-set static maps of the search environment for the drone, enabling it to correct its search trajectory in real time based on the static map during the search process.
[0003] However, the above solutions cannot overcome the negative impact of the dynamic changes of the aforementioned factors on the drone search process, resulting in low efficiency of drones searching for target objects. Summary of the Invention
[0004] Based on the above technical problems, this application provides a cooperative control system for UAVs in complex interference environments, which can achieve high-precision and high-efficiency search for target objects in complex interference environments.
[0005] The technical solution provided in this application is as follows: This application provides a cooperative control system for unmanned aerial vehicles (UAVs) in complex interference environments, including: The acquisition module is used to acquire multimodal k-th source data; wherein, the k-th source data includes the k-th meteorological data, the k-th obstacle distribution state, and the k-th object state of the complex interference environment in which the UAV is located at the k-th time; k is an integer greater than or equal to 1; The fusion module is used to perform spatiotemporal fusion on the data in the k-th source data to obtain the k-th environmental situation map; wherein, the k-th environmental situation map includes the regional risk level of at least some areas in the complex interference environment; The decision module is used to determine the (k+1)th control strategy based on the kth device state of the UAV and the kth environmental situation map using PER-D3QN; wherein the (k+1)th control strategy includes the (k+1)th search path and the (k+1)th action set. The collaborative control module is used to control the UAV to search for target objects in the complex interference environment at time k+1, based on the (k+1)th search path and the (k+1)th action set.
[0006] The collaborative control system for UAVs in complex interference environments provided in this application embodiment has at least the following beneficial effects: In the collaborative control system for UAVs in complex interference environments provided in this application embodiment, the acquisition module is used to acquire multimodal k-th source data. The k-th source data includes the k-th meteorological data, the k-th obstacle distribution status, and the k-th object status of the complex interference environment in which the UAV is located at time k. k is an integer greater than or equal to 1. Thus, as the value of k changes continuously, it is possible to achieve comprehensive, real-time, and flexible tracking and acquisition of meteorological data, obstacle distribution data, and object status of the complex interference environment in which the UAV is located. Furthermore, the fusion module is used to perform spatiotemporal fusion of the data in the k-th source data to obtain the k-th environmental situation map of the complex interference environment. The k-th environmental situation map includes the regional risk level of at least some areas in the complex interference environment. Thus, the fusion module can intuitively fuse various multimodal, discrete, and isolated data in the k-th source data, so that the k-th environmental situation map can comprehensively and intuitively display the complex interference environment. The system determines the regional risk level of at least some areas. Based on this, the decision-making module uses PER-D3QN to determine the (k+1)th control strategy, including the (k+1)th search path and the (k+1)th action set, based on the (k)th device status of the UAV and the (k)th environmental situation map. This leverages PER-D3QN's ability to cope with complex and changing environments and its efficient and accurate adaptive predictive control strategy, improving the completeness and accuracy of the (k+1)th control strategy and increasing its determination efficiency. On the other hand, the collaborative control module, based on the (k+1)th search path and the (k+1)th action set, controls the UAV to search for target objects in a complex interference environment at time (k+1). This mitigates the negative impact of dynamic factors such as meteorological data on the target object search process in complex interference environments, enabling automated, intelligent, high-precision, and highly efficient flexible search of target objects. Attached Figure Description
[0007] Figure 1 A schematic diagram of the structure of a cooperative control system for a drone under complex interference environment provided in an embodiment of this application; Figure 2 Another structural schematic diagram of the collaborative control system provided in the embodiments of this application; Figure 3 This is another structural schematic diagram of the collaborative control system provided in the embodiments of this application. Detailed Implementation
[0008] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0009] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0010] In complex and disruptive environments such as emergency rescue, disaster sites, urban operations, and field exploration, using drones to search for targets has become a widely adopted technical solution. However, in real-world search environments, adverse weather conditions, including heavy rain, dense fog, and strong winds, can make it difficult for drones to perceive these environments. Factors such as the distribution of dynamic obstacles and interference from strong light and noise can also make it difficult for drones to accurately plan search paths.
[0011] To address the above technical issues, the following technical solutions are provided: pre-configure a static map of the search environment for the drone, enabling the drone to obtain the distribution of various obstacles in the search environment in advance. During the specific object search process, the drone's search trajectory can be controlled manually to improve the efficiency of drone search.
[0012] However, the above technical solutions cannot overcome the negative impact of dynamic changes in factors such as weather and dynamic obstacles in the search environment on the UAV search process.
[0013] Based on the above technical problems, this application provides a cooperative control system for unmanned aerial vehicles (UAVs) in complex interference environments. Figure 1 This is a schematic diagram of the structure of a cooperative control system for a drone under complex interference environments, provided as an embodiment of this application. Figure 1 As shown, the cooperative control system 100 may include: The acquisition module 101 is used to acquire the k-th source data of multimodal data; wherein, the k-th source data includes the k-th meteorological data, the k-th obstacle distribution state, and the k-th object state of the complex interference environment in which the UAV is located at the k-th time; k is an integer greater than or equal to 1; The fusion module 102 is used to perform spatiotemporal fusion on the data in the k-th source data to obtain the k-th environmental situation map; wherein the k-th environmental situation map includes the regional risk level of at least some areas in the complex interference environment; The decision module 103 is used to determine the (k+1)th control strategy based on the kth device state and the kth environmental situation map of the UAV through PER-D3QN; wherein the (k+1)th control strategy includes the (k+1)th search path and the (k+1)th action set. The collaborative control module 104 is used to control the UAV to search for target objects in a complex interference environment at time k+1, based on the k+1 search path and the k+1 action set.
[0014] In some embodiments, a complex interference environment may include an environment where the severity of weather conditions is greater than or equal to a first threshold, the complexity of terrain is greater than or equal to a second threshold, and the complexity of obstacle distribution is greater than or equal to a third threshold. For example, a complex interference environment may include an environment where natural disasters or man-made disasters occur. For instance, a complex interference environment may include an environment where fires or floods occur, a battlefield environment, or an environment for emergency medical treatment, emergency equipment maintenance, and material supply.
[0015] In some embodiments, the k-th meteorological data can describe the meteorological conditions within the data acquisition range of the UAV in a complex and interfering environment at time k. For example, the meteorological conditions may include wind speed, precipitation status, humidity, and visibility. For example, the precipitation status may include rain or snow, and may also include descriptive data on precipitation amounts, such as heavy rain or light snow. Accordingly, the k-th meteorological data can be acquired in real time by meteorological sensors associated with the UAV.
[0016] In some embodiments, the k-th obstacle distribution state may include the distribution state of obstacles within the data acquisition range of the UAV at time k; for example, the distribution state of obstacles may include at least one of the density state of obstacles, the type, volume and quantity of obstacles; correspondingly, the k-th obstacle distribution state can be obtained by analyzing data acquired by thermal imagers, infrared cameras or radar, for example, the k-th obstacle distribution state can be determined by analyzing point cloud data acquired by radar.
[0017] In some embodiments, the state of the k-th object may include the distribution state of target objects within the data acquisition range of the UAV at time k. Exemplarily, the target object may be a person or animal, and may also include equipment or devices; for example, the target object may include a wounded person or a vehicle. Exemplarily, the distribution state of the target objects may include at least one of the following: whether a target object exists within the data acquisition range, the number of target objects within the data acquisition range, the type of target object, the attitude of the target object, and its location. Accordingly, the state of the k-th object can be obtained by analyzing data acquired by a thermal imager, infrared camera, or radar, etc.
[0018] In some embodiments, when the target object is wearing a smart monitoring device, the state of the kth object may further include the target object's physiological indicators and rescue needs data; wherein, the physiological indicators data may include data such as the target object's heart rate, blood pressure and blood oxygen, while the rescue needs data may include data on the target object's needs for drinking water, food and other tools and equipment; for example, the smart monitoring device may include a smart wearable device.
[0019] Specifically, the distribution state of the kth obstacle and / or the state of the kth object can also be obtained in any of the following ways: PointNet processes the k-th original point cloud corresponding to the radar echo to obtain the k-th obstacle distribution state and / or the k-th object state. For example, the input layer of PointNet can receive the k-th original point cloud, and the shared multilayer perceptron (MLP) of PointNet can perform point-by-point feature transformation and dimensionality upscaling on the k-th original point cloud through multiple fully connected layers. The pooling layer of PointNet can perform pooling processing on the result of feature transformation and dimensionality upscaling to obtain the global feature vector of the k-th original point cloud. Then, the fully connected layer of PointNet concatenates the above global feature vector to obtain the k-th obstacle distribution state and / or the k-th object state.
[0020] ResNet receives the k-th image or k-th thermal infrared image acquired by the image acquisition device. It then performs downsampling, primary feature extraction, and deep semantic feature extraction on the k-th image or k-th thermal infrared image through the convolutional layers, pooling layers, and residual blocks of ResNet. Finally, it generates global image features through pooling layers and identifies the global image features to obtain the k-th obstacle distribution state and / or the k-th object state.
[0021] In some embodiments, the acquisition module may include a collection of sensor devices associated with the UAV; for example, the sensor devices may include image acquisition devices and radar, wherein the image acquisition devices may include infrared cameras, thermal imagers and RGB cameras, etc.
[0022] In some embodiments, after the UAV switches to object search mode, the acquisition module can continuously acquire source data at specified time intervals to obtain a source data sequence; wherein, the source data sequence includes the kth source data, and the object search mode can include a working mode in which the UAV searches for target objects.
[0023] In some embodiments, the k-th environmental situation map can represent and label the regional risk level of the area corresponding to the data acquisition range of the UAV in a gridded form; for example, the k-th environmental situation map can correspond to a target coordinate system, which may be different from the radar coordinate system corresponding to the radar and the camera coordinate system corresponding to the image acquisition device, and the target coordinate system can be pre-specified or adjusted; for example, the k-th environmental situation map can be obtained in the following ways: By processing the k-th source data, a k-th environmental map is obtained. Risk identification is performed on the k-th environmental map to obtain the regional risk level of each pixel region in the k-th environmental map, and the pixel regions in the k-th environmental map are labeled based on the regional risk level. Based on the device position of the UAV at time k, a high-precision digital map corresponding to the data collection area is obtained, and the k-th environmental map labeled with regional risk levels is spatially overlaid and fused with the high-precision digital map to obtain a k-th environmental situation map. For example, the regional risk level can include the degree of threat posed by the weather conditions and / or obstacles in the pixel region to the life safety of the target object. For example, the k-th environmental situation map can use different colors to intuitively present the regional risk level of different pixel regions.
[0024] For example, the k-th environment graph can be obtained in the following way: The k-th environment map is obtained by spatiotemporal fusion of data from the k-th source data through PI-Fusion. For example, the k-th environment map can be a two-dimensional image or a three-dimensional image. Specifically, the feature alignment module of PI-Fusion can spatially align the k-th meteorological data with the k-th obstacle distribution state and the k-th object state. The feature mapping module of PI-Fusion can perform feature mapping on the features corresponding to the k-th meteorological data, the k-th obstacle distribution state and the k-th object state respectively through a cross-attention mechanism or a gated fusion network to obtain the k-th mapped data. Then, the k-th mapped data is integrated and fused to obtain the k-th environment map.
[0025] Specifically, the feature mapping module can process the feature distance and spatial distance of the features corresponding to the k-th meteorological state, the k-th obstacle distribution state, and the k-th object state, respectively, to determine the comprehensive distance between the features corresponding to the above data. Then, based on the comprehensive distance, feature mapping is performed on the features corresponding to the k-th meteorological state, the k-th obstacle distribution state, and the k-th object state, respectively, to obtain the k-th mapping data. Then, a four-parameter planar model is used to fuse the features corresponding to the k-th meteorological state, the k-th obstacle distribution state, and the k-th object state into the target coordinate system. Specifically, it can be shown in equations (1) to (2): (1) (2) in, for and The characteristic distance between them for and Spatial distance between them For the overall distance, For the normalization parameter, ( ) represents the different features in the k-th meteorological state, the k-th obstacle distribution state, and the k-th object state. , ) is the coordinate system of the target and ( The corresponding feature points, and For translation parameters, For scale parameters, These are rotation parameters.
[0026] In some embodiments, the k-th environment graph can also be obtained in the following ways: A transformation matrix is established between the radar coordinate system corresponding to the point cloud data and the image coordinate system corresponding to the image data. The k-th meteorological data is then transformed using the transformation matrix to obtain the k-th transformed meteorological data. The k-th transformed meteorological data is then fused with the image data corresponding to the k-th obstacle distribution state and the k-th object state to obtain the k-th fused image. Feature recognition is then performed on the k-th fused image to obtain the k-th recognition result. Finally, the k-th recognition result is combined with the three-dimensional spatial range corresponding to the k-th meteorological data to obtain the k-th environmental map. The k-th recognition result may include the distribution range and location of obstacles and target objects within the data acquisition range at the k-th time.
[0027] The transformation of the k-th meteorological data using the transformation matrix can be calculated using equation (3): (3) in,( , , () represents point cloud data in the radar coordinate system. , () represents the pixels in the image data. This is the transformation matrix.
[0028] In some embodiments, the (k+1)th search path may include the flight trajectory of the UAV at time (k+1) when it searches for the target object.
[0029] In some embodiments, the (k+1)th action set may include at least one or more actions that the UAV needs to perform when flying along the (k+1)th search path in order to search for the target object at the (k+1)th time; for example, the (k+1)th action set may include at least one of the following actions: flying forward, flying backward, flying left, flying right, hovering, climbing, and descending.
[0030] In some embodiments, the state of the k-th device may include the load state, remaining energy, and fault state of the drone at the k-th time. For example, the load state may include the usage of the drone's data processing resources at the k-th time, and may also include whether the drone is in an overload state. For example, the remaining energy may include the remaining available power of the drone at the k-th time, and the fault state may include whether the drone is in a fault state at the k-th time, and the number and / or type of drone device faults.
[0031] In some embodiments, the (k+1)th control policy can be determined in the following way: By analyzing the historical search records of the UAV using PER-D3QN, the historical search range of the UAV is determined. Based on the degree of regional similarity between the historical search range and the k-th range represented by the k-th environmental situation map, the (k+1)-th search path is determined. Then, based on the regional risk level of the k-th environmental situation map, the spatial risk level of the current location of the UAV is determined. Finally, based on the spatial risk level and the status of the k-th device, the (k+1)-th action set is determined.
[0032] Specifically, if the regional similarity is less than or equal to the fourth threshold, it indicates that there is basically no overlap between the historical search range of the UAV and the k-th range. In this case, the trajectory formed by the location in the k-th environmental situation map where the regional risk level is less than or equal to the risk threshold can be determined as the k+1 search path. If the regional similarity is greater than the fifth threshold, it indicates that the historical search range of the UAV and the k-th range have a high overlap rate. In this case, the trajectory formed by the edge location in the k-th environmental situation map where the regional risk level is less than or equal to the risk area can be determined as the k+1 search path.
[0033] Specifically, based on the spatial risk level of the current location of the UAV, the initial set of actions when the UAV flies along the search path at time k+1 can be predicted. The set of actions in the initial set that have a probability of damage to the UAV body less than or equal to a probability threshold and can be supported or carried by the kth device state can be determined as the k+1 action set. For example, the action set that can be supported or carried by the kth device state may include at least one of the following: actions that can be supported by the remaining available power of the UAV, and actions that can be performed by the UAV when it is in a non-faulty state.
[0034] In some embodiments, there may be a one-to-one correspondence between the (k+1)th search path and the (k+1)th action set; for example, the (k+1)th search path may include at least two paths, and each path in the (k+1)th search path may correspond one-to-one with an action in the (k+1)th action set.
[0035] In some embodiments, the collaborative control module can control the drone to search for target objects in the following ways: At time k+1, the drone is controlled to execute the action in the action set k+1, and continues to fly according to the flight trajectory corresponding to the search path k+1, so as to continuously search for the target object.
[0036] It should be noted that while the collaborative control module controls the UAV to search for the target object, the acquisition module can collect the (k+1)th source data and obtain the (k+1)th environmental situation map through the method provided in the aforementioned embodiment. Then, the (k+2)th control strategy is determined, and the UAV is controlled to continue searching for the target object based on the (k+2)th control strategy. The above operations are executed recursively in this way to achieve autonomous collaborative control of the UAV's search process for the target object in a complex interference environment.
[0037] It should be noted that when there are multiple target objects, the urgency level can be determined based on the physiological indicator data and / or rescue demand data uploaded by the intelligent monitoring devices of the target objects. The drone can then be controlled to search for the target object with the highest urgency level, and the k+1 control strategy for that target object can be determined first.
[0038] As can be seen from the above, in the collaborative control system for UAVs in complex interference environments provided in this application embodiment, the acquisition module is used to acquire multimodal k-th source data. The k-th source data includes the k-th meteorological data, the k-th obstacle distribution status, and the k-th object status of the complex interference environment where the UAV is located at time k. k is an integer greater than or equal to 1. Thus, as the value of k changes continuously, it is possible to achieve comprehensive, real-time, and flexible tracking and acquisition of meteorological data, obstacle distribution data, and object status of the complex interference environment where the UAV is located. Furthermore, the fusion module is used to perform spatiotemporal fusion of the data in the k-th source data to obtain the k-th environmental situation map of the complex interference environment. The k-th environmental situation map includes the regional risk level of at least some areas in the complex interference environment. Thus, the fusion module can intuitively fuse various multimodal, discrete, and isolated data in the k-th source data, so that the k-th environmental situation map can comprehensively and intuitively display the complex interference environment. The regional risk level of at least some areas in the interference environment is determined. Based on this, the decision module uses PER-D3QN to determine the (k+1)th control strategy, including the (k+1)th search path and the (k+1)th action set, based on the (k)th device state of the UAV and the (k)th environmental situation map. In this way, by leveraging the advantages of PER-D3QN in coping with complex and changing environments and possessing efficient and accurate adaptive predictive search strategies, the integrity and accuracy of the (k+1)th control strategy can be improved, as well as the determination efficiency of the (k+1)th control strategy. On the other hand, the cooperative control module is used to control the UAV to search for target objects in the complex interference environment at time (k+1) based on the (k+1)th search path and the (k+1)th action set. In this way, the negative impact of dynamic changing factors such as meteorological data on the target object search process can be reduced in the complex interference environment, thereby achieving automated, intelligent, high-precision, and highly efficient flexible search of target objects.
[0039] Based on the foregoing embodiments, in the collaborative control system for UAVs under complex interference environments provided in this application, the decision module is used to determine path constraints and, through Model Predictive Control (MPC) associated with the UAV, determine the (k+1)th candidate strategy based on the path constraints and the state of the kth device; wherein, the (k+1)th candidate strategy includes a set of (k+1)th candidate paths and a set of (k+1)th candidate actions corresponding to the set of (k+1)th candidate paths. The decision module is also used to evaluate the (k+1)th candidate action set using the PER-D3QN value function to obtain the (k+1)th value set, filter the (k+1)th candidate action set based on the (k+1)th value set to obtain the (k+1)th action set, and determine the path in the (k+1)th candidate strategy that corresponds to the (k+1)th action set as the (k+1)th search path.
[0040] In some embodiments, path constraints may include pre-determined conditions, or they may vary depending on the object type of the target object, the priority of the search for the target object, the environmental type of the complex interference environment, and the device status of the UAV. For example, the object type may include biological objects and device objects, the environmental type may include natural and social environments, and the device status may include fault status and non-fault status, etc.
[0041] In some embodiments, path constraints may include at least one condition for balancing the device safety of the UAV with the search efficiency for the target object.
[0042] In some embodiments, the (k+1)th candidate strategy may include multiple strategies, and each strategy in the (k+1)th candidate strategy may follow path constraints.
[0043] In some embodiments, the k-th device state may include the k-th spatial position and k-th spatial action of the UAV at time k; for example, the k-th spatial position may include the coordinate position of the UAV in the target coordinate system at time k, and the k-th spatial action may include the components of the UAV's flight speed in various directions of the target coordinate system, such as the k-th device state. It can be shown in equation (4): (4) in, It can be the k-th spatial location. It can be an action in the k-th space.
[0044] In some embodiments, the (k+1)th candidate strategy can be determined in the following way: Taking the kth spatial position in the kth device state as the initial state, the control MPC predicts the kth initial state based on the path constraints through the motion state equation of the UAV, obtains the (k+1)th candidate path set, determines the motion trend between the kth spatial action and the trajectory points represented by the paths in the (k+1)th candidate path set, and decomposes the above motion trend to obtain the (k+1)th candidate action set.
[0045] In some embodiments, the value in the (k+1)th value set can characterize the action advantage of a candidate action in the (k+1)th candidate action set; for example, the action advantage can characterize the difference between the search reward corresponding to a certain action in the (k+1)th candidate action set performed by the UAV and the search reward corresponding to all other actions in the (k+1)th candidate action set.
[0046] Accordingly, the value in the (k+1)th value set can be calculated in the following way: The state gain of the UAV's current position is determined by PER-D3QN. Given the UAV's current position, the action advantage of performing any action from the (k+1)th candidate action set is determined. The state gain and action advantage are calculated using a value function to obtain the value corresponding to any of the above actions. Correspondingly, by traversing the actions in the (k+1)th candidate action set, the (k+1)th value set can be obtained.
[0047] In some embodiments, the (k+1)th action set can be obtained in the following way: The set of values in the (k+1)th value set that are greater than or equal to the value threshold is determined as the (k+1)th target value. Then, the candidate actions in the (k+1)th candidate action set that correspond to the (k+1)th target value are determined as the (k+1)th action set.
[0048] In some embodiments, there may be a one-to-one correspondence between the action set in the (k+1)th candidate action set and the path in the (k+1)th candidate path set. Thus, after determining the (k+1)th action set, the (k+1)th search path can be determined from the (k+1)th candidate path set based on the (k+1)th action set and the aforementioned correspondence.
[0049] As can be seen from the above, in the collaborative control system for UAVs under complex interference environments provided in this application embodiment, the decision module is used to determine path constraints and, based on the path constraints and the state of the kth device, determines the (k+1)th candidate strategy through the MPC associated with the UAV. Thus, leveraging the advantages of MPC in adapting to dynamic time-varying data processing, high precision, and high flexibility, the relevance and accuracy of the (k+1)th candidate path set and the (k+1)th candidate action set in the (k+1)th candidate strategy can be improved. Furthermore, the (k+1)th candidate action set is evaluated using the PER-D3QN value function to obtain the (k+1)th value set, thereby improving the accuracy of the value in the (k+1)th value set. On the other hand, by filtering the (k+1)th candidate action set based on the (k+1)th value set, the (k+1)th action set is obtained, achieving targeted filtering of the (k+1)th candidate action set. Simultaneously, by determining the path corresponding to the (k+1)th action set in the (k+1)th candidate strategy as the (k+1)th search path, the efficiency of determining the (k+1)th search path is improved.
[0050] Based on the foregoing embodiments, in the collaborative control system of UAV under complex interference environment provided in this application embodiment, the decision module is used to determine the extreme value of the advantage of the action of the action of the (k+1)th candidate action set, determine the value of the kth state corresponding to the kth device state, and process the value of the kth state, the extreme value of the advantage of the action of the (k+1)th action, and the action advantage of the action of the action of the (k+1)th candidate action set through a value function to obtain the (k+1)th value set.
[0051] In some embodiments, the action advantage of an action in the (k+1)th candidate action set can be obtained by processing the action in the (k+1)th candidate action set through the advantage branch of Dueling DQN in PER-D3QN; correspondingly, the value of the kth state can be obtained by processing the kth device state through the value branch of Dueling DQN in PER-D3QN.
[0052] In some embodiments, the extreme value of the (k+1)th action advantage may include the maximum value of the action advantage corresponding to the action in the (k+1)th candidate action set; correspondingly, the extreme value of the (k+1)th action advantage can be obtained by statistically analyzing the values of the action advantages in the (k+1)th candidate action set.
[0053] In some embodiments, the m-th value in the (k+1)-th value set It can be calculated using equation (5): (5) in, The value of the k-th state. The action advantage of the m-th action in the (k+1)-th candidate action set. Let m be the extreme value of the advantage of the (k+1)th action, where m is an integer greater than or equal to 1.
[0054] In related technologies, the value function obtains the action score by directly superimposing the state value and the action advantage. However, in the technical solution provided in this application embodiment, the value function of the related technologies is modified by introducing the (k+1)th action advantage extreme value, which can reflect the relative advantage of actions in the candidate action set.
[0055] As can be seen from the above, in the collaborative control system for UAVs under complex interference environments provided in this application embodiment, the decision module is used to determine the (k+1)th action advantage extreme value of actions in the (k+1)th candidate action set, determine the value of the kth state corresponding to the kth device state, and process the kth state value, the (k+1)th action advantage extreme value, and the action advantage of actions in the (k+1)th candidate action set through a value function to obtain the (k+1)th value set. Thus, by introducing the (k+1)th action advantage extreme value, the (k+1)th value set calculated by the value function can comprehensively represent the relative advantages of actions in the (k+1)th candidate action set over a wider range.
[0056] Based on the foregoing embodiments, in the collaborative control system for UAVs under complex interference environments provided in this application embodiment, the path constraints are associated with the UAV's flight performance indicators, the UAV's energy consumption status, the k-th communication status of the complex interference environment, and the regional risk level in the k-th environmental situation map.
[0057] In some embodiments, flight performance indicators may include extreme values of the magnitude of flight maneuvers that the UAV can perform; for example, flight maneuvers may include flight speed and flight attitude; for example, extreme values of magnitude may include the UAV's maximum turning radius, maximum rate of climb, etc.
[0058] In some embodiments, the power status of the drone may include the remaining available power of the drone.
[0059] In some embodiments, the k-th communication state may include the signal strength and interference state of wireless communication in a complex interference environment; for example, the interference state may include electromagnetic interference, multipath fading interference, and thermal noise interference.
[0060] In some embodiments, path constraints can be determined based on thresholds corresponding to flight performance indicators, energy consumption status, k-th communication status, and regional risk level, respectively. For example, path constraints may include the UAV's turning radius being less than the extreme value of the turning radius in the flight performance indicators, the remaining available power being greater than or equal to the power threshold, the signal strength represented by the k-th communication status being greater than or equal to the strength threshold, and the regional risk level being greater than or equal to the level threshold.
[0061] For example, the paths in the (k+1)th candidate path set can vary depending on the energy consumption state and the kth communication state. For instance, if the energy consumption state indicates that the remaining available power of the drone is less than the power threshold, and the kth communication state indicates that the signal strength is less than or equal to the strength threshold, then the (k+1)th candidate path set can be determined as: the trajectory points formed by the area corresponding to the low risk level in the area risk level between the current position of the drone and the target position. At the same time, the (k+1)th candidate actions can be determined to include: the flight speed and attitude adjustment range included in the flight performance indicators.
[0062] In some embodiments, path constraints may also be related to the visibility or degree of weather disturbance corresponding to the k-th meteorological data; for example, the degree of weather disturbance may include wind resistance associated with the amplitude of airflow and visibility, and visibility may include the maximum visual detection distance corresponding to rainfall, air pollution or light pollution; for example, path constraints may include visibility being greater than or equal to a visibility threshold and wind resistance characterized by the degree of weather disturbance being less than or equal to a drag threshold.
[0063] For example, the (k+1)th candidate path set may also include: a set of spatial locations with wind resistance less than or equal to the resistance threshold, visibility greater than or equal to the visibility threshold, and regional risk level of low risk.
[0064] As can be seen from the above, in the cooperative control system for UAVs under complex interference environments provided in this application embodiment, the path constraints are related to the UAV's flight performance indicators, the UAV's energy consumption status, the k-th communication status of the complex interference environment, and the regional risk level in the k-th environmental situation map. Thus, through the above limitations, the comprehensiveness and completeness of the path constraints are improved, thereby providing precise constraints for determining the (k+1)-th candidate strategy, and further enhancing the effectiveness and relevance of the (k+1)-th candidate strategy.
[0065] Based on the foregoing embodiments, in the collaborative control system of UAV under complex interference environment provided in this application embodiment, the k-th source data also includes the k-th communication state of the complex interference environment; The decision module is also used to determine the control device based on the k-th communication state; The collaborative control module is used to control the UAV to search for target objects based on the (k+1)th search path and the (k+1)th action set through the control device.
[0066] In some embodiments, the kth communication state can be obtained in the following way: The wireless communication module associated with the UAV tracks and detects the wireless communication signal and its changing state to obtain the k-th communication state.
[0067] In some embodiments, the k-th communication state may include the change state and / or interference state of the wireless communication state within the data acquisition range at time k; for example, the k-th communication state may include the signal strength of the wireless communication signal at time k, the interference strength against the wireless communication signal and its change state, etc.
[0068] In some embodiments, the k-th communication state may further include the stability of the wireless communication link between the drone and the devices in the device set; exemplarily, the device set may include server devices and edge devices; specifically, the server devices may be deployed in the cloud or at a remote server, while the edge devices may be deployed at a location less than or equal to a distance threshold from a complex interference environment.
[0069] In some embodiments, the control device may include at least one electronic device ultimately used to control the drone to search for a target object; for example, the control device may be determined in the following ways: If the stability of the wireless communication link between the UAV and the server device, which represents the k-th communication state, is greater than or equal to the stability threshold, then the server device is identified as the control device. If the stability of the wireless communication link between the UAV and the server device, which represents the k-th communication state, is less than the stability threshold, and the stability of the wireless communication link between the UAV and the edge device is greater than or equal to the stability threshold, then the edge device is identified as the control device.
[0070] In some embodiments, the autonomous collaborative control system provided in this application can be deployed synchronously in each device of the device set. After the control device is determined, the autonomous collaborative control system deployed in the control device can be used to collaboratively control the UAV to search for target objects.
[0071] As can be seen from the above, in the collaborative control system for UAVs under complex interference environments provided in this application embodiment, the k-th source data also includes the k-th communication state of the complex interference environment. This expands the types of data in the k-th source data and improves the comprehensiveness of the k-th source data. Furthermore, the decision module is used to determine the control device based on the k-th communication state, and control the UAV to search for target objects based on the (k+1)-th search path and the (k+1)-th action set through the control device. This not only enables flexible determination and switching of the control device, but also enables diversified and flexible control of the UAV's search for target objects.
[0072] Based on the foregoing embodiments, in the collaborative control system of UAV under complex interference environment provided in this application embodiment, the kth communication state includes the kth bit error rate and the kth signal-to-noise ratio; The decision module is used to determine the control device from the device set based on the bit error rate interval where the k-th bit error rate is located and the signal-to-noise ratio interval where the k-th signal-to-noise ratio is located.
[0073] The equipment suite includes a central controller, edge controllers, and drones.
[0074] In some embodiments, the central controller may be the server device in the foregoing embodiments, and the edge controller may include the edge device in the foregoing embodiments.
[0075] In some embodiments, the bit error rate interval can be a subset of the first interval; for example, the first interval may include multiple intervals with different bit error rates.
[0076] In some embodiments, the signal-to-noise ratio interval may be a subset of the second interval; for example, the interval may include multiple intervals with different signal-to-noise ratios.
[0077] In some embodiments, a mapping relationship can be pre-established between the intervals in the first interval, the intervals in the second interval, and the devices in the device set. In this way, after determining the bit error rate interval where the k-th bit error rate is located and the signal-to-noise ratio interval where the k-th signal-to-noise ratio is located, the devices in the device set that correspond to the bit error rate interval and the signal-to-noise ratio interval can be determined as control devices by combining the above mapping relationship, the bit error rate interval, and the signal-to-noise ratio interval.
[0078] For example, when the signal-to-noise ratio of the kth bit is greater than or equal to 10 dB and the bit error rate of the kth bit is less than or equal to 10 dB. The control device can be a central controller, while edge controllers and drones can be used as auxiliary controllers; for example, if the k-th signal-to-noise ratio is greater than or equal to 5dB and less than 10dB, and the k-th bit error rate is greater than or equal to... and less than or equal to Then the edge controller can be identified as the control device; for example, if the k-th signal-to-noise ratio is less than 5dB while the k-th bit error rate is greater than... If the duration of the above states is greater than or equal to the time period threshold, the drone can be identified as a control device, and at this time, the drone switches to an offline self-control state.
[0079] As can be seen from the above, in the collaborative control system for UAVs under complex interference environments provided in this application embodiment, the k-th communication state includes the k-th bit error rate and the k-th signal-to-noise ratio (SNR). Furthermore, the decision module is used to determine the control device from the central controller, edge controller, and UAV based on the bit error rate range where the k-th bit error rate falls and the SNR range where the k-th SNR falls. Thus, by linking the determination process of the control device with the bit error rate and SNR detected by the UAV in real time, the effectiveness of the control device in controlling the UAV can be improved, and the probability of the UAV becoming uncontrollable due to fluctuations in the bit error rate and SNR can be reduced, thereby improving the stability of the UAV's search for target objects.
[0080] Based on the foregoing embodiments, in the collaborative control system of UAV under complex interference environment provided in this application embodiment, the decision module is used to determine the (k+1)th correction data, and correct at least part of the data in the (k+1)th control strategy based on the (k+1)th correction data to obtain the (k+1)th correction strategy. The collaborative control module is used to control the UAV to search for target objects based on the (k+1)th correction strategy.
[0081] In some embodiments, the (k+1)th correction data may include at least one type of data for optimizing the (k+1)th search path in the (k+1)th control strategy and / or adjusting the magnitude of actions in the (k+1)th action set.
[0082] In some embodiments, the (k+1)th correction data can be determined in the following way: Historical meteorological data and obstacle change status before time k are obtained. Based on the historical meteorological data and obstacle change status, the (k+1)th meteorological state and obstacle distribution status at time k+1 are predicted. The influence of the (k+1)th meteorological state and obstacle distribution status on the (k+1)th search path and / or (k+1)th action set is analyzed. Then, the (k+1)th correction data is determined based on the influence level. For example, if the (k+1)th meteorological state indicates that the wind resistance in the first direction is greater than or equal to the resistance threshold, the flight driving force of the (k+1)th action set along the first direction can be modified from the first force to the second force to counteract the wind resistance in the first direction. Alternatively, the flight of the (k+1)th action set along the first direction can be modified to the flight along the second direction, where the second direction may include the tailwind direction, the first direction may include the headwind direction, and the first force may be less than the second force.
[0083] Accordingly, at this time, the (k+1)th correction data corrects at least a portion of the data in the (k+1)th control strategy to overcome the negative impact of the (k+1)th weather state and the (k+1)th obstacle distribution state on at least a portion of the data in the (k+1)th control strategy.
[0084] In some embodiments, the (k+1)th correction strategy can be obtained in the following way: Based on the correction object corresponding to the (k+1)th correction data, determine the data to be corrected from the (k+1)th correction strategy. Then, based on the correction magnitude and / or correction direction represented by the (k+1)th correction data, perform corresponding corrections on the data to be corrected. At the same time, keep the data in the (k+1)th control strategy other than the data to be corrected unchanged. At this point, the data to be corrected and the data in the (k+1)th control strategy other than the data to be corrected can be determined as the (k+1)th correction strategy.
[0085] As can be seen from the above, in the collaborative control system for UAVs under complex interference environments provided in this application embodiment, the decision module is used to determine the (k+1)th correction data and correct at least some data in the (k+1)th control strategy based on the (k+1)th correction data to obtain the (k+1)th correction strategy. In this way, the adjustment and correction of the (k+1)th control strategy can be realized. On this basis, the collaborative control module controls the UAV to search for target objects based on the (k+1)th correction strategy, which can realize flexible and controllable correction of the control strategy for the UAV to search for target objects, thereby improving the controllability and flexibility of the UAV to search for target objects.
[0086] Based on the foregoing embodiments, in the collaborative control system of UAV under complex interference environment provided in this application embodiment, the k-th source data also includes the k-th communication state of the complex interference environment.
[0087] The decision module is used to obtain the k-th threat response rule and the historical search tasks of the UAV, and determine the (k+1)-th correction data based on the k-th threat response rule and the historical search tasks.
[0088] Among them, the kth threat response rule corresponds to the kth communication state.
[0089] In some embodiments, the k-th threat response rule can be a proper subset of the threat response rule set; for example, the rules in the threat response rule set may include a set of rules for improving the communication stability between the UAV and the control device in the event of deterioration of wireless communication status in a complex interference environment.
[0090] In some embodiments, the rules in the threat response rule set may include a set of wireless resource management operations performed by the drone in different communication states in order to improve the communication stability between the drone and the control device; for example, wireless resource management operations may include frequency hopping, cell handover, and power control, etc.
[0091] In some embodiments, the relationships between rules in the threat response rule set and multiple communication states can be pre-established; thus, the k-th threat response rule can be determined in the following way: After obtaining the k-th communication state, the target communication state that matches the k-th communication state can be determined from the association based on the matching status between the k-th communication state and the communication states in the association. The rule in the threat response rule set that corresponds to the target communication state is then determined as the k-th threat response rule.
[0092] For example, the k-th threat response rule can be as follows: IF threat_type="electronic_jamming" AND threat_strength>threshold_1THEN action="change_frequency_band" AND path_adjustment="avoid_circle(radius=100m)" In some embodiments, a historical search task may include regional characteristics of a historical search area by the UAV within at least one historical period; for example, regional characteristics may include the weather conditions, communication conditions, location range, obstacle distribution conditions, and the status of the target object in the historical search area.
[0093] In some embodiments, the (k+1)th correction data can be determined in the following way: Based on the task matching degree between the historical search task and the UAV search task at time k, a first correction weight is determined for the (k+1)th search path, and / or a second correction weight is determined based on the (k)th threat response rule. Then, the first correction weight and / or the second correction weight are determined as the (k+1)th correction data.
[0094] For example, a first correction weight is determined based on the meteorological matching degree between the meteorological state of the historical search task and the meteorological state represented by the k-th meteorological data, the regional matching degree between the historical search area of the historical search task and the search area of the UAV at time k, and the obstacle matching degree between the obstacle distribution state of the historical search task and the obstacle distribution state at time k. For example, the first correction weight can increase as the above matching degrees decrease, or it can decrease as the above matching degrees increase. For example, the task matching degree can include meteorological matching degree, regional matching degree, and obstacle matching degree.
[0095] For example, the second correction weight can be determined based on the communication matching degree between the historical communication state of the historical search task and the k-th communication state; for example, the second correction weight can decrease as the communication matching degree increases, or it can increase as the communication matching degree decreases.
[0096] As can be seen from the above, in the collaborative control system for UAVs under complex interference environments provided in this application embodiment, the decision module is used to obtain the k-th threat response rule corresponding to the k-th communication state, and determine the (k+1)-th correction data based on the k-th threat response rule and the UAV's historical search tasks. This achieves the correlation between the threat response rule and the communication state, improves the correlation between the (k+1)-th correction data and the threat response rule and the UAV's historical search tasks, thereby improving the correlation between the (k+1)-th correction data and the k-th communication state. This enhances the reference value of historical search tasks for the UAV's search operation at time k+1, and also characterizes the degree of influence of the k-th communication state on the (k+1)-th correction data.
[0097] Based on the foregoing embodiments, in the collaborative control system of UAV under complex interference environment provided in this application embodiment, the decision module is used to determine the (k+1)th search weight based on the historical search tasks of the UAV, modify the (k+1)th action set based on the (k+1)th search weight to obtain the (k+1)th modified action, and modify the (k)th threat response rule based on the (k+1)th rule weight corresponding to the (k+1)th search weight to obtain the (k+1)th modified rule. The decision module is also used to determine the (k+1)th correction strategy based on the (k+1)th correction action, the (k+1)th correction rule, and the (k+1)th search path.
[0098] In some embodiments, the k+1th search weight can be determined in the following way: Based on the meteorological matching degree between the meteorological state of the historical search task and the meteorological state represented by the k-th meteorological data, the regional matching degree between the historical search area of the historical search task and the search area of the UAV at time k, the obstacle matching degree between the obstacle distribution state of the historical search task and the obstacle distribution state at time k, and the communication matching degree between the historical communication state of the historical search task and the communication state at time k, the familiarity or experience level of the UAV for the search task at time k is determined. Then, based on this familiarity or experience level, the (k+1)-th search weight is determined. For example, the (k+1)-th search weight... It can be calculated using equation (6): (6) in, This represents the drone's familiarity or experience level with the search task at time k. The weights used to adjust the above levels of familiarity or experience.
[0099] Accordingly, the direction and / or magnitude of actions in the (k+1)th action set can be modified based on the (k+1)th search weight to obtain the (k+1)th modified action.
[0100] In some embodiments, the sum of the (k+1)th rule weight and the (k+1)th search weight can be 1; for example, the (k+1)th rule weight It can be calculated using equation (7): (7) In some embodiments, the (k+1)th correction rule can be obtained in the following way: Based on the weight of the (k+1)th rule, the importance of the kth threat response rule is adjusted to obtain the (k+1)th modified rule.
[0101] In some embodiments, the set of the (k+1)th correction action, the (k+1)th correction rule, and the (k+1)th search path can be determined as the (k+1)th correction strategy.
[0102] As can be seen from the above, in the collaborative control system for UAVs under complex interference environments provided in this application embodiment, the decision module is used to determine the (k+1)th search weight based on historical search tasks, and to modify the (k+1)th action set based on the (k+1)th search weight to obtain the (k+1)th modified action. In this way, by associating the (k+1)th search weight with the historical search tasks of the UAV, and modifying the (k+1)th action set based on the (k+1)th search weight, the coherence of the UAV's search actions can be improved. Furthermore, based on the (k+1)th rule weight corresponding to the (k+1)th search weight, the (k)th threat response rule is modified to obtain the (k+1)th modified rule, realizing the linkage modification of the (k+1)th action set and the (k)th threat response rule, thereby improving the integrity and comprehensiveness of the (k+1)th modification strategy.
[0103] Based on the foregoing embodiments, in the collaborative control system for UAVs in complex interference environments provided in this application embodiment, the collaborative control module is further used to track the object distance between the UAV and the target object during the process of the UAV searching for the target object, determine the search score associated with the object distance, and adjust the flight direction of the UAV based on the search score.
[0104] In some embodiments, object distance can be determined in the following ways: By extracting features and identifying targets from the environmental situation map at the k-th moment or other times, the location of the target object is determined. Then, the location of the UAV is determined by the positioning device associated with the UAV. Finally, the distance between the UAV and the target object is determined and updated in real time based on the distance between the UAV and the target object.
[0105] In some embodiments, the search score can be used to characterize the extent to which a drone continuously approaches a target object during its search for that object.
[0106] In some embodiments, the search score can be determined in the following ways: Determine the distance between the UAV and the target object at time k, determine the maximum search distance of the UAV, and determine the search score based on the UAV's search mission reward, safe search reward, the k-th distance, and the maximum search distance. Specifically, it can be shown in equation (8): (8) in, As a reward for the search task, As a reward for safe search, To maximize the search distance, The distance to the k-th object. Weights for safe search.
[0107] It should be noted that the search task reward and the safe search reward can be related to the search needs of the UAV; for example, the search needs can be determined by adjusting the weights of the objective function related to the UAV's kinematic dynamic equations; wherein, the objective function can be as shown in equation (9): (9) in, As a reward for safe search, In order to reward search efficiency, To stabilize search rewards, As a reward for communication quality, ( , , , () represents the set of weights for the objective function. The task reward value is calculated using the objective function.
[0108] Specifically, the safe search reward is negatively correlated with the number of times the UAV enters high-risk areas during the search for the target object and the distance to obstacles; the search efficiency reward is positively correlated with the completion progress of the search task and the flight speed, and negatively correlated with the energy consumption rate of the UAV; the stable search reward is negatively correlated with the flight altitude, flight speed and attitude change rate of the UAV during flight; and the communication quality reward is positively correlated with the signal-to-noise ratio of the communication link.
[0109] For example, by adjusting the above set of weights, the focus or priority of the search demand can be reflected. For instance, if the weight corresponding to the stable search reward is greater than other weights in the set of weights, it indicates that the current search task for the target object is more focused on the stable search for the target object.
[0110] In some embodiments, the flight direction of a drone can be controlled in the following ways: The search path in the correction strategy updated at each time moment is segmented to obtain a set of path segments. After the end of any segment path, the distance to the object is tracked and detected in real time, and the search score is calculated according to the object distance using equation (8). After the end of the search period corresponding to the search path, the object distance and search score corresponding to the search path are accumulated to obtain the path search score corresponding to the search path. Then, the flight direction of the UAV is fine-tuned according to the size of the path search score so that the UAV can gradually approach the target object. For example, the adjustment angle of the flight direction of the UAV can be reduced as the path search score increases.
[0111] As can be seen from the above, in the collaborative control system for UAVs in complex interference environments provided in this application embodiment, the collaborative control module is also used to track the distance between the UAV and the target object during the UAV's search for the target object. In this way, the tracking and detection of the object distance is realized during the UAV's search for the target object. Furthermore, by determining the search score associated with the object distance and adjusting the UAV's flight direction based on the search score, the UAV's flight direction can meet the overall search requirements for the target object, thereby improving the search efficiency for the target object.
[0112] Figure 2 Another structural schematic diagram of the collaborative control system provided in the embodiments of this application is shown below. Figure 2 As shown, the collaborative control system 100 may include a data acquisition and input module 201, a communication status monitoring module 202, a risk perception module 203, a path planning module 204, and an execution control module 205.
[0113] For example, the data acquisition and input module and the communication status monitoring module can be equivalent to the acquisition module in the foregoing embodiments, and are used to acquire the k-th source data of multimodal data.
[0114] For example, the risk perception module, which is equivalent to the fusion module in the aforementioned embodiments, is used to generate the k-th environmental situation map.
[0115] For example, the path planning module can be equivalent to the decision module in the aforementioned embodiments, used to generate the (k+1)th control strategy, and can also modify the (k+1)th control strategy to obtain the (k+1)th modified strategy.
[0116] For example, the execution control module can be equivalent to the cooperative control module in the foregoing embodiments, and is used to control the UAV to continuously search for the target object at time k+1 based on the k+1 correction strategy.
[0117] Through the collaborative data processing of the above modules, high-precision and high-efficiency searching of target objects can be achieved.
[0118] Figure 3 This is another structural schematic diagram of the collaborative control system provided in the embodiments of this application, as shown below. Figure 3 As shown, the data acquisition and input module 201 is used to perform the following operations: acquire multi-source data through multi-source sensors, acquire risk data, and obtain UAV performance parameters.
[0119] For example, the multi-source sensor may include radar, thermal imager, RGB image acquisition device and positioning device, and correspondingly, the multi-source data may include radar echo, thermal image, RGB image and the current position of the UAV; and by analyzing the above various data in the multi-source data, the distribution state of the kth obstacle and the state of the kth object within the data acquisition range with the current position of the UAV as the geometric center can be determined.
[0120] For example, risk data may include data that poses a risk of hindering the flight status of the drone; for example, risk data may include k-th meteorological data, which can be collected by meteorological sensors.
[0121] For example, drone performance parameters may include the drone's flight performance indicators.
[0122] For example, the risk perception module 203 can perform data mapping, data fusion, and grid division operations; wherein, data mapping can perform feature mapping in the foregoing embodiments, or, through the transformation matrix in the foregoing embodiments, map different multi-source data to the target coordinate system, and data fusion can include combining the point cloud data corresponding to the radar echo to convert the image data in the target coordinate system into the k-th environment map, and through grid division, the k-th environment map can be converted into the k-th environment situation map.
[0123] For example, the risk perception module can also use the Apriori algorithm to mine the confidence and support of the appearance of target objects in high-level regional risk areas, and determine the task priority based on the confidence and support, so as to trigger the collaborative control module to control the search process of the UAV according to the task priority.
[0124] For example, the path planning module 204 obtains the (k+1)th candidate strategy by performing the operation of determining the candidate strategy, and then it can also modify the candidate strategy to obtain the (k+1)th modified strategy.
[0125] For example, the execution control module 205 can control the drone to search based on the (k+1)th correction strategy to continuously search for the target object. Furthermore, during the continuous search for the target object, the drone can also track and calculate the search score to adjust the search direction of the drone, so that the search path of the drone can continuously approach the target object.
[0126] For example, during the process of the UAV searching for a target object, if the collaborative control system receives a new task to be processed, it can generate a task processing instruction and control the UAV to respond to the task processing instruction, search for the target object using the method provided in the aforementioned embodiments, and execute the target operation corresponding to the task processing instruction after the target object is found. For example, the task to be processed may include performing equipment maintenance operations on the equipment associated with the target object, delivering emergency medical supplies to the location of the target object, etc. Correspondingly, the target operation may include delivering equipment or tools corresponding to the equipment maintenance operation, and delivering emergency medical supplies, etc.
[0127] The steps and processes executed by each of the above modules achieve the integration of data collection, risk perception, path planning, and collaborative control, thereby improving the search efficiency for target objects.
[0128] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0129] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0130] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0131] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0132] It should be noted that the aforementioned computer-readable storage media can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various electronic devices including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0133] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0134] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0135] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware nodes. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0136] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0137] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0138] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0139] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A cooperative control system for unmanned aerial vehicles (UAVs) under complex interference environments, characterized in that, include: The acquisition module is used to acquire multimodal k-th source data; wherein, the k-th source data includes the k-th meteorological data, the k-th obstacle distribution state, and the k-th object state of the complex interference environment in which the UAV is located at the k-th time; k is an integer greater than or equal to 1; The fusion module is used to perform spatiotemporal fusion on the data in the k-th source data to obtain the k-th environmental situation map; wherein, the k-th environmental situation map includes the regional risk level of at least some areas in the complex interference environment; The decision module is used to determine the (k+1)th control strategy based on the kth device state of the UAV and the kth environmental situation map using PER-D3QN; wherein the (k+1)th control strategy includes the (k+1)th search path and the (k+1)th action set. The collaborative control module is used to control the UAV to search for target objects in the complex interference environment at time k+1, based on the (k+1)th search path and the (k+1)th action set.
2. The system according to claim 1, characterized in that, The decision module is used to determine path constraints and, based on the path constraints and the state of the kth device, determine the (k+1)th candidate strategy through the MPC associated with the UAV; wherein the (k+1)th candidate strategy includes a set of (k+1)th candidate paths and a set of (k+1)th candidate actions corresponding to the set of (k+1)th candidate paths. The decision module is further configured to evaluate the (k+1)th candidate action set using the PER-D3QN value function to obtain the (k+1)th value set, filter the (k+1)th candidate action set based on the (k+1)th value set to obtain the (k+1)th action set, and determine the path in the (k+1)th candidate strategy that corresponds to the (k+1)th action set as the (k+1)th search path.
3. The system according to claim 2, characterized in that, The decision module is used to determine the extreme value of the (k+1)th action advantage of the actions in the (k+1)th candidate action set, determine the value of the kth state corresponding to the kth device state, and process the value of the kth state, the extreme value of the (k+1)th action advantage, and the action advantage of the actions in the (k+1)th candidate action set through the value function to obtain the (k+1)th value set.
4. The system according to claim 2, characterized in that, The path constraints are associated with the UAV's flight performance indicators, the UAV's energy consumption status, the k-th communication status of the complex interference environment, and the regional risk level in the k-th environmental situation map.
5. The system according to claim 1, characterized in that, The k-th source data also includes the k-th communication state of the complex interference environment; The decision module is also used to determine the control device based on the kth communication state; The collaborative control module is used to control the UAV to search for the target object based on the (k+1)th search path and the (k+1)th action set through the control device.
6. The system according to claim 5, characterized in that, The kth communication state includes the kth bit error rate and the kth signal-to-noise ratio; The decision module is used to determine the control device from the device set based on the bit error rate interval where the k-th bit error rate is located and the signal-to-noise ratio interval where the k-th signal-to-noise ratio is located; wherein the device set includes a central controller, an edge controller and the UAV.
7. The system according to claim 1, characterized in that, The decision module is used to determine the (k+1)th correction data, and based on the (k+1)th correction data, correct at least a portion of the data in the (k+1)th control strategy to obtain the (k+1)th correction strategy. The collaborative control module is used to control the UAV to search for the target object based on the (k+1)th correction strategy.
8. The system according to claim 7, characterized in that, The k-th source data also includes the k-th communication state of the complex interference environment; The decision module is used to obtain the k-th threat response rule and the historical search tasks of the UAV, and determine the (k+1)-th correction data based on the k-th threat response rule and the historical search tasks; wherein the k-th threat response rule corresponds to the k-th communication state.
9. The system according to claim 7, characterized in that, The decision module is used to determine the (k+1)th search weight based on the historical search tasks of the UAV, modify the (k+1)th action set based on the (k+1)th search weight to obtain the (k+1)th modified action, and modify the (k+1)th threat response rule based on the (k+1)th rule weight corresponding to the (k+1)th search weight to obtain the (k+1)th modified rule. The decision module is further configured to determine the (k+1)th correction strategy based on the (k+1)th correction action, the (k+1)th correction rule, and the (k+1)th search path.
10. The system according to any one of claims 1 to 9, characterized in that, The collaborative control module is also used to track the object distance between the UAV and the target object during the process of the UAV searching for the target object, determine the search score associated with the object distance, and adjust the flight direction of the UAV based on the search score.