Multi-uav target coverage method and device based on visual language model

By fusing drone status and mission commands through a visual language model, motion commands are generated, solving the problem of matching the coverage range of multiple drones with high-level intentions and achieving real-time and accurate coverage.

CN122151958APending Publication Date: 2026-06-05INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF AUTOMATION CHINESE ACAD OF SCI
Filing Date
2026-05-07
Publication Date
2026-06-05

Smart Images

  • Figure CN122151958A_ABST
    Figure CN122151958A_ABST
Patent Text Reader

Abstract

The application provides a multi-unmanned aerial vehicle target coverage method and device based on a visual language model, applied to the technical field of unmanned aerial vehicles, and the method comprises the following steps: acquiring a static overhead view of a task area where a group of unmanned aerial vehicles are located, first state information of each unmanned aerial vehicle at a current time, second state information of a plurality of observation objects, and a task instruction; fusing the static overhead view, the first state information of each unmanned aerial vehicle, and the second state information of each observation object to obtain target fusion features; processing the target fusion features and the task instruction by using a visual language model to obtain movement instructions of each unmanned aerial vehicle; and each unmanned aerial vehicle moves according to the movement instruction thereof. In the application, the visual language model can determine the movement instructions of each unmanned aerial vehicle after understanding the states of each unmanned aerial vehicle and each observation object in the task area and understanding the task instruction issued by a user, each unmanned aerial vehicle moves according to the movement instruction thereof, and the coverage range of the multi-unmanned aerial vehicle is matched with a high-level dynamic intention in real time and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) technology, and in particular to a method and apparatus for multi-UAV target coverage based on a visual language model. Background Technology

[0002] In key areas such as smart city management, public safety, and emergency rescue, utilizing drone swarms for collaborative management, communication relay, or material delivery of multiple dispersed ground targets has become an efficient technological means. Typical application scenarios include: conducting simultaneous aerial reconnaissance of multiple affected areas after natural disasters (such as earthquakes and floods) to quickly assess the distribution of the disaster and locate trapped individuals; providing real-time, three-dimensional control of venue entrances and exits, core gathering areas, and surrounding transportation hubs during large-scale events or gatherings to ensure public safety; or rapidly constructing an aerial mobile communication network during temporary events to provide stable broadband signal coverage for multiple temporary command points or user groups. The essence of these tasks requires a swarm of multiple drones to autonomously, collaboratively, persistently, and efficiently provide "coverage" services to multiple static or mobile targets in dynamic urban scenarios with numerous buildings and intricate road networks.

[0003] Compared to the traditional single-drone single-point operation mode, the above-mentioned multi-drone collaborative coverage mission for multiple targets faces more challenges, such as the challenge of mission dynamism: the target points to be covered may have different priorities, or their status (such as location) may change dynamically over time. Commanders may flexibly adjust the mission focus through natural language commands based on the real-time situation (e.g., "immediately strengthen the control of the target on the east side"), requiring the decision-making system to have the ability to interpret and respond to the intentions of higher authorities in real time.

[0004] Traditional optimization methods, such as model predictive control based on optimal control or mixed-integer programming based on operations research, typically abstract the problem into a mathematical model that seeks the optimal solution under specific constraints. These methods have significant drawbacks. For example, they cannot understand and process flexible, semantically rich instructions given by commanders in natural language (such as "prioritize communication in the hospital area and avoid the shadows of tall buildings"). A gap exists between system behavior and human intent, making it difficult to achieve real-time and accurate matching between the coverage area of ​​multiple drones and dynamic intentions at higher levels. Summary of the Invention

[0005] This invention provides a multi-UAV target coverage method and apparatus based on a visual language model to solve the shortcomings of existing technologies where the inability to understand task instructions makes it difficult to match the coverage range of multiple UAVs with the dynamic intent of higher layers in real time and accurately. After the visual language model understands the state of each UAV and each observed object in the task area based on target fusion features and understands the task instructions issued by the user, it determines the motion instructions of each UAV. Each UAV moves according to its own motion instructions, thus realizing the real-time and accurate matching of the coverage range of multiple UAVs with the dynamic intent of higher layers.

[0006] This invention provides a method for multi-UAV target coverage based on a visual language model, comprising the following steps: Obtain a static top-down view of the mission area where the UAV group is located, the first state information of each UAV in the UAV group at the current moment, the second state information of several observed objects at the current moment, and the mission instructions issued by the user; The static top view, each of the first state information and each of the second state information are fused to obtain the target fusion feature; The target fusion features and the task instructions are processed using a visual language model to obtain the motion instructions for each of the UAVs; Control each of the aforementioned drones to move according to the corresponding motion commands.

[0007] According to the present invention, a multi-UAV target coverage method based on a visual language model is provided, wherein fusing the static top view, each of the first state information and each of the second state information to obtain target fusion features includes: Acquire the target environment features of the static top view, the first encoded features of each first state information, and the second encoded features of each second state information; For each of the first encoded features, cross-attention processing is performed on the target environment features and the first encoded features to obtain the first-order fusion features corresponding to the first encoded features; Self-attention processing is performed on the first feature group obtained by combining the first-order fusion features to obtain the second-order fusion features corresponding to each of the first coding features; For each of the second-order fusion features, the second-order fusion feature is cross-attention processed with each of the second coding features to obtain the third-order fusion feature corresponding to the second-order fusion feature; The target fusion feature is obtained based on each of the first-order fusion features, each of the second-order fusion features, and each of the third-order fusion features.

[0008] According to the present invention, a multi-UAV target coverage method based on a visual language model is provided, wherein obtaining the target environment features of the static top view includes: Feature extraction is performed on the static top view to obtain the initial environmental features of the static top view, which include sub-features of multiple regions in the static top view; Determine the first proportion of pixels belonging to roads and the second proportion of pixels belonging to buildings in each region of the static top view; For each region, the sub-features of the region are updated based on a first proportion of pixels belonging to roads and a second proportion of pixels belonging to buildings within the region. Based on the latest sub-features of each region, the target environment features of the static top view are obtained.

[0009] According to a multi-UAV target coverage method based on a visual language model provided by the present invention, for each region, the sub-features of the region are updated according to a first proportion of pixels belonging to roads and a second proportion of pixels belonging to buildings within the region, including: It was determined that the second proportion within the region was greater than or equal to the preset proportion. The sub-features of the region are set to preset values, which are used to characterize that the region is impassable.

[0010] According to the present invention, a multi-UAV target coverage method based on a visual language model is provided, wherein obtaining the target environment features of the static top view based on the latest sub-features of each region includes: For each region, the latest sub-features of the region, the first proportion, and the second proportion are concatenated to obtain the environmental features of the region. The environmental features of each region are combined according to the positional relationship between the regions to obtain the target environmental features.

[0011] According to the present invention, a multi-UAV target coverage method based on a visual language model is provided, wherein obtaining the target fusion feature based on each of the first-order fusion features, each of the second-order fusion features, and each of the third-order fusion features includes: For each UAV, the first-order fusion feature, second-order fusion feature and third-order fusion feature corresponding to the UAV are spliced ​​together to obtain spliced ​​data. Each of the spliced ​​data is fused using a fully connected neural network to obtain the spatiotemporal feature data corresponding to each of the spliced ​​data. The spatiotemporal feature data are combined to obtain the target fusion feature.

[0012] According to the present invention, a multi-UAV target coverage method based on a visual language model is provided, wherein the method processes the target fusion features and the task instructions using a visual language model to obtain the motion instructions of each UAV, including: Obtain the current real-time top view of the task area; The real-time top-down view is encoded using the visual encoder in the visual language model to obtain visual features, and the task instructions are encoded using the text encoder in the visual language model to obtain text features. The text features, the visual features, and the target fusion features are fused together to obtain multimodal fusion features; Based on the multimodal fusion features, motion commands for each UAV are obtained.

[0013] According to the present invention, a method for multi-UAV target coverage based on a visual language model is provided, the method further comprising a training step of the visual language model, the training step comprising: Initialize the parameters of the visual language model and the course distribution parameters, wherein the course distribution parameters are used to indicate the sampling probability of the training dataset within different difficulty ranges; The preset training steps are executed repeatedly until the visual language model converges or reaches the maximum number of training steps. The preset training steps include: Based on the task distribution parameters required for this training, the initial dataset required for this training is obtained, and the task distribution parameters include the latest course distribution parameters; The trajectory data of a virtual drone moving according to training motion commands in a simulation environment is determined. The training motion commands are motion commands planned by the visual language model to be trained based on the initial dataset. The parameters of the visual language model are adjusted based on the trajectory data; The course distribution parameters are adjusted at preset intervals of a certain number of training cycles.

[0014] The present invention also provides a multi-UAV target coverage device based on a visual language model, comprising the following modules: The data acquisition module is used to acquire a static top view of the mission area where the UAV group is located, the first state information of each UAV in the UAV group at the current moment, the second state information of several observed objects at the current moment, and the mission instructions issued by the user. The feature fusion module is used to fuse the static top view, each of the first state information and each of the second state information to obtain the target fused features; The motion command generation module is used to process the target fusion features and the task commands using a visual language model to obtain motion commands for each of the UAVs; The drone control module is used to control each of the drones to move according to the corresponding motion commands.

[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multi-UAV target coverage method based on the visual language model as described above.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-UAV target coverage method based on a visual language model as described above.

[0017] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the multi-UAV target coverage method based on a visual language model as described above.

[0018] This invention provides a multi-UAV target coverage method and device based on a visual language model. By fusing the first state information of each UAV at the current moment, the second state information of several observed objects, and a static top view of the task area, the visual language model can understand the state of each UAV and each observed object in the task area based on the target fusion features. After understanding the task instructions, it can determine the movement instructions of each UAV, so that after each UAV moves according to its own movement instructions, the coverage range of each UAV is matched with the dynamic intent of the upper layer in real time and accurately. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 This is one of the flowcharts of the multi-UAV target coverage method based on a visual language model provided by the present invention.

[0021] Figure 2 This is the second flowchart of the multi-UAV target coverage method based on a visual language model provided by the present invention.

[0022] Figure 3This is the third flowchart of the multi-UAV target coverage method based on visual language model provided by the present invention.

[0023] Figure 4 This is a schematic diagram of the structure of the multi-UAV target coverage device based on a visual language model provided by the present invention.

[0024] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention.

[0025] Figure label: 400: Multi-UAV target coverage device based on visual language model; 401: Data acquisition module; 402: Feature fusion module; 403: Motion command generation module; 404: UAV control module; 510: Processor; 520: Communication interface; 530: Memory; 540: Communication bus. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0027] The following is combined with Figures 1 to 3 This invention describes a multi-UAV target coverage method based on a visual language model. Figure 1 This is one of the flowcharts illustrating the multi-UAV target coverage method based on a visual language model provided by the present invention, such as... Figure 1 As shown, the method includes the following steps: Step 101: Obtain a static top-down view of the mission area where the UAV group is located, the first state information of each UAV in the UAV group at the current moment, the second state information of several observed objects at the current moment, and the mission instructions issued by the user.

[0028] A drone swarm comprises multiple drones. The multi-drone target coverage method based on a visual language model provided by this invention can be applied to one of the drones in the swarm or to an electronic device that establishes a communication connection with the drone swarm.

[0029] The mission area includes the area that the drone group needs to cover. That is, the mission area can be the area that the drone group needs to cover or an area larger than the area that the drone group needs to cover.

[0030] A static top-down view can be obtained by fusing a pre-acquired top-down view of the mission area or by combining top-down views collected by each UAV at the current moment. The pre-acquired top-down view may not include temporary moving obstacles within the mission area at the current moment, such as vehicles temporarily parked within the mission area. This embodiment uses a pre-acquired top-down view of the mission area as an example of a static top-down view.

[0031] The initial status information of each drone at the current moment may include, but is not limited to, one or more of the following: the drone's position, speed, heading, and battery level.

[0032] "Several observation objects" refers to one or more observation objects, which are the target objects that each UAV needs to observe. Target objects can include static objects and / or dynamic objects. For example, a static object can be an object that cannot be moved from a fixed location, such as a certain area or building, while a dynamic object can be an object that can be moved. The second state information of each observation object at the current moment can include the position and priority of the observation object.

[0033] Mission instructions can be issued by the drone team commander via voice or text.

[0034] Step 102: Fuse the static top view, each first state information and each second state information to obtain the target fusion feature.

[0035] Optionally, step 102 can be implemented by: acquiring the target environment features of a static top-down view, the first encoded features of each first state information, and the second encoded features of each second state information. Then, the target environment features, each first encoded feature, and each second encoded feature are fused to obtain the target fusion features. For example, a three-layer cascaded attention fusion network can be used to sequentially guide each UAV to focus on the environmental region most relevant to it (first layer), understand the intentions and states of other UAVs in the cluster to achieve implicit negotiation (second layer), and identify the observation objects that need to be focused on at the moment (third layer).

[0036] Step 103: Process the target fusion features and task commands using a visual language model to obtain the motion commands for each UAV.

[0037] The visual language model can be pre-trained, with inputs including target fusion features and task commands, and outputs motion commands for each UAV or a selection of motion commands for each UAV. The motion command for each UAV is determined from the selection of motion commands available to that UAV.

[0038] Step 104: Control each drone to move according to the corresponding motion command.

[0039] Step 104 can be implemented by sending motion commands to each drone. After receiving the motion commands, each drone can move according to the received motion commands, thereby adjusting the coverage area of ​​each drone.

[0040] In some application scenarios, the coverage objective of multiple drones includes image acquisition of the mission area. In this case, the coverage area of ​​each drone can be the image acquisition range covered by the camera mounted on the drone. In other application scenarios, the coverage objective of multiple drones is to build broadband coverage for the mission area. Therefore, the coverage area of ​​the drones during the process of building broadband signal coverage can be the broadband signal coverage range of the drones. It can be seen that different application scenarios have different coverage objectives, and different coverage objectives correspond to different coverage ranges.

[0041] In the above scheme, by fusing the first state information of each UAV at the current moment, the second state information of several observed objects, and the static top view of the task area, the visual language model can understand the state of each UAV and each observed object in the task area based on the target fusion features. After understanding the task instructions, it can determine the motion instructions of each UAV, so that after each UAV moves according to its own motion instructions, the coverage of each UAV matches the dynamic intent of the high layer in real time and accurately.

[0042] In one possible embodiment, step 102 described above may include, for example: Figure 2 The following steps are shown: Step 201: Obtain the target environment features of the static top view, the first encoding features of each first state information, and the second encoding features of each second state information.

[0043] Here, taking the first state information, which includes the drone's position, speed, and heading in the horizontal plane, as an example, we define the first... The status of the drone is ,in For the current moment Next The position of the drone in the horizontal plane. Indicates the current time Next The speed of the drone Indicates the current time Next The flight path of the drone.

[0044] The second state information includes the position and priority of the observed object. Define the first... The state of each observed object is ,in Indicates the first The observed object at the current time Two-dimensional coordinates, Indicates the first The observed object at the current time Priority. Specifically, a two-dimensional coordinate system can be pre-constructed for the task area, with the position in the two-dimensional coordinate system corresponding to the pixel point in the top view. Knowing the two-dimensional coordinates of the observed object allows us to determine the position of the corresponding pixel point in the static top view.

[0045] Two independent state information encoders can be used to encode each first state information and each second state information, respectively. These two independent state information encoders can each be a fully connected neural network. For example, the first state information encoder encodes the first state information... The first encoded feature obtained by encoding the first state information of the UAV The second state information encoder for the first The second encoded feature is obtained by encoding the second state information of the observed objects. .

[0046] In one possible embodiment, the implementation of obtaining the target environment features of the static top view in step 201 may include the following steps: First, feature extraction is performed on the static top view to obtain the initial environmental features of the static top view. These initial environmental features include sub-features of multiple regions within the static top view. Each region can be a grid region obtained by dividing the static top view into a mesh. For example, the feature map... The space is evenly divided into OK The columns are arranged in a regular grid, with each grid corresponding to a region in the static top view.

[0047] Optionally, define a static top view. ,in Superscript represents the set of real numbers. and These represent the height and width in pixels of the static top view, respectively, with 3 representing the red, green, and blue color channels.

[0048] A pre-trained deep convolutional neural network can be used as the backbone feature extractor. The backbone feature extractor Will As input, the output is a dense visual feature map. ,in superscript and It refers to the spatial height and width of the feature map. It is the number of feature channels.

[0049] The average pooling method can be used to process the visual feature map to obtain sub-features within each region. For example, for the feature map located at the th region... line, number For a column region, the following steps can be performed: Utilize spatial average pooling from... Extract the sub-features corresponding to this region , This refers to the number of feature channels. The sub-features here can be in the form of vectors.

[0050] Secondly, determine the first proportion of pixels belonging to roads and the second proportion of pixels belonging to buildings in each area of ​​the static top view.

[0051] Optionally, the first proportion can be the number of pixels belonging to roads in the area or the proportion of pixels belonging to roads in the area. The second proportion can be the number of pixels belonging to buildings in the area or the proportion of pixels belonging to buildings in the area. In this embodiment, the first proportion is the proportion of pixels belonging to roads in the area, and the second proportion is the proportion of pixels belonging to buildings in the area.

[0052] For example, a static top-down view is segmented to obtain the probability that each pixel in the static top-down view belongs to a road and the probability that it belongs to a building. In one application scenario, to obtain more explicit geographic semantic information, this invention uses two pre-trained semantic segmentation models to segment the static top-down view. Processing is performed. This includes the road segmentation model. Output a probability map of each pixel belonging to a road. Architectural segmentation model Output a probability map of each pixel belonging to a building. Probability diagram The graph records the probability that each pixel in the static top-down view belongs to a road. The data records the probability that each pixel in the static top-down view belongs to a building.

[0053] Then, based on the probability that each pixel belongs to a road and the probability that it belongs to a building, the category of each pixel is determined. For example, for the pixel located at the... line, number In terms of the column region, it can be seen from the probability map. and probability graph The calculation determines the percentage of pixels belonging to roads within the specified area. and the percentage of pixels belonging to the building .

[0054] Then, for each region, the sub-features of the region are updated based on the first proportion of pixels belonging to roads and the second proportion of pixels belonging to buildings within the region.

[0055] Optionally, for each region, the sub-features of the region are updated based on a first proportion of pixels belonging to roads and a second proportion of pixels belonging to buildings within the region. This includes determining that the second proportion within the region is greater than or equal to a preset proportion. The sub-features of the region are then set to preset values, which are used to characterize the region as impassable.

[0056] The preset value can be a given extremely large positive number. For example, for the value located at the _____, ... line, number For a given area, if the percentage of pixels belonging to buildings within that area... Exceeding a given threshold Then the sub-features of this region Setting a given extremely large positive number indicates that the grid is impassable. If the proportion of pixels belonging to buildings in the region exceeds a given threshold, it indicates that the building is too close to the drone's camera, meaning the building is too tall for the drone to pass through. If the second proportion in the region is determined to be less than the preset proportion, the sub-features of the region can be kept unchanged.

[0057] Based on the latest sub-features of each region, the target environment features of the static top view are obtained.

[0058] This can be achieved by combining the sub-features according to the positional relationships between the regions to obtain the target environment features of the static top-down view. Alternatively, the implementation of obtaining the target environment features of the static top-down view based on the latest sub-features of each region may include the following steps: For each region, the latest sub-features, the first proportion, and the second proportion are concatenated to obtain the region's environmental features; the environmental features of each region are then combined according to their positional relationships to obtain the target environmental features.

[0059] In other words, for each region, if its sub-features have been updated, the updated sub-features, the first proportion, and the second proportion are concatenated to obtain the environmental features of that region; if the sub-features of that region have not been updated, the sub-features, the first proportion, and the second proportion are concatenated to obtain the environmental features of that region.

[0060] For example, for the position located at the line, number For the region of the column, the above features are aggregated into a comprehensive environmental feature vector. ,in This indicates a splicing operation.

[0061] Optionally, the environmental characteristics of each region can be reduced in dimensionality. For example, this can be achieved using a fully connected layer. The environmental features of each region are fused and dimensionality reduced to obtain the final environmental features of that region. These environmental features can be in the form of semantic feature vectors, located at the... line, number The final environmental characteristics of the column area ,in These are preset environmental characteristic dimensions. The final environmental characteristics of each region... The target environmental characteristics are obtained by combining the regions according to their spatial relationships. .

[0062] Step 202: For each first coding feature, perform cross-attention processing on the target environment feature and the first coding feature to obtain the first-order fusion feature corresponding to the first coding feature.

[0063] Specifically, a first cross-attention layer can be used to perform cross-attention processing on the target environment features and the first encoded features to obtain the first-order fusion features corresponding to the first encoded features. For example, for each UAV, the first encoded features after encoding the UAV's first state information can be used as the source of the query, and the target environment features can be used as a set of key-value pairs. Then, the query vector, key vector, and value vector are calculated. The key vector and value vector are then flattened in the spatial dimension to obtain the corresponding matrix, and the attention weight of each position in the matrix is ​​calculated. Here, each position in the matrix corresponds to a region in a static top-down view.

[0064] For example, for the first The drone, in its coded state As the source of the query. Target environment characteristics. Treating it as a set of key-value pairs, the query vector corresponding to the first cross-attention layer can be calculated. Key vector Value vector ,in , and These are the learnable weight matrices. It is the dimension of the attention space corresponding to the first cross-attention layer. Indicates the dimension of each first encoded feature. This represents the preset environmental feature dimensions. and Flattened in spatial dimension The matrix is ​​such that each position corresponds to a region. The attention weights can be calculated using formula (1): Formula (1) In formula (1) Indicates for the first For drones, the area ( Attention weights for environmental features. It is a query vector. It is a region The corresponding key vector. This indicates transpose. It is the dimension of the attention space corresponding to the first cross-attention layer.

[0065] For all regions Perform probabilistic normalization to obtain the normalized attention weights. The normalization formula can be found in formula (2): Formula (2) In formula (2) Indicates for the first For drones, the area Attention weights normalized to environmental characteristics Indicates for the first For drones, the area The attention weights before environmental feature normalization, the summation range of each region in formula (2) is the first ( ) to the indivual.

[0066] The first-order fusion feature corresponding to the first encoded feature is obtained by weighted summation, as shown in formula (3): Formula (3) In formula (3) Indicates the first The first-order fusion feature corresponding to the first encoded feature of the drone Indicates for the first For drones, the area The normalized attention weights at this position, Indicates the region The value vector of environmental features.

[0067] Step 203: Perform self-attention processing on the first feature group obtained by combining each first-order fusion feature to obtain the second-order fusion feature corresponding to each first coding feature.

[0068] Specifically, a multi-head self-attention layer can be used to perform self-attention processing on the first feature group obtained by combining the first-order fusion features, resulting in second-order fusion features corresponding to each first-order encoded feature. This promotes information interaction and intent coordination among drones in the drone group, enabling each drone to perceive the focus of attention of its companions. The first-order fusion features corresponding to each first-order encoded feature output by the aforementioned cross-attention layer are stacked to form the first feature group. , Indicates the number of first-order fusion features. This refers to the dimension of the attention space. Inputting it into a standard multi-head self-attention layer outputs the second feature set. Second feature group The Middle The line represents the first The second-order fusion feature corresponding to the first coding feature of each drone In this multi-head attention layer, the query vector, key vector, and value vector all originate from the same input matrix. Therefore, it can be seen that the second-order fusion feature incorporates the collaborative information between the drones in the drone group.

[0069] Step 204: For each second-order fusion feature, perform cross-attention processing on the second-order fusion feature and each second coding feature to obtain the third-order fusion feature corresponding to the second-order fusion feature.

[0070] A second cross-attention layer can be used to perform cross-attention processing on each second-order fusion feature and each second-order encoding feature to obtain the third-order fusion feature corresponding to the second-order fusion feature.

[0071] Specifically, for each UAV, the second-order fusion feature corresponding to that UAV is used as the source of the query vector, and the set of second-order encoded features of each observed object is used as the key and value. The query vector, key vector, and value vector corresponding to the second cross-attention layer are calculated. In this way, the third-order fusion feature corresponding to each second-order fusion feature is determined.

[0072] For example, for the first The drone, with its second-order fusion characteristics The set of second-encoded features of each observed object serves as the source of the query vector. Using keys and values, calculate the query vector corresponding to the second cross-attention layer. Key vector Sum value vector ,in, , and These are the learnable weight matrices. , It is the dimension of the attention space corresponding to the second cross-attention layer. These are the dimensions of each second encoded feature. It is the dimension of the attention space corresponding to the first cross-attention layer mentioned above.

[0073] The calculation method for the attention weights in this cross-attention layer can be found in formula (4): Formula (4) In formula (4) Indicates for the first Regarding drones Attention weights for the second encoded feature, This represents the calculated query vector. Indicates the first The key vector corresponding to the second encoded feature This indicates transpose. This is the dimension of the attention space corresponding to the second cross-attention layer. Probability normalization is performed to obtain .

[0074] The weighted summation yields the third-order fusion feature corresponding to the second-order fusion feature, as detailed in formula (5): Formula (5) In formula (5) Indicates the first The third-order fusion feature corresponding to the second-order fusion feature of the drone. Indicates for the first Regarding drones The attention weights after normalization of the second encoded feature Indicates the first The value vector corresponding to each second encoded feature. It is the dimension of the attention space corresponding to the second cross-attention layer.

[0075] Step 205: Based on each first-order fusion feature, each second-order fusion feature, and each third-order fusion feature, obtain the target fusion feature.

[0076] In one possible embodiment, step 205 above may include the following steps: For each UAV, the first-order, second-order, and third-order fusion features corresponding to the UAV are concatenated to obtain concatenated data. Each concatenated data is then fused using a fully connected neural network to obtain the spatiotemporal feature data corresponding to each concatenated data. The spatiotemporal feature data are then combined to obtain the target fusion feature.

[0077] In other words, for each drone, the first-order fusion feature, second-order fusion feature, and third-order fusion feature corresponding to that drone are concatenated and then processed through a fully connected neural network. The target fusion sub-features of each UAV are obtained. Target fusion sub-features of drones The target fusion features are obtained by combining the fusion sub-features of each target. . This represents the dimension of the target fusion sub-features.

[0078] In one possible embodiment, step 103 above may include the following steps: Obtain the current real-time top-down view of the task area; encode the real-time top-down view using the visual encoder in the visual language model to obtain visual features, and encode the task instructions using the text encoder in the visual language model to obtain text features; fuse the text features, visual features, and target fusion features to obtain multimodal fusion features; and obtain the motion instructions of each UAV based on the multimodal fusion features.

[0079] The current real-time top-down view of the mission area can be obtained by stitching or other fusion processing of images acquired by each UAV in the mission area at the current moment, or by rendering from a simulation environment. Define the real-time top-down view. The dimension is , and This indicates the height and width in pixels of the real-time top-down view, with 3 representing the three color channels: red, green, and blue.

[0080] In one application scenario, the visual encoder module Image Encoding to obtain visual features Text encoder Task instructions Word segmentation and encoding yield text features Ultimately, visual features Text features and the aforementioned target fusion features Fusion is performed to form multimodal fusion features .

[0081] The above-mentioned method for obtaining motion commands for each UAV based on multimodal fusion features can be: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] The input decoder obtains the motion feature vector for each drone. For each drone, the visual language model assigns a unique decision head. For example, for the... For a single drone, the motion feature vector Enter a special one for the first Decision-making head of drone design The drone's actions are defined as discrete actions, including forward, backward, left, right, and stop. The decision head's output is a probability distribution over the action space. Determine the... The drone at the current moment Movement instructions The method can be found in formula (6): Formula (6) Where A = {forward, backward, left, right, stop} is the set of discrete actions. The action with the highest probability is selected. As the first The drone at the current moment Movement instructions.

[0082] For example, the visual language model takes text features, visual features, and target fusion features as inputs, and uses its internal cross-modal attention mechanism to achieve deep alignment and reasoning of visual signals, language semantics, and internal scene understanding to obtain motion commands for each UAV.

[0083] In one possible embodiment, the visual language model can be trained based on automatic curriculum learning. The method also includes a training step for the visual language model, which includes: The parameters of the visual language model and the course distribution parameters are initialized. The course distribution parameters indicate the sampling probability of the training dataset within different difficulty ranges. Pre-set training steps are executed iteratively until the visual language model converges or reaches the maximum number of training steps. These pre-set training steps include: obtaining the initial dataset required for this training based on the task distribution parameters, which include the latest course distribution parameters; determining the trajectory data of the virtual drone moving according to training motion commands in the simulation environment, where the training motion commands are the motion commands planned by the visual language model to be trained based on the initial dataset; adjusting the parameters of the visual language model based on the trajectory data; and adjusting the course distribution parameters at preset intervals of a certain number of training cycles. Each initial dataset includes top-view samples, target fusion feature samples, and task command samples.

[0084] For example, a parameterized task space is defined. An example of a training task (training dataset). Defined by a set of difficulty parameters, including the intensity of environmental dynamics. (0 represents a completely static target, 1 represents a target containing a large number of fast-moving targets), number of observed objects Number of drones , Represents the set of positive integers. Initial training task. Composed of the simplest parameters, such as The current task distribution parameters include the intensity of environmental dynamics. and course distribution parameters In other words, task distribution. ,in These are distribution parameters (such as the sampling probability for different difficulty ranges). Initialize model parameters. and course distribution parameters This means that the visual language model is first trained using the simplest task instances.

[0085] The above describes how the initial dataset required for this training is obtained based on the task distribution parameters. The task distribution parameters can include the latest course distribution parameters in the form of sampling from the task space based on the task distribution parameters to obtain the initial dataset required for this training, which is equivalent to obtaining a batch of training task instances.

[0086] Then, in the simulation environment corresponding to each sampled training task instance, the visual language model is run, and trajectory data is collected. This trajectory data may include top-down view samples, task command samples, target fusion feature samples, and performance parameters. These performance parameters can be determined based on performance metrics obtained from testing the visual language model on test task instances, specifically by considering target coverage and UAV flight costs. The trajectory data can be stored in an experience replay pool.

[0087] Trajectory data is sampled from the experience replay pool, and the parameters of the visual language model are updated using a proximal policy optimization algorithm. .

[0088] The above-mentioned method of adjusting the course distribution parameters at preset intervals of a certain number of training periods can include the following operations at preset intervals of a certain number of training periods: First, perform performance evaluation: The visual language model (with parameters...) In a set of verification task instances covering different difficulty levels Test on, This indicates the number of verification task instances. For each verification task instance... Calculate the performance index, i.e., the reward value. ,in Indicates target coverage (the higher the coverage) The larger ( Indicates the cost of drone flight (the higher the flight cost) (The larger the value). Then, perform difficulty measurement: for each difficulty dimension (i.e., number of observed objects, number of drones, environmental dynamics), analyze the relationship between performance and difficulty. For example, for the dimension of the number of observed objects, find the largest... So that in all If the average performance on the verification task is greater than or equal to the performance threshold, then this performance boundary is denoted as . Similarly, the dynamic boundary of the environment is obtained. and the boundary of the number of drones Then update the curriculum: based on the current skill level. , and The scheduler updates the task sampling distribution. The principle of updating is to sample tasks with a higher probability whose difficulty parameters are slightly higher than the current capability boundary. Specifically, heuristic rules can be used, for example, sampling... The probability distribution is concentrated in Within the interval, It is a preset difficulty adjustment threshold. Adjustments should be made based on this principle.

[0089] Please see Figure 3 The multi-UAV target coverage method based on a visual language model provided by the present invention may further include the following steps: Step 301: Obtain a static top-down view of the mission area where the UAV group is located, the first state information of each UAV in the UAV group at the current moment, the second state information of several observed objects at the current moment, and the mission instructions issued by the user.

[0090] The implementation method of step 301 can be found in the implementation method of step 101 above, and will not be repeated here.

[0091] Step 302: Obtain the target environment features of the static top view, the first encoding features of each first state information, and the second encoding features of each second state information.

[0092] The implementation method of step 302 can be found in the implementation method of step 201 above, and will not be repeated here.

[0093] Step 303: Input the target environment features and each first coding feature into the cross-attention layer of the first layer to obtain the first-order fusion features corresponding to each first coding feature.

[0094] The implementation method of step 303 can be found in the implementation method of step 202 above, and will not be repeated here.

[0095] Step 304: Input the first feature group obtained by combining the first-order fusion features into the multi-head self-attention layer of the second layer to obtain the second-order fusion features corresponding to each first coding feature.

[0096] The implementation method of step 304 can be found in the implementation method of step 203 above, and will not be repeated here.

[0097] Step 305: Input each second-order fusion feature and each second-order encoding feature into the cross-attention layer of the third layer to obtain the third-order fusion feature corresponding to each second-order fusion feature.

[0098] The implementation method of step 305 can be found in the implementation method of step 204 above, and will not be repeated here.

[0099] Step 306: Based on each first-order fusion feature, each second-order fusion feature, and each third-order fusion feature, obtain the target fusion feature.

[0100] The implementation method of step 306 can be found in the implementation method of step 205 above, and will not be repeated here.

[0101] Step 307: Input the real-time top view of the mission area, target fusion features, and mission commands into the visual encoder to obtain the motion commands of each UAV.

[0102] The implementation method of step 307 can be found in the implementation method of step 103 above, and will not be repeated here.

[0103] Step 308: Control each drone to move according to the corresponding motion command.

[0104] The implementation method of step 308 can be found in the implementation method of step 104 above, and will not be repeated here.

[0105] The following describes the multi-UAV target coverage device based on a visual language model provided by the present invention. The multi-UAV target coverage device based on a visual language model described below can be referred to in correspondence with the multi-UAV target coverage method based on a visual language model described above. Figure 4 As shown, the multi-UAV target coverage device 400 based on a visual language model includes the following modules: Data acquisition module 401 is used to acquire a static top view of the task area where the UAV group is located, the first state information of each UAV in the UAV group at the current moment, the second state information of several observed objects at the current moment, and the task instructions issued by the user. The feature fusion module 402 is used to fuse the static top view, each of the first state information and each of the second state information to obtain the target fusion feature; The motion command generation module 403 is used to process the target fusion features and the task commands using a visual language model to obtain the motion commands of each of the UAVs; The UAV control module 404 is used to control each of the UAVs to move according to the corresponding motion commands.

[0106] According to the present invention, a multi-UAV target coverage device 400 based on a visual language model includes a feature fusion module 402 that fuses the static top view, each of the first state information and each of the second state information to obtain target fusion features, including: Acquire the target environment features of the static top view, the first encoded features of each first state information, and the second encoded features of each second state information; For each of the first encoded features, cross-attention processing is performed on the target environment features and the first encoded features to obtain the first-order fusion features corresponding to the first encoded features; Self-attention processing is performed on the first feature group obtained by combining the first-order fusion features to obtain the second-order fusion features corresponding to each of the first coding features; For each of the second-order fusion features, the second-order fusion feature is cross-attention processed with each of the second coding features to obtain the third-order fusion feature corresponding to the second-order fusion feature; The target fusion feature is obtained based on each of the first-order fusion features, each of the second-order fusion features, and each of the third-order fusion features.

[0107] According to the present invention, a multi-UAV target coverage device 400 based on a visual language model includes a feature fusion module 402 for acquiring target environment features of the static top-view, comprising: Feature extraction is performed on the static top view to obtain the initial environmental features of the static top view, which include sub-features of multiple regions in the static top view; Determine the first proportion of pixels belonging to roads and the second proportion of pixels belonging to buildings in each region of the static top view; For each region, the sub-features of the region are updated based on a first proportion of pixels belonging to roads and a second proportion of pixels belonging to buildings within the region. Based on the latest sub-features of each region, the target environment features of the static top view are obtained.

[0108] According to a multi-UAV target coverage device 400 based on a visual language model provided by the present invention, the feature fusion module 402 updates the sub-features of each region based on a first proportion of pixels belonging to roads and a second proportion of pixels belonging to buildings within the region, including: It was determined that the second proportion within the region was greater than or equal to the preset proportion. The sub-features of the region are set to preset values, which are used to characterize that the region is impassable.

[0109] According to the present invention, a multi-UAV target coverage device 400 based on a visual language model includes a feature fusion module 402 that obtains target environment features of the static top view based on the latest sub-features of each region, including: For each region, the latest sub-features of the region, the first proportion, and the second proportion are concatenated to obtain the environmental features of the region. The environmental features of each region are combined according to the positional relationship between the regions to obtain the target environmental features.

[0110] According to the present invention, a multi-UAV target coverage device 400 based on a visual language model includes a feature fusion module 402 that obtains the target fusion features based on each of the first-order fusion features, each of the second-order fusion features, and each of the third-order fusion features, including: For each UAV, the first-order fusion feature, second-order fusion feature and third-order fusion feature corresponding to the UAV are spliced ​​together to obtain spliced ​​data. Each of the spliced ​​data is fused using a fully connected neural network to obtain the spatiotemporal feature data corresponding to each of the spliced ​​data. The spatiotemporal feature data are combined to obtain the target fusion feature.

[0111] According to the present invention, a multi-UAV target coverage device 400 based on a visual language model includes a motion command generation module 403 that processes the target fusion features and the task commands using a visual language model to obtain motion commands for each UAV, including: Obtain the current real-time top view of the task area; The real-time top-down view is encoded using the visual encoder in the visual language model to obtain visual features, and the task instructions are encoded using the text encoder in the visual language model to obtain text features. The text features, the visual features, and the target fusion features are fused together to obtain multimodal fusion features; Based on the multimodal fusion features, motion commands for each UAV are obtained.

[0112] According to the present invention, a multi-UAV target coverage device 400 based on a visual language model is provided, the multi-UAV target coverage device 400 based on a visual language model further includes a training module ( Figure 4 (Not shown in the image), the training module can be used to perform the training steps of the visual language model, the training steps including: Initialize the parameters of the visual language model and the course distribution parameters, wherein the course distribution parameters are used to indicate the sampling probability of the training dataset within different difficulty ranges; The preset training steps are executed repeatedly until the visual language model converges or reaches the maximum number of training steps. The preset training steps include: Based on the task distribution parameters required for this training, the initial dataset required for this training is obtained, and the task distribution parameters include the latest course distribution parameters; The trajectory data of a virtual drone moving according to training motion commands in a simulation environment is determined. The training motion commands are motion commands planned by the visual language model to be trained based on the initial dataset. The parameters of the visual language model are adjusted based on the trajectory data; The course distribution parameters are adjusted at preset intervals of a certain number of training cycles.

[0113] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a multi-UAV target coverage method based on a visual language model. This method includes: acquiring a static top-down view of the task area where the UAV group is located, first state information of each UAV in the UAV group at the current moment, second state information of several observed objects at the current moment, and task instructions issued by the user; fusing the static top-down view, each of the first state information, and each of the second state information to obtain target fusion features; processing the target fusion features and the task instructions using a visual language model to obtain motion instructions for each UAV; and controlling each UAV to move according to the corresponding motion instructions.

[0114] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0115] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-UAV target coverage method based on a visual language model provided by the above methods. The method includes: acquiring a static top view of the task area where the UAV group is located, first state information of each UAV in the UAV group at the current moment, second state information of several observed objects at the current moment, and task instructions issued by the user; fusing the static top view, each of the first state information and each of the second state information to obtain target fusion features; processing the target fusion features and the task instructions using a visual language model to obtain motion instructions for each UAV; and controlling each UAV to move according to the corresponding motion instructions.

[0116] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the multi-UAV target coverage method based on a visual language model provided by the above methods. The method includes: acquiring a static top view of the task area where the UAV group is located, first state information of each UAV in the UAV group at the current moment, second state information of several observed objects at the current moment, and task instructions issued by the user; fusing the static top view, each of the first state information and each of the second state information to obtain target fusion features; processing the target fusion features and the task instructions using a visual language model to obtain motion instructions for each UAV; and controlling each UAV to move according to the corresponding motion instructions.

[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0118] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

[0120] All actions involving the acquisition of images, status information, or data in this application are carried out in accordance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.

[0121] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the individual's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, when multiple drones need to collect images of the mission area, clear and prominent signs will be set at the entry and exit points of the target area to inform users that they have entered the personal information collection area and that personal information will be collected. If an individual voluntarily enters the collection area (mission area), it is considered that they have consented to the collection of their personal information.

Claims

1. A method for multi-UAV target coverage based on a visual language model, characterized in that, include: Obtain a static top-down view of the mission area where the UAV group is located, the first state information of each UAV in the UAV group at the current moment, the second state information of several observed objects at the current moment, and the mission instructions issued by the user; The static top view, each of the first state information and each of the second state information are fused to obtain the target fusion feature; The target fusion features and the task instructions are processed using a visual language model to obtain the motion instructions for each of the UAVs; Control each of the aforementioned drones to move according to the corresponding motion commands.

2. The method according to claim 1, characterized in that, The process of fusing the static top view, each of the first state information and each of the second state information to obtain the target fusion feature includes: Acquire the target environment features of the static top view, the first encoded features of each first state information, and the second encoded features of each second state information; For each of the first encoded features, cross-attention processing is performed on the target environment features and the first encoded features to obtain the first-order fusion features corresponding to the first encoded features; Self-attention processing is performed on the first feature group obtained by combining the first-order fusion features to obtain the second-order fusion features corresponding to each of the first coding features; For each of the second-order fusion features, the second-order fusion feature is cross-attention processed with each of the second coding features to obtain the third-order fusion feature corresponding to the second-order fusion feature; The target fusion feature is obtained based on each of the first-order fusion features, each of the second-order fusion features, and each of the third-order fusion features.

3. The method according to claim 2, characterized in that, The acquisition of the target environment features of the static top view includes: Feature extraction is performed on the static top view to obtain the initial environmental features of the static top view, which include sub-features of multiple regions in the static top view; Determine the first proportion of pixels belonging to roads and the second proportion of pixels belonging to buildings in each region of the static top view; For each region, the sub-features of the region are updated based on a first proportion of pixels belonging to roads and a second proportion of pixels belonging to buildings within the region. Based on the latest sub-features of each region, the target environment features of the static top view are obtained.

4. The method according to claim 3, characterized in that, For each region, the sub-features of the region are updated based on a first proportion of pixels belonging to roads and a second proportion of pixels belonging to buildings within the region, including: It was determined that the second proportion within the region was greater than or equal to the preset proportion. The sub-features of the region are set to preset values, which are used to characterize that the region is impassable.

5. The method according to claim 3, characterized in that, The process of obtaining the target environment features of the static top view based on the latest sub-features of each region includes: For each region, the latest sub-features of the region, the first proportion, and the second proportion are concatenated to obtain the environmental features of the region. The environmental features of each region are combined according to the positional relationship between the regions to obtain the target environmental features.

6. The method according to any one of claims 2 to 5, characterized in that, The process of obtaining the target fusion feature based on each of the first-order fusion features, each of the second-order fusion features, and each of the third-order fusion features includes: For each UAV, the first-order fusion feature, second-order fusion feature and third-order fusion feature corresponding to the UAV are spliced ​​together to obtain spliced ​​data. Each of the spliced ​​data is fused using a fully connected neural network to obtain the spatiotemporal feature data corresponding to each of the spliced ​​data. The spatiotemporal feature data are combined to obtain the target fusion feature.

7. The method according to any one of claims 1 to 5, characterized in that, The process of using a visual language model to process the target fusion features and the task commands to obtain the motion commands for each of the UAVs includes: Obtain the current real-time top view of the task area; The real-time top-down view is encoded using the visual encoder in the visual language model to obtain visual features, and the task instructions are encoded using the text encoder in the visual language model to obtain text features. The text features, the visual features, and the target fusion features are fused together to obtain multimodal fusion features; Based on the multimodal fusion features, motion commands for each UAV are obtained.

8. The method according to any one of claims 1 to 5, characterized in that, The method further includes a training step for the visual language model, the training step comprising: Initialize the parameters of the visual language model and the course distribution parameters, wherein the course distribution parameters are used to indicate the sampling probability of the training dataset within different difficulty ranges; The preset training steps are executed repeatedly until the visual language model converges or reaches the maximum number of training steps. The preset training steps include: Based on the task distribution parameters required for this training, the initial dataset required for this training is obtained, and the task distribution parameters include the latest course distribution parameters; The trajectory data of a virtual drone moving according to training motion commands in a simulation environment is determined. The training motion commands are motion commands planned by the visual language model to be trained based on the initial dataset. The parameters of the visual language model are adjusted based on the trajectory data; The course distribution parameters are adjusted at preset intervals of a certain number of training cycles.

9. A multi-UAV target coverage device based on a visual language model, characterized in that, include: The data acquisition module is used to acquire a static top view of the mission area where the UAV group is located, the first state information of each UAV in the UAV group at the current moment, the second state information of several observed objects at the current moment, and the mission instructions issued by the user. The feature fusion module is used to fuse the static top view, each of the first state information and each of the second state information to obtain the target fused features; The motion command generation module is used to process the target fusion features and the task commands using a visual language model to obtain motion commands for each of the UAVs; The drone control module is used to control each of the drones to move according to the corresponding motion commands.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multi-UAV target coverage method based on a visual language model as described in any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-UAV target coverage method based on a visual language model as described in any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-UAV target coverage method based on a visual language model as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • CN120469480A

  • CN120704350A

  • CN120722955A

  • CN121052331A

  • CN121742525A