Multi-unmanned aerial vehicle cooperative target sensing method and system, electronic equipment and storage medium
Through the collaborative target perception method of multiple drones, feature extraction and matching is used for image data of multiple drones and deep learning networks, solving the problem of low search efficiency of single drones and achieving a larger range and efficient target search.
Patent Information
- Application Number
- CN202411920314.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
Single drones have limitations in coverage, endurance and data processing capabilities, resulting in low overall search efficiency.
Through the multi-drone collaborative target perception method, multiple sets of image data captured by multiple drones are obtained, and image features are extracted and matched using twin feature extraction networks and top-level networks, and the Euro-style distance vectors and image matching probability are calculated to realize multi-drone collaborative target perception.
Make full use of space-time asynchronous information between multiple drones, improve the search range and efficiency of targets in cities, villages, and streets, and overcome the limitations of a single drone.
Smart Images

Figure CN120047850A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision perception, and more particularly, to a multi-UAV collaborative target perception method, system, electronic device, and storage medium. Background Art
[0002] The UAV target perception technology is a key technology required for UAVs to perform tasks such as reconnaissance. In related technologies, target information is obtained only by a single UAV. In this way, a single UAV has great limitations in terms of coverage, endurance, and data processing capabilities, and the overall search efficiency is relatively low. Summary of the Invention
[0003] To solve or improve at least one of the above technical problems, an object of the present invention is to provide a multi-UAV collaborative target perception method.
[0004] Another object of the present invention is to provide a multi-UAV collaborative target perception system.
[0005] Another object of the present invention is to provide an electronic device.
[0006] Another object of the present invention is to provide a readable storage medium.
[0007] To achieve the above object, a first aspect of the present invention provides a multi-UAV collaborative target perception method, the steps of which include:
[0008] First, obtain multiple sets of image data captured by multiple UAVs; wherein, the multiple sets of image data include first image data and second image data, the first image data comes from one of the multiple UAVs, and the second image data comes from another one of the multiple UAVs.
[0009] Second, based on a siamese feature extraction network, divide the first image data and the second image data into multiple image blocks and perform feature extraction to obtain multiple image features.
[0010] Third, calculate the Euclidean distance vector between the multiple image features.
[0011] Fourth, based on a top-level network, calculate the image matching probability between the multiple image blocks according to the Euclidean distance vector, and the image matching probability is used to represent the probability that two image blocks belong to the same target.
[0012] Fifth, in the case where the image matching probability is greater than a first threshold, classify at least two image blocks into the same group of matching image pairs.
[0013] Step 6: Based on multiple groups of matching image pairs, determine the association information between the targets detected by multiple UAVs to achieve multi-UAV collaborative target perception.
[0014] The present invention aims to provide a multi-UAV collaborative target perception method. By obtaining multiple groups of image data through multiple UAVs and performing data processing, multi-UAV collaborative target perception is achieved. Multi-UAV collaborative perception helps to make full use of the spatio-temporal asynchronous information between multiple UAVs to search for targets in a larger range in cities, rural areas, and streets, and can overcome the limitations of a single UAV in terms of coverage, endurance, data processing ability, etc., improving the overall search efficiency.
[0015] In addition, the above technical solution provided by the present invention may also have the following additional technical features:
[0016] In some technical solutions, optionally, the steps of the multi-UAV collaborative target perception method further include: before dividing the first image data and the second image data into multiple image blocks and performing feature extraction based on the Siamese feature extraction network to obtain multiple image features, performing data cleaning and data augmentation on the first image data and the second image data.
[0017] In this technical solution, data cleaning is used to remove noise and outliers in the first image data and / or the second image data. Data cleaning is also used to remove image defects caused by possible interference factors during the UAV shooting process, such as lens stains, data loss or errors during image transmission, etc.
[0018] By data cleaning, the accuracy of subsequent feature extraction and target perception can be improved, and the possibility of misjudgment caused by incorrect or low-quality data can be effectively reduced.
[0019] Data augmentation is used to expand the data set. Data augmentation is used to perform operations such as flipping, cropping, and brightness adjustment on the first image data and / or the second image data to generate more diverse image samples. This data processing method is beneficial to improving the accuracy of subsequent feature extraction and target perception.
[0020] In some technical solutions, optionally, the first image data includes a first target image and a first environmental image; and / or the second image data includes a second target image and a second environmental image.
[0021] In this technical solution, the first target image and the second target image are used to reflect the shape and color features of the target. The first environmental image and the second environmental image are used to reflect the environmental features of the area where the target is located.
[0022] There are differences between the first image data and the second image data in terms of image perspective, coverage range, and the target details presented. Multiple sets of image data captured by multiple drones from different perspectives and different positions provide data support for data processing in subsequent steps.
[0023] In some technical solutions, optionally, based on the siamese feature extraction network, the first image data and the second image data are divided into multiple image patches and feature extraction is performed to obtain multiple image features. The steps include: based on the siamese feature extraction network, dividing the images in the first image data into multiple first image patches and performing feature extraction to obtain multiple first image features; based on the siamese feature extraction network, dividing the images in the second image data into multiple second image patches and performing feature extraction to obtain multiple second image features.
[0024] In this technical solution, the siamese feature extraction network includes a convolutional layer, a pooling layer, and a fully connected layer. The siamese feature extraction network performs feature extraction on multiple image patches through the convolutional layer, the pooling layer, and the fully connected layer to obtain multiple image features.
[0025] The convolutional layer is used to capture local features in the image patch, such as edge, texture, and other information. The pooling layer is used to downsample the local features. The fully connected layer is used to integrate the local features extracted by the convolutional layer and the pooling layer, and determine the image features based on the local features, and finally output the deep feature description vector of the image patch.
[0026] By dividing the first image data and the second image data into multiple image patches, feature extraction can be performed for each individual target or key parts of the target (such as the external contour of a vehicle, the facial features of a person, etc.), which is beneficial to improving the accuracy and effectiveness of the image features.
[0027] In some technical solutions, optionally, based on the siamese feature extraction network, dividing the images in the first image data into multiple first image patches and performing feature extraction to obtain multiple first image features, the steps include: classifying multiple images in the first image data into the first image group to be matched; inputting the first image group to be matched into the first sub-network of the siamese feature extraction network; based on the first sub-network of the siamese feature extraction network, dividing the images in the first image group to be matched into multiple first image patches and performing feature extraction to obtain multiple first image features.
[0028] In this technical solution, by grouping multiple images in the first image data, it helps to uniformly manage and process these image data, so that the first image group to be matched serves as the input data of the first sub-network of the siamese feature extraction network.
[0029] Processing the image data from a specific source through a dedicated sub-network can achieve the shunt processing of image data from different sources (from different drones), which is beneficial to improving the processing efficiency and the logic of the system.
[0030] By dividing the images in the first image data into multiple first image blocks and performing feature extraction, feature extraction can be carried out for each individual target or the key parts of the target, which is beneficial to improving the accuracy and effectiveness of the first image features.
[0031] In some technical solutions, optionally, based on the siamese feature extraction network, divide the images in the second image data into multiple second image blocks and perform feature extraction to obtain multiple second image features. The steps include: classifying multiple images in the second image data into a second image group to be matched; inputting the second image group to be matched into the second sub-network of the siamese feature extraction network; based on the second sub-network of the siamese feature extraction network, divide the images in the second image group to be matched into multiple second image blocks and perform feature extraction to obtain multiple second image features.
[0032] In this technical solution, by grouping multiple images in the second image data together, it helps to uniformly manage and process these image data, so that the second image group to be matched can be used as the input data of the second sub-network of the siamese feature extraction network.
[0033] Processing the image data from a specific source through a dedicated sub-network can achieve the shunt processing of image data from different sources (from different drones), which is beneficial to improving the processing efficiency and the logic of the system.
[0034] By dividing the images in the second image data into multiple second image blocks and performing feature extraction, feature extraction can be carried out for each individual target or the key parts of the target, which is beneficial to improving the accuracy and effectiveness of the second image features.
[0035] In some technical solutions, optionally, calculate the Euclidean distance vector between multiple image features. The steps include: calculating the Euclidean distance between each first image feature and each second image feature, and determining the Euclidean distance vector based on multiple Euclidean distances.
[0036] In this technical solution, the Euclidean distance is used to compare the similarity between two image features. The smaller the Euclidean distance, the more similar the two image features are, and the more likely the corresponding image blocks belong to the same target; conversely, the larger the perspective distance, the greater the difference between the two image features, and the lower the possibility of belonging to the same target.
[0037] The Euclidean distance vector is a set of vectors containing multiple Euclidean distance values. In the actual application scenario of multi-UAV collaborative target perception, when multiple image features need to be compared with each other, by calculating the Euclidean distance between every two image features, combining multiple Euclidean distance values together forms the Euclidean distance vector.
[0038] The Euclidean distance vector is used to reflect the distance relationship between all image features in the feature space, providing a comprehensive data basis for subsequent analysis and judgment.
[0039] The second aspect of the present invention provides a multi-UAV collaborative target perception system, including an image data acquisition module, a feature extraction module, a first calculation module, a second calculation module, an image matching module, and an information determination module.
[0040] The image data acquisition module is used to acquire multiple groups of image data captured by multiple UAVs. Among them, the multiple groups of image data include first image data and second image data. The first image data comes from one of the multiple UAVs, and the second image data comes from another one of the multiple UAVs.
[0041] The feature extraction module is used to divide the first image data and the second image data into multiple image blocks and perform feature extraction based on the Siamese feature extraction network to obtain multiple image features.
[0042] The first calculation module is used to calculate the Euclidean distance vector between multiple image features.
[0043] The second calculation module is used to calculate the image matching probability between multiple image blocks based on the top-level network according to the Euclidean distance vector. The image matching probability is used to represent the probability that two image blocks belong to the same target.
[0044] The image matching module is used to classify at least two image blocks into the same group of matching image pairs when the image matching probability is greater than the first threshold.
[0045] The information determination module is used to determine the association information between the targets detected by multiple UAVs based on multiple groups of matching image pairs to achieve multi-UAV collaborative target perception.
[0046] The present invention aims to provide a multi-UAV collaborative target perception system, which acquires multiple groups of image data through multiple UAVs and performs data processing to achieve multi-UAV collaborative target perception. Multi-UAV collaborative perception helps to make full use of the spatio-temporal asynchronous information between multiple UAVs to search for targets in a larger range in cities, villages, and streets, and can overcome the limitations of a single UAV in terms of coverage, endurance, data processing ability, etc., improving the overall search efficiency.
[0047] The third aspect of the present invention provides an electronic device, including a memory and a processor. Among them, a program or instruction that can run on the processor is stored on the memory. When the processor executes the program or instruction, the steps of the multi-UAV collaborative target perception method in any of the above technical solutions are implemented. Therefore, the electronic device has the beneficial effects of any of the above technical solutions, which will not be elaborated here.
[0048] The fourth aspect of the present invention provides a readable storage medium. The readable storage medium stores a program or instruction. When the program or instruction is executed by a processor, the steps of the multi-UAV collaborative target perception method in any of the above technical solutions are implemented. Therefore, the readable storage medium has the beneficial effects of any of the above technical solutions, which will not be elaborated here.
[0049] The additional aspects and advantages of the technical solutions of the present invention will become apparent in the following description section or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 The flowchart of the multi-UAV collaborative target perception method according to an embodiment of the present invention is shown;
[0051] Figure 2 The flowchart of the multi-UAV collaborative target perception method according to another embodiment of the present invention is shown;
[0052] Figure 3 The flowchart of the multi-UAV collaborative target perception method according to another embodiment of the present invention is shown;
[0053] Figure 4 The flowchart of the multi-UAV collaborative target perception method according to another embodiment of the present invention is shown;
[0054] Figure 5 The flowchart of the multi-UAV collaborative target perception method according to another embodiment of the present invention is shown;
[0055] Figure 6 The flowchart of the multi-UAV collaborative target perception method according to another embodiment of the present invention is shown;
[0056] Figure 7 The structural block diagram of the multi-UAV collaborative target perception system according to an embodiment of the present invention is shown;
[0057] Figure 8 The structural block diagram of the electronic device according to an embodiment of the present invention is shown;
[0058] Figure 9 The schematic diagram of data processing based on the twin feature extraction network according to an embodiment of the present invention is shown.
[0059] Among them, Figure 7 and Figure 8 the corresponding relationship between the reference numerals and the component names in the drawings is as follows:
[0060] 200: Multi-UAV collaborative target perception system; 210: Image data acquisition module; 220: Feature extraction module; 230: First calculation module; 240: Second calculation module; 250: Image matching module; 260: Information determination module; 300: Electronic device; 310: Memory; 320: Processor. Detailed implementation manners
[0061] In order to more clearly understand the above objects, features, and advantages of the embodiments of the present invention, the embodiments of the present invention will be further described in detail below with reference to the drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0062] In the following description, many specific details are set forth in order to fully understand the present invention. However, the embodiments of the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the limitations of the specific embodiments disclosed below.
[0063] Next, refer to Figures 1 to 9 to describe a multi-UAV collaborative target perception method, system, electronic device, and storage medium provided according to some embodiments of the present invention.
[0064] In an embodiment according to the present invention, as Figure 1 shown, the steps of the multi-UAV collaborative target perception method include:
[0065] S102, obtaining multiple sets of image data captured by multiple UAVs; among them, the multiple sets of image data include first image data and second image data, the first image data comes from one of the multiple UAVs, and the second image data comes from another one of the multiple UAVs.
[0066] Obtain multiple sets of image data through the camera ends of multiple UAVs. Optionally, the image data is visible light images. Multiple UAVs respectively take pictures of the ground scene from different perspectives and different positions to obtain multiple sets of image data, and send the multiple sets of image data to the ground terminal.
[0067] Multiple UAVs perform image acquisition on the ground scene (target area) according to a preset flight path and shooting task to obtain multiple sets of image data. Each UAV is equipped with an advanced image acquisition device (such as a camera end), which can obtain high-resolution and high-definition visible light images.
[0068] The first image data is a video or image sequence captured by the first unmanned aerial vehicle (one of the unmanned aerial vehicle clusters). The second image data is a video or image sequence captured by the second unmanned aerial vehicle (another unmanned aerial vehicle in the unmanned aerial vehicle clusters).
[0069] There are differences between the first image data and the second image data in terms of image perspective, coverage, and the target details presented.
[0070] Multiple sets of image data captured by multiple unmanned aerial vehicles from different perspectives and positions provide data support for data processing in subsequent steps. The subsequent steps sort out and integrate the multiple sets of image data to achieve accurate identification, positioning, and comprehensive perception of the target.
[0071] In a specific embodiment, the multiple sets of image data include multiple target images and multiple environment images. The target images are used to reflect the shape and color characteristics of the target. The environment images are used to reflect the environmental characteristics of the area where the target is located.
[0072] S104, based on the siamese feature extraction network, divide the first image data and the second image data into multiple image patches and perform feature extraction to obtain multiple image features.
[0073] In an actual scenario, a complete image (target image or environment image) may contain multiple targets or multiple parts of a target. By dividing the image into patches, the features of each local area can be captured more carefully. By performing feature extraction on multiple image patches respectively, multiple image features can be obtained, and these image features will serve as the core basis for subsequent judgment of the correlation degree between image patches and determination of target identity.
[0074] It should be emphasized that by dividing the first image data and the second image data into multiple image patches, feature extraction can be performed for each individual target or key parts of the target (such as the external contour of a vehicle, the facial features of a person, etc.), which is beneficial to improving the accuracy and effectiveness of the image features.
[0075] After obtaining the first image data and the second image data from multiple unmanned aerial vehicles, in-depth processing is carried out based on the pre-constructed and trained siamese feature extraction network.
[0076] The siamese feature extraction network is a deep learning network architecture mainly used to extract features of images. Optionally, the siamese feature extraction network adopts the basic network structure of Resnet-18, and through the combination of convolutional layers, pooling layers, and fully connected layers, it can effectively mine the deep feature information in the images.
[0077] It should be noted that Resnet-18 is a deep residual network architecture and belongs to a type of convolutional neural network.
[0078] Optionally, the siamese feature extraction network includes a convolutional layer, a pooling layer, and a fully-connected layer. The siamese feature extraction network extracts features from the image patches through the convolutional layer, the pooling layer, and the fully-connected layer.
[0079] The convolutional layer extracts features from the image patches through a series of convolutional kernels. These convolutional kernels slide on the image patches according to the set stride for convolutional operations, thereby extracting local features. The convolutional layer is used to capture local features in the image patches, such as edge, texture, and other information.
[0080] The pooling layer is used to downsample the local features. This method can reduce the data dimension while retaining key feature information to reduce the subsequent calculation amount.
[0081] The fully-connected layer is used to integrate the local features extracted by the convolutional layer and the pooling layer, and determine the image features based on the local features, and finally output the deep feature description vector of the image patch.
[0082] Optionally, the siamese feature extraction network is composed of two sub-networks with the same structure, and these two sub-networks share parameters. The two sub-networks respectively process the first image data and the second image data.
[0083] S106, calculate the Euclidean distance vector between multiple image features.
[0084] Determine the image feature vector according to the image features. Specifically, determine the first image feature vector according to the first image feature; determine the second image feature vector according to the second image feature.
[0085] Based on the first image feature vector and the second image feature vector, calculate the Euclidean distance. And determine the Euclidean distance vector according to multiple Euclidean distances.
[0086] It should be noted that the Euclidean distance is a measurement method used to measure the straight-line distance between two points in Euclidean space. In the context of multi-UAV collaborative target perception, it is used to measure the distance between two image feature vectors.
[0087] In a specific embodiment, the first image feature vector is A = (a 1 , a 2 , …, an ), the second image feature vector is B = (b 1 , b 2 , …, b n ).
[0088] The calculation formula of the Euclidean distance is
[0089] The Euclidean distance is used to compare the similarity between two image features. The smaller the Euclidean distance, the more similar the two image features are, and the more likely the corresponding image patches belong to the same target; conversely, the larger the perspective distance, the greater the difference between the two image features, and the lower the possibility of belonging to the same target.
[0090] For example, when the detection target of vehicle is included in both the first image patch and the second image patch, the smaller the Euclidean distance, the closer the image features of the vehicle in the first image patch and the vehicle in the second image patch are in terms of shape, color or texture, etc., and the vehicles in the two image patches are very likely to be the same target.
[0091] The Euclidean distance vector is a vector set containing multiple Euclidean distance values. In the actual application scenario of multi-UAV collaborative target perception, when multiple image features need to be compared with each other, by calculating the Euclidean distance between every two image features, a set of multiple Euclidean distance values is combined to form the Euclidean distance vector.
[0092] The Euclidean distance vector is used to reflect the distance relationship between all image features in the feature space, providing a comprehensive data basis for subsequent analysis and judgment.
[0093] S108. Based on the top-level network, calculate the image matching probability between multiple image patches according to the Euclidean distance vector, and the image matching probability is used to represent the probability that two image patches belong to the same target.
[0094] The Euclidean distance vector is used as the input data of the top-level network, and the top-level network calculates the image matching probability according to the Euclidean distance vector to determine the probability that every two image patches belong to the same target.
[0095] The top-level network mainly plays the role of information conversion or data conversion. The top-level network can convert the relatively abstract feature space distance measurement information (Euclidean distance vector) into relatively intuitive quantitative values (image matching probability). The Euclidean distance vector can only reflect the degree of difference of image features in space, but cannot directly indicate whether the image patches belong to the same target, while the image matching probability intuitively gives the quantitative value of this possibility.
[0096] S110. In the case where the image matching probability is greater than the first threshold, classify at least two image patches into the same group of matching image pairs.
[0097] When the image matching probability is greater than a certain set threshold (the first threshold), it is determined that the corresponding image patches belong to the same target, and then these image patches are grouped into the same set of matching image pairs.
[0098] Considering the accuracy of target perception, data processing efficiency, and other factors, the first threshold is flexibly set according to actual requirements.
[0099] Optionally, the first threshold is from 0.4 to 0.6.
[0100] In a specific embodiment, the first threshold is 0.5. When the image matching probability is greater than 0.5, at least two image patches are grouped into the same set of matching image pairs.
[0101] S112. Based on multiple sets of matching image pairs, determine the association information between the targets detected by multiple UAVs to achieve multi-UAV collaborative target perception.
[0102] Due to the differences in spatial positions, shooting angles, and time points of different UAVs, the acquired image data has diversity and complementarity. By integrating the sets of matching image pairs, a complete portrait of the target from the perspectives of multiple UAVs can be constructed, clarifying the presentation forms of each target in the fields of view of different UAVs and their mutual relationships.
[0103] Optionally, based on multiple sets of matching image pairs, determine the position distribution and predicted motion trajectories of the targets detected by multiple UAVs, as well as the association information between multiple targets.
[0104] This data processing method not only helps to eliminate the limitations of single-UAV perception but also enables more accurate and comprehensive perception of the entire target group, thus achieving the ultimate goal of multi-UAV collaborative target perception and providing a solid information basis for subsequent task execution (such as pursuit, search and rescue, etc.).
[0105] Optionally, each target is defined by an identifier (ID, Identifier) and the target is tracked.
[0106] The present invention aims to provide a multi-UAV collaborative target perception method. By acquiring multiple sets of image data through multiple UAVs and performing data processing, multi-UAV collaborative target perception is achieved. Multi-UAV collaborative perception helps to make full use of the spatio-temporal asynchronous information between multiple UAVs to search for targets in a larger range in cities, rural areas, and streets, and can overcome the limitations of single UAVs in terms of coverage, endurance, data processing ability, etc., improving the overall search efficiency.
[0107] It should be emphasized that the multi-UAV collaborative target perception method can achieve multi-UAV collaborative target perception in scenarios such as urban streets, rural areas, and mountainous areas, complete tasks such as pursuit and search and rescue, overcome the limitations of a single UAV in terms of coverage, endurance, data processing ability, etc., and improve the overall search efficiency.
[0108] In some embodiments, optionally, as Figure 2 shown, before S104 (dividing the first image data and the second image data into multiple image blocks and performing feature extraction based on the twin feature extraction network to obtain multiple image features), the steps of the multi-UAV collaborative target perception method further include:
[0109] S103, performing data cleaning and data augmentation on the first image data and the second image data.
[0110] The purpose of this step is to preprocess the image data, and the preprocessing methods include but are not limited to data cleaning and data augmentation.
[0111] Data cleaning is used to remove noise and outliers in the first image data and / or the second image data. Data cleaning is also used to remove image defects caused by possible interference factors during the UAV shooting process, such as lens stains, data loss or errors during image transmission, etc.
[0112] By data cleaning, the accuracy of subsequent feature extraction and target perception can be improved, and the possibility of misjudgment caused by incorrect or low-quality data can be effectively reduced.
[0113] Data augmentation is used to expand the data set. Data augmentation is used to perform operations such as flipping, cropping, and brightness adjustment on the first image data and / or the second image data to generate more diverse image samples. This data processing method is beneficial to improving the accuracy of subsequent feature extraction and target perception.
[0114] In some embodiments, optionally, the first image data includes a first target image and a first environment image.
[0115] The first image data is a video or image sequence obtained by shooting from the first UAV (one of the UAVs in the multi-UAV cluster).
[0116] The first target image is used to reflect the shape and color features of the target. The first environment image is used to reflect the environmental features of the area where the target is located.
[0117] In some embodiments, optionally, the second image data includes a second target image and a second environment image.
[0118] The second image data is a video or image sequence captured by a second unmanned aerial vehicle (another unmanned aerial vehicle in a cluster of multiple unmanned aerial vehicles).
[0119] The second target image is used to reflect the shape and color features of the target. The second environmental image is used to reflect the environmental features of the area where the target is located.
[0120] There are differences between the first image data and the second image data in terms of image perspective, coverage range, and the details of the presented target, etc.
[0121] Multiple sets of image data captured by multiple unmanned aerial vehicles from different perspectives and different positions provide data support for data processing in subsequent steps. The subsequent steps sort out and integrate the multiple sets of image data to achieve precise identification, positioning, and comprehensive perception of the target.
[0122] In some embodiments, optionally, as Figure 3 shown, the steps of S104 (dividing the first image data and the second image data into multiple image blocks and performing feature extraction based on a siamese feature extraction network to obtain multiple image features) include:
[0123] S1042, based on the siamese feature extraction network, divide the images in the first image data into multiple first image blocks and perform feature extraction to obtain multiple first image features.
[0124] The siamese feature extraction network includes a convolutional layer, a pooling layer, and a fully connected layer. The siamese feature extraction network performs feature extraction on multiple first image blocks through the convolutional layer, the pooling layer, and the fully connected layer to obtain multiple first image features.
[0125] The convolutional layer is used to capture the first local features in the first image block, such as edge, texture, and other information. The pooling layer is used to downsample the first local features. The fully connected layer is used to integrate the first local features extracted by the convolutional layer and the pooling layer, and determine the first image features based on the first local features, and finally output the deep feature description vector of the first image block.
[0126] By dividing the images in the first image data into multiple first image blocks and performing feature extraction, feature extraction can be performed for each individual target or key part of the target (such as the external contour of a vehicle, the facial features of a person, etc.), which is beneficial to improving the accuracy and effectiveness of the first image features.
[0127] S1044, based on the siamese feature extraction network, divide the images in the second image data into multiple second image blocks and perform feature extraction to obtain multiple second image features.
[0128] The twin feature extraction network extracts features from multiple second image patches through convolutional layers, pooling layers, and fully connected layers, obtaining multiple second image features.
[0129] The convolutional layer is used to capture second local features in the second image patch, such as information on edges, textures, etc. The pooling layer is used to downsample the second local features. The fully connected layer is used to integrate the second local features extracted by the convolutional layer and the pooling layer, and determine the second image features based on the second local features, finally outputting the deep feature description vector of the second image patch.
[0130] By dividing the images in the second image data into multiple second image patches and performing feature extraction, feature extraction can be carried out for each individual target or key parts of the target (such as the outer contour of a vehicle, the facial features of a person, etc.), which is beneficial to improving the accuracy and effectiveness of the second image features.
[0131] In some embodiments, optionally, as Figure 4 shown, the steps of S1042 (based on the twin feature extraction network, dividing the images in the first image data into multiple first image patches and performing feature extraction, obtaining multiple first image features) include:
[0132] S1045, classifying multiple images in the first image data into the first image group to be matched.
[0133] The purpose of this step is to construct the first image group to be matched based on the first image data. By grouping multiple images in the first image data together, it helps to uniformly manage and process these image data, so that the first image group to be matched serves as the input data of the first sub-network of the twin feature extraction network.
[0134] S1046, inputting the first image group to be matched into the first sub-network of the twin feature extraction network.
[0135] This data processing method, by using a dedicated sub-network to process image data from a specific source, can achieve the shunt processing of image data from different sources (from different drones), which is beneficial to improving the processing efficiency and the logic of the system.
[0136] S1047, based on the first sub-network of the twin feature extraction network, dividing the images in the first image group to be matched into multiple first image patches and performing feature extraction, obtaining multiple first image features.
[0137] By dividing the images in the first image data into multiple first image patches and performing feature extraction, feature extraction can be carried out for each individual target or key parts of the target, which is beneficial to improving the accuracy and effectiveness of the first image features.
[0138] In some embodiments, optionally, asFigure 5 As shown, the steps of S1044 (dividing the images in the second image data into multiple second image blocks and performing feature extraction based on the twin feature extraction network to obtain multiple second image features) include:
[0139] S1048, classifying multiple images in the second image data into the second image group to be matched.
[0140] The purpose of this step is to construct the second image group to be matched based on the second image data. By classifying multiple images in the second image data into a group, it helps to uniformly manage and process these image data so that the second image group to be matched can be used as the input data of the second sub-network of the twin feature extraction network.
[0141] S1049, inputting the second image group to be matched into the second sub-network of the twin feature extraction network.
[0142] This data processing method can achieve the shunt processing of image data from different sources (from different drones) by using a dedicated sub-network to process image data from specific sources, which is beneficial to improving the processing efficiency and the logic of the system.
[0143] S1050, based on the second sub-network of the twin feature extraction network, dividing the images in the second image group to be matched into multiple second image blocks and performing feature extraction to obtain multiple second image features.
[0144] By dividing the images in the second image data into multiple second image blocks and performing feature extraction, it is possible to perform feature extraction for each individual target or key part of the target, which is beneficial to improving the accuracy and effectiveness of the second image features.
[0145] In some embodiments, optionally, as Figure 6 shown, the steps of S106 (calculating the Euclidean distance vector between multiple image features) include:
[0146] S1062, calculating the Euclidean distance between each first image feature and each second image feature, and determining the Euclidean distance vector based on multiple Euclidean distances.
[0147] The Euclidean distance is used to compare the similarity between two image features. The smaller the Euclidean distance, the more similar the two image features are, and the more likely the corresponding image blocks belong to the same target; conversely, the larger the perspective distance, the greater the difference between the two image features, and the lower the possibility of belonging to the same target.
[0148] For example, when the detection target of "vehicle" is included in both the first image block and the second image block, the smaller the Euclidean distance, the closer the image features of the vehicle in the first image block and the vehicle in the second image block are in terms of shape, color, texture, etc., and the vehicles in the two image blocks are very likely to be the same target.
[0149] The Euclidean distance vector is a vector set containing multiple Euclidean distance values. In the actual application scenario of multi-UAV collaborative target perception, when multiple image features need to be compared with each other, by calculating the Euclidean distance between every two image features, a set of multiple Euclidean distance values is combined to form the Euclidean distance vector.
[0150] The Euclidean distance vector is used to reflect the distance relationship between all image features in the feature space, providing a comprehensive data basis for subsequent analysis and judgment.
[0151] In some embodiments, optionally, before S108 (based on the top-level network, calculating the image matching probability between multiple image blocks according to the Euclidean distance vector, where the image matching probability is used to represent the probability that two image blocks belong to the same target), the steps of the multi-UAV collaborative target perception method further include:
[0152] S107, training the top-level network to determine the mapping relationship between the Euclidean distance vector and the image matching probability.
[0153] Based on the top-level network, according to the Euclidean distance vector and the mapping relationship, determining the image matching probability between multiple image blocks.
[0154] In an embodiment according to the present invention, the steps of the multi-UAV collaborative target perception method include:
[0155] S1, obtaining visible light images (image data) from the camera ends of multiple UAVs and sending them back to the ground end (ground terminal).
[0156] S2, based on the Siamese feature extraction network, extracting image features to obtain the deep feature description vectors of the image blocks.
[0157] S3, combining the feature vectors and inputting them into the top-level network to compare the image descriptors, obtaining the image matching result, and outputting the probability that two image blocks belong to the same target (image matching probability).
[0158] S4, integrating the probabilities (image matching probabilities) of each group of image blocks to obtain the association information between the targets detected by multiple UAVs, and realizing multi-UAV collaborative target perception.
[0159] The drone target perception technology is a key technology required for drones to perform tasks such as reconnaissance. Cooperative perception of multiple drones helps to make full use of the spatio-temporal asynchronous information between multiple drones and multiple sensors, search for targets in a larger range in cities, villages, and streets, and can overcome the limitations of a single drone or single sensor in terms of coverage, endurance, and data processing capabilities, thereby improving the overall search efficiency.
[0160] In an embodiment according to the present invention, the steps of the multi-drone cooperative target perception method include:
[0161] S201, Two drones respectively take pictures of the ground scene and send multiple images to the ground terminal.
[0162] S202, Multiple images obtained by the two drones are respectively input into two branches (sub-networks) of a Siamese detection network (Siamese feature extraction network) for feature extraction.
[0163] Optionally, the Siamese feature extraction network adopts the basic network structure of Resnet-18.
[0164] S203, Calculate the Euclidean distance between the image features and input it into the top-level network to calculate the image matching probability.
[0165] It should be noted that for the Siamese feature extraction network, the input data is the image group to be matched; the output data is the group of matched image pairs.
[0166] Among them, the image group to be matched is I = {I 1 , I 2 …I n}, I' = {I 1 ', I 2 '…I n '}. The group of matched image pairs is {M} = { {I 1q , I 2l}, {I 1k , I 2b}…}.
[0167] Figure 9 The following is a schematic diagram of data processing based on the Siamese feature extraction network. The following is a code example (simplified version) of the data processing algorithm:
[0168] Input: Image group to be matched I = {I 1 , I 2 …I n}, I' = {I 1 ', I 2 '…I n '}
[0169] Output: The set of matching image pairs {M} = { {I 1q , I 2l}, {I 1k , I 2b}...}
[0170]
[0171] Return {M}.
[0172] S204. Output the association information between each target image, assign an id to each individual target, and obtain the perception results of all targets on the field.
[0173] In a specific embodiment, the drone is a quadrotor small drone. The parameters of the mounted optoelectronic pod are as follows: the resolution of the visible light module is not less than 1920×1080. The on-board computing unit of the drone uses NVIDIA Jetson Xavier NX (a control system module) and is powered by 15w.
[0174] For the ground terminal, the GPU (Graphics Processing Unit) of its computing unit uses GTX2080Ti (a model), and the CPU (Central Processing Unit) uses i5-10400 (a model).
[0175] In an embodiment according to the present invention, as Figure 7 shown, the multi-drone collaborative target perception system 200 includes an image data acquisition module 210, a feature extraction module 220, a first computing module 230, a second computing module 240, an image matching module 250, and an information determination module 260.
[0176] The image data acquisition module 210 is used to acquire multiple sets of image data captured by multiple drones. Among them, the multiple sets of image data include first image data and second image data. The first image data comes from one of the multiple drones, and the second image data comes from another one of the multiple drones.
[0177] Multiple sets of image data are acquired through the camera ends of multiple drones. Optionally, the image data is visible light images. Multiple drones take pictures of the ground scene from different perspectives and different positions respectively, acquire multiple sets of image data, and send the multiple sets of image data to the ground terminal.
[0178] Multiple drones collect images of the ground scene (target area) according to the preset flight paths and shooting tasks, obtaining multiple sets of image data. Each drone is equipped with advanced image acquisition equipment (such as a camera end) that can acquire high-resolution and high-clarity visible light images.
[0179] The first image data is a video or image sequence captured by the first drone (one of the multiple drones in the drone cluster). The second image data is a video or image sequence captured by the second drone (another drone in the multiple drones in the drone cluster).
[0180] There are differences between the first image data and the second image data in terms of image perspective, coverage range, and the target details presented.
[0181] Multiple sets of image data captured by multiple drones from different perspectives and positions provide data support for data processing in subsequent steps. The subsequent steps sort out and integrate the multiple sets of image data to achieve accurate identification, positioning, and comprehensive perception of the target.
[0182] In a specific embodiment, the multiple sets of image data include multiple target images and multiple environmental images. The target images are used to reflect the shape and color characteristics of the target. The environmental images are used to reflect the environmental characteristics of the area where the target is located.
[0183] The feature extraction module 220 is used to divide the first image data and the second image data into multiple image blocks and perform feature extraction based on the twin feature extraction network, obtaining multiple image features.
[0184] In an actual scenario, a complete image (target image or environmental image) may contain multiple targets or multiple parts of a target. By dividing the image blocks, the features of each local area can be captured more carefully. By performing feature extraction on multiple image blocks respectively, multiple image features can be obtained, and these image features will serve as the core basis for subsequent judgment of the correlation degree between image blocks and determination of target identity.
[0185] It should be emphasized that by dividing the first image data and the second image data into multiple image blocks, feature extraction can be performed for each individual target or key parts of the target (such as the external contour of a vehicle, the facial features of a person, etc.), which is beneficial to improving the accuracy and effectiveness of the image features.
[0186] After obtaining the first image data and the second image data from multiple drones, in-depth processing is carried out based on the pre-constructed and trained twin feature extraction network.
[0187] The Siamese feature extraction network is a deep learning network architecture mainly used for extracting features of images. Optionally, the Siamese feature extraction network adopts the basic network structure of Resnet-18. Through the combination of convolutional layers, pooling layers, and fully-connected layers, it can effectively mine the deep feature information in images.
[0188] It should be noted that Resnet-18 is a deep residual network architecture and belongs to a type of convolutional neural network.
[0189] Optionally, the Siamese feature extraction network includes convolutional layers, pooling layers, and fully-connected layers. The Siamese feature extraction network extracts features from image patches through convolutional layers, pooling layers, and fully-connected layers.
[0190] The convolutional layer extracts features from the image patch through a series of convolutional kernels. These convolutional kernels slide on the image patch according to the set stride for convolutional operations, thereby extracting local features. The convolutional layer is used to capture local features in the image patch, such as edge, texture, and other information.
[0191] The pooling layer is used to downsample the local features. This method can reduce the data dimension while retaining the key feature information to reduce the subsequent calculation amount.
[0192] The fully-connected layer is used to integrate the local features extracted by the convolutional layer and the pooling layer, and determine the image features based on the local features. Finally, it outputs the deep feature description vector of the image patch.
[0193] Optionally, the Siamese feature extraction network consists of two sub-networks with the same structure, and these two sub-networks share parameters. The two sub-networks respectively process the first image data and the second image data.
[0194] The first calculation module 230 is used to calculate the Euclidean distance vector between multiple image features.
[0195] Determine the image feature vector according to the image features. Specifically, determine the first image feature vector according to the first image feature; determine the second image feature vector according to the second image feature.
[0196] Based on the first image feature vector and the second image feature vector, calculate the Euclidean distance. And determine the Euclidean distance vector according to multiple Euclidean distances.
[0197] It should be noted that the Euclidean distance is a metric method used to measure the straight-line distance between two points in Euclidean space. In the context of multi-UAV collaborative target perception, it is used to measure the distance between two image feature vectors.
[0198] In a specific embodiment, the first image feature vector is A = (a 1 , a 2 , …, a n ), and the second image feature vector is B = (b 1 , b 2 , …, b n ).
[0199] The calculation formula for the Euclidean distance is
[0200] The Euclidean distance is used to compare the similarity between two image features. The smaller the Euclidean distance, the more similar the two image features are, and the more likely the corresponding image patches belong to the same target; conversely, the larger the perspective distance, the greater the difference between the two image features, and the lower the probability of belonging to the same target.
[0201] For example, when both the first image patch and the second image patch contain the detection target of a vehicle, the smaller the Euclidean distance, the closer the image features of the vehicle in the first image patch and the vehicle in the second image patch are in terms of shape, color, or texture, and the vehicles in the two image patches are very likely to be the same target.
[0202] The Euclidean distance vector is a vector set containing multiple Euclidean distance values. In the actual application scenario of multi-UAV collaborative target perception, when multiple image features need to be compared with each other, by calculating the Euclidean distance between every two image features, a set of multiple Euclidean distance values is combined to form the Euclidean distance vector.
[0203] The Euclidean distance vector is used to reflect the distance relationship between all image features in the feature space, providing a comprehensive data basis for subsequent analysis and judgment.
[0204] The second calculation module 240 is used to calculate the image matching probability between multiple image patches based on the top-level network. The image matching probability is used to represent the probability that two image patches belong to the same target.
[0205] The Euclidean distance vector serves as the input data for the top-level network, and the top-level network calculates the image matching probability based on the Euclidean distance vector to determine the probability that every two image patches belong to the same target.
[0206] The top-level network mainly plays the role of information conversion or data conversion. The top-level network can convert relatively abstract feature space distance metric information (Euclidean distance vector) into relatively intuitive quantization values (image matching probability). The Euclidean distance vector can only reflect the degree of difference in image features in space, but cannot directly indicate whether an image patch belongs to the same target, while the image matching probability intuitively gives a quantization value of this possibility.
[0207] The image matching module 250 is used to group at least two image patches into the same group of matching image pairs when the image matching probability is greater than the first threshold.
[0208] When the image matching probability is greater than a certain set threshold (the first threshold), it is determined that the corresponding image patches belong to the same target, and then these image patches are grouped into the same group of matching image pairs.
[0209] Considering the accuracy of target perception, data processing efficiency, and other factors, the first threshold is flexibly set according to actual needs.
[0210] Optionally, the first threshold is 0.4 to 0.6.
[0211] In a specific embodiment, the first threshold is 0.5. When the image matching probability is greater than 0.5, at least two image patches are grouped into the same group of matching image pairs.
[0212] The information determination module 260 is used to determine the association information between the targets detected by multiple UAVs based on multiple groups of matching image pairs to achieve multi-UAV collaborative target perception.
[0213] Due to the differences in spatial position, shooting angle, and time point of different UAVs, the acquired image data has diversity and complementarity. By integrating the group of matching image pairs, a complete portrait of the target from the perspectives of multiple UAVs can be constructed, clarifying the presentation forms of each target in the fields of view of different UAVs and their mutual relationships.
[0214] Optionally, based on multiple groups of matching image pairs, determine the position distribution and predicted motion trajectories of the targets detected by multiple UAVs, as well as the association information between multiple targets.
[0215] This data processing method not only helps to eliminate the limitations of single-UAV perception but also enables more accurate and comprehensive perception of the entire target group, thus achieving the ultimate goal of multi-UAV collaborative target perception and providing a solid information basis for subsequent task execution (such as pursuit, search and rescue, etc.).
[0216] Optionally, each target is defined by an identifier (ID, Identifier) and the target is tracked.
[0217] The present invention aims to provide a multi-UAV collaborative target perception system 200, which obtains multiple sets of image data through multiple UAVs and performs data processing, so as to achieve multi-UAV collaborative target perception. Multi-UAV collaborative perception helps to make full use of the spatio-temporal asynchronous information between multiple UAVs to search for targets in a larger range in cities, villages, and streets, and can overcome the limitations of a single UAV in terms of coverage, endurance, data processing ability, etc., improving the overall search efficiency.
[0218] It should be emphasized that the multi-UAV collaborative target perception system 200 can achieve multi-UAV collaborative target perception in scenarios such as urban streets, villages, and mountainous areas, complete tasks such as pursuit and search and rescue, overcome the limitations of a single UAV in terms of coverage, endurance, data processing ability, etc., and improve the overall search efficiency.
[0219] In an embodiment according to the present invention, as Figure 8 shown, the electronic device 300 includes a memory 310 and a processor 320. Among them, a program or instruction that can run on the processor 320 is stored on the memory 310, and when the processor 320 executes the program or instruction, the steps of the multi-UAV collaborative target perception method in any of the above embodiments are implemented. Therefore, the electronic device 300 has the beneficial effects of any of the above embodiments, which will not be elaborated here.
[0220] In an embodiment according to the present invention, a readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the multi-UAV collaborative target perception method in any of the above embodiments are implemented. Therefore, the readable storage medium has the beneficial effects of any of the above embodiments, which will not be elaborated here.
[0221] In the present invention, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance; the term "multiple" means two or more, unless otherwise clearly defined. Terms such as "installation", "connection", "connection", and "fixation" should all be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; "connection" can be a direct connection or an indirect connection through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0222] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "front", "rear", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or unit referred to must have a specific direction, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0223] In the description of this specification, the description of terms such as "one embodiment", "some embodiments", "specific embodiments", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or instance. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0224] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-UAV cooperative target perception method, characterized in that: include: Acquire multiple sets of image data taken by multiple drones; wherein the multiple sets of image data include first image data and second image data, the first image data comes from one of the multiple drones, and the second image data comes from another of the multiple drones; Based on a twin feature extraction network, the first image data and the second image data are divided into a plurality of image blocks and feature extraction is performed to obtain a plurality of image features; Calculating a Euclidean distance vector between a plurality of said image features; Based on the top-level network, calculating the image matching probability between the plurality of image blocks according to the Euclidean distance vector, wherein the image matching probability is used to represent the probability that two image blocks belong to the same object; When the image matching probability is greater than a first threshold, at least two of the image blocks are classified into the same group of matching image pairs; Based on the multiple groups of matching image pairs, the association information between the targets detected by the multiple drones is determined to achieve multi-drone collaborative target perception.
2. The multi-UAV cooperative target perception method according to claim 1 is characterized in that: Also includes: In the twin feature extraction network-based method, the first image data and the second image data are divided into multiple image blocks and feature extraction is performed. Before obtaining multiple image features, data cleaning and data enhancement are performed on the first image data and the second image data.
3. The multi-UAV cooperative target perception method according to claim 1 is characterized in that: The first image data includes a first target image and a first environment image; and / or The second image data includes a second target image and a second environment image.
4. The multi-UAV cooperative target perception method according to any one of claims 1 to 3, characterized in that: The step of dividing the first image data and the second image data into a plurality of image blocks and performing feature extraction based on the twin feature extraction network to obtain a plurality of image features includes: Based on the twin feature extraction network, the image in the first image data is divided into a plurality of first image blocks and feature extraction is performed to obtain a plurality of first image features; Based on the twin feature extraction network, the image in the second image data is divided into multiple second image blocks and feature extraction is performed to obtain multiple second image features.
5. The multi-UAV cooperative target perception method according to claim 4 is characterized in that: The step of dividing the image in the first image data into a plurality of first image blocks and performing feature extraction based on the twin feature extraction network to obtain a plurality of first image features includes: Classifying a plurality of images in the first image data into a first to-be-matched image group; Inputting the first to-be-matched image group into the first sub-network of the twin feature extraction network; Based on the first sub-network of the twin feature extraction network, the images in the first to-be-matched image group are divided into a plurality of first image blocks and feature extraction is performed to obtain a plurality of first image features.
6. The multi-UAV cooperative target perception method according to claim 4 is characterized in that: The step of dividing the image in the second image data into a plurality of second image blocks and performing feature extraction based on the twin feature extraction network to obtain a plurality of second image features includes: classifying a plurality of images in the second image data into a second group of images to be matched; Inputting the second to-be-matched image group into the second sub-network of the twin feature extraction network; Based on the second sub-network of the twin feature extraction network, the images in the second group of images to be matched are divided into a plurality of second image blocks and feature extraction is performed to obtain a plurality of second image features.
7. The multi-UAV cooperative target perception method according to claim 4 is characterized in that: The calculating of the Euclidean distance vectors between the plurality of image features comprises: The Euclidean distance between each of the first image features and each of the second image features is calculated, and the Euclidean distance vector is determined according to a plurality of the Euclidean distances.
8. A multi-UAV cooperative target perception system, characterized in that: include: An image data acquisition module (210) is used to acquire multiple sets of image data captured by multiple drones; wherein the multiple sets of image data include first image data and second image data, the first image data comes from one of the multiple drones, and the second image data comes from another of the multiple drones; A feature extraction module (220), configured to divide the first image data and the second image data into a plurality of image blocks and perform feature extraction based on a twin feature extraction network to obtain a plurality of image features; A first calculation module (230), used for calculating the Euclidean distance vectors between a plurality of the image features; A second calculation module (240) is used to calculate the image matching probability between the plurality of image blocks based on the top network and the Euclidean distance vector, wherein the image matching probability is used to represent the probability that two image blocks belong to the same object; An image matching module (250) is used to classify at least two of the image blocks into the same group of matching image pairs when the image matching probability is greater than a first threshold; The information determination module (260) is used to determine the association information between the targets detected by the multiple drones based on the multiple sets of matching image pairs, so as to achieve multi-drone collaborative target perception.
9. An electronic device, characterized in that: include: A memory (310) and a processor (320), wherein the memory (310) stores a program or instruction that can be run on the processor (320), and when the processor (320) executes the program or the instruction, the steps of the multi-UAV collaborative target perception method according to any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or an instruction, and when the program or the instruction is executed by the processor, the steps of the multi-UAV collaborative target perception method according to any one of claims 1 to 7 are implemented.