A clustering method for objects in drone aerial images

Through deep learning algorithms and twin neural networks to calculate the difference in image similarity and latitude and longitude coordinates, the difficulty of identifying the same object in aerial images in drone aerial images is solved, and effective clustering and recognition of objects in multi-view images is achieved.

CN114743114BActive Publication Date: 2025-05-13ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210264446.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-05-13
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

In the aerial images of the drone, the same object shows great differences in different images due to factors such as shooting posture, angle and occlusion, which makes it difficult to determine the same object in different aerial images.

Method used

Deep learning algorithm combined with twin neural networks is used to calculate image similarity and combine latitude and longitude coordinate differences to judge the distance reachability between images, thereby achieving clustering of objects in multi-view images.

Benefits of technology

It improves the accuracy of the recognition of the same object in the aerial image of the drone, solves the problem that the image results cannot be mapped to the actual objects, and achieves a clearer semantics of the target detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114743114B_ABST
    Figure CN114743114B_ABST
Patent Text Reader

Abstract

A method for clustering objects in drone aerial images can realize clustering of objects existing in multi-view aerial images, make the semantics of target detection results clearer, and solve the problem that image results in drone inspections cannot be mapped to actual objects, including: (1) cutting out two sub-images, one large and one small, according to the marking box of the object in the image; (2) pairing each group of sub-images as the input of the central-ring double-stream twin neural network; (3) calculating the similarity distance through the output of the neural network and the longitude and latitude coordinates of the center of the object in each corresponding sub-image, and the image pairs with a similarity distance less than a threshold are considered to be distance-reachable; (4) transferring distance reachability to obtain a set of images whose internal elements are mutually reachable as a clustering result. The present invention has the following advantages: the method has strong universality; high accuracy; high computational efficiency; and concise results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of image processing and relates to a method for clustering objects in multi-view images. Background Art

[0002] UAV is a remote sensing platform that is easy to operate and transfer. UAV aerial images have the advantages of high definition, large scale, small area, and high currentness. They are widely used in national ecological environment protection, mineral resource exploration, marine environment monitoring, land use survey, water resource development, crop growth monitoring and yield estimation, agricultural operations, natural disaster monitoring and evaluation, urban planning and municipal management, forest pest protection and monitoring, digital earth and other fields. Especially in the field of inspection, the inspection model of UAV plus UAV nest is gradually being promoted.

[0003] Usually, drone aerial images have a certain degree of overlap, and the target object will appear in multiple aerial images. At present, it is relatively easy to mark specific targets in a single aerial image through deep learning or manual means, but the same object in different images often shows great differences due to factors such as shooting posture, angle and occlusion. There is still an urgent need for an effective clustering method to identify the same object in different aerial images. Summary of the invention

[0004] In order to solve the clustering problem of the same object in different aerial images, the present invention designs a clustering method for objects in drone aerial images, which realizes the clustering of objects that appear repeatedly in different drone aerial images. Under the condition of using deep learning algorithms to detect targets in drone aerial images or manually annotating targets, this method calculates the image similarity of the small image where the object is located through a twin neural network, and combines the difference in longitude and latitude coordinates to calculate whether the two objects are reachable by distance, and clusters them according to the distance accessibility between the objects. The specific steps are as follows:

[0005] 1) According to the marked box of the object in the image, two sub-images, one large and one small, are cut out;

[0006] 2) Pair each group of subgraphs separately as the input of the mid-loop two-stream twin neural network;

[0007] 3) The similarity distance is calculated through the output of the neural network and the longitude and latitude coordinates of the center of the corresponding sub-image. The image pairs with a similarity distance less than the threshold are considered to be reachable;

[0008] 4) Transfer distance reachability to obtain a set of images whose internal elements are mutually reachable as the clustering result.

[0009] The clustering method of objects in drone aerial images relies on the object labeling results, mines the similarities between images through the twin neural network, and combines the GPS distance to determine whether the objects in different images are reachable, thereby further realizing clustering.

[0010] In the above-mentioned method for clustering objects in drone aerial images, the positions of objects in the aerial images are marked manually or by a machine learning algorithm.

[0011] Preferably, in the above-mentioned method for clustering objects in drone aerial images, in step (1), a sub-image is cut out according to the object marking frame and recorded as x′, and then the marking frame is enlarged to cut out a slightly larger sub-image and recorded as x.

[0012] In the above-mentioned method for clustering objects in drone aerial images, in step (1), when the border is enlarged to obtain the sub-image x, if it touches the boundary of the original image, the enlargement of the border is stopped.

[0013] Preferably, in the above-mentioned method for clustering objects in drone aerial images, the central-ring dual-stream twin neural network described in step (2) refers to a central stream and surround stream twin neural network composed of two sets of convolutional layers and pooling layers, the input of the central stream is a pair of subgraphs (x′1, x′2), and the input of the surround stream is a pair of subgraphs (x1, x2).

[0014] The similarity distance calculated by the decision layer of the above-mentioned middle-loop two-stream twin neural network is expressed as:

[0015]

[0016] The learning objective function of the middle-loop two-stream twin neural network is:

[0017]

[0018] Among them, ω represents the weight of the neural network; f ω represents the output of the neural network; m represents the number of samples; y i Indicates whether the i-th sample is similar, if similar, it is 1, otherwise it is 0; β represents the similarity threshold, D ω If it exceeds β, it is considered to be dissimilar. is the regularization term.

[0019] The activation function of the middle-loop two-stream twin neural network is the ReLu function.

[0020] Furthermore, in the above-mentioned method for clustering objects in drone aerial images, in step (2), when allocating image pairs, a pair of sub-images originating from the same aerial image is considered to be unreachable and does not enter the neural network for calculation; if the center longitude and latitude coordinate paradigm distance of the image pair exceeds 3 times the average positioning error of the longitude and latitude positioning, it is considered to be unreachable and does not enter the neural network for calculation.

[0021] Furthermore, the mid-loop two-stream twin neural network can accept input images of different resolutions.

[0022] Preferably, in the above-mentioned method for clustering objects in drone aerial images, in step (3), the decision layer of the neural network uses a convolutional layer instead of a fully connected layer, and the similarity distance is calculated using the paradigm distance.

[0023] In the above-mentioned method for clustering objects in drone aerial images, in step (3), the paradigm distance of the longitude and latitude coordinates of the image pairs is calculated, and the distance accessibility is determined in combination with the similarity distance output by the neural network.

[0024] Preferably, in the above-mentioned method for clustering objects in drone aerial images, in the process of transferring distance reachability in step (4), if a sub-graph is simultaneously reachable to two or more sub-graphs originating from the same aerial image, the one with the smallest similarity distance is selected for reachability transfer, and its distance from the remaining sub-graphs is changed to be unreachable.

[0025] The technical concept of the present invention is: to identify specific targets in aerial images through deep learning algorithms; to calculate image similarity using twin neural networks; to improve calculation accuracy using central flow and surround flow in the neural network; to improve the accuracy of accessibility judgment by combining longitude and latitude coordinate differences; and to obtain specific classification of objects through accessibility transfer. The present invention can realize the clustering of objects existing in multi-view aerial images, making the semantics of target detection results clearer, and solving the problem that image results in drone inspections cannot be mapped to actual objects.

[0026] The advantages of the present invention are: the method has strong universality; high accuracy; high computational efficiency; and concise results. In the present invention, a ring-shaped dual-stream twin neural network is used to judge similarity, and the image features of the marking box where the object is located are more concerned with the implicit central stream in the network. For objects whose features are not obvious, or occluded or non-rigidly deformed objects caused by changes in the perspective of drone shooting, the accuracy of similarity judgment can be improved through the environmental information in the surrounding stream; the convolution layer is used in the neural network instead of the fully connected layer, so that the network can accept images of different resolutions as input; the longitude and latitude information corresponding to the image pixel coordinates is added to improve the clustering accuracy and reduce unnecessary similarity calculations; the clustering scheme is designed based on the characteristics of drone shooting using distance accessibility instead of density accessibility, which ensures the clustering effect of areas with low overlap in aerial photography coverage, while avoiding excessive clustering of images with high density caused by the actual number of target objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a specific input image cropping schematic diagram of the present invention.

[0028] Figure 2 This is the middle-loop dual-stream twin neural network architecture diagram of the present invention.

[0029] Figure 3 It is a display diagram of the specific clustering result of the present invention in three-dimensional geographic data. DETAILED DESCRIPTION

[0030] With reference to the accompanying drawings, the present invention is further described by taking the clustering of dying or dead pine trees in an aerial image as an example:

[0031] The clustering method of objects in drone aerial images uses deep learning-based target detection algorithms such as YOLO to identify specific targets in aerial images, calculates the image similarity of the small image where the object is located through the twin neural network, and combines the difference in longitude and latitude coordinates to calculate whether the two images are reachable by distance, and divides all images into several sets with no repeated elements and reachable distances between elements in the set according to the distance reachability between images.

[0032] In this example, the drone aerial photos were taken by oblique photography with five flight routes, with a set heading overlap of 80% and a lateral overlap of 70%. The dead tree identification frame was identified by using the YOLO5 framework. After multiple tests, the recognition accuracy of the target detection algorithm was above 86%, which means that most dead trees could be detected. On average, each dead tree appeared in 6 different aerial images. The clustering method was used to describe the image results of target detection as objects with actual physical meaning.

[0033] A method for clustering objects in drone aerial images comprises the following steps:

[0034] 1) According to the marked box of the object in the image, two sub-images, one large and one small, are cut out;

[0035] 2) Pair each group of subgraphs separately as the input of the mid-loop two-stream twin neural network;

[0036] 3) The similarity distance is calculated by the output of the neural network and the longitude and latitude coordinates of the center of the corresponding sub-image. The image pairs with a similarity distance less than the threshold are considered to be reachable;

[0037] 4) Transfer distance reachability to obtain a set of images whose internal elements are mutually reachable as the clustering result.

[0038] In step (1), a sub-image is cut out according to the object marking frame and recorded as x′, and then the marking frame is enlarged to cut out a slightly larger sub-image and recorded as x.

[0039] Furthermore, the enlarged mark frame refers to a frame with the original mark center as the center, and each side is enlarged at the same proportion. If the boundary of the original image is touched when the side length is enlarged, the expansion is stopped.

[0040] Specifically, the magnification factor is determined by whether the features of the object itself are obvious and whether the surrounding environment information is obvious. Usually, it can be magnified to 2-5 times the original size.

[0041] Figure 1 It is a schematic diagram of cutting two sub-images, one large and one small, from the original image, and the central pixels of the two sub-images are the same.

[0042] In step (2), when allocating image pairs, a pair of sub-images from the same aerial image is considered to be unreachable and will not be included in the neural network for calculation; if the center longitude and latitude coordinate paradigm distance of an image pair exceeds 3 times the average positioning error of the longitude and latitude positioning, it is considered to be unreachable and will not be included in the neural network for calculation.

[0043] Furthermore, allocating image pairs refers to dividing the images into two groups of large and small images (x′1, x1) and (x′2, x2), a total of 4 images, and dividing them into (x′1, x′2) and (x1, x2) according to the large image to the large image and the small image to the small image.

[0044] The central-ring dual-stream twin neural network described in step (2) refers to a central stream and a surrounding stream twin neural network composed of two sets of convolutional layers and pooling layers. The input of the central stream is a pair of subgraphs (x′1, x′2), and the input of the surrounding stream is a pair of subgraphs (x1, x2). The twin networks in the central stream and the surrounding stream share weights. The output enters the decision layer with two convolutional layers replacing the fully connected layer.

[0045] Furthermore, the similarity distance calculated by the decision layer of the middle-loop two-stream twin neural network is expressed as:

[0046]

[0047] The learning objective function is:

[0048]

[0049] Specifically, ω represents the weight of the neural network; f ω represents the output of the neural network; m represents the number of samples; y i Indicates whether the i-th sample is similar, if similar, it is 1, otherwise it is 0; β represents the similarity threshold, D ω If it exceeds β, it is considered to be dissimilar. is the regularization term.

[0050] Preferably, a ReLu function is used as the activation function.

[0051] Figure 2 Schematic diagram of the architecture of the mid-loop two-stream twin neural network.

[0052] The longitude and latitude distance between the image pairs described in step (3) is expressed as:

[0053] D lla =(lon1-lon2) 2 +(lat1-lat2) 2 (3)

[0054] Specifically, (lon1, lat1) and (lon2, lat2) are the latitude and longitude coordinates of the image pair approximated to meters. For my country, the latitude and longitude coordinates are multiplied by 10. 5 An approximate solution can be obtained.

[0055] The neural network output D ω Perform min-max normalization according to [0,β], denoted as D' ω ;D lla Then the min-max normalization is performed with 3 times the positioning error as the boundary, denoted as D' lla , according to D reachable =α1D' ω +α2D' lla (α1, α2 are weight coefficients, satisfying α1+α2=1) to obtain the reachable distance, β reachable is the distance reachable threshold, satisfying D reachable ≤β reachable is considered to be reachable by distance.

[0056] In step (4), when object 1 is reachable from object 2, object 2 is reachable from object 3, and object 1, object 2, and object 3 are from different aerial images, it is considered that object 1 is also reachable from object 3. Object 1, object 2, and object 3 belong to the same reachable set.

[0057] Preferably, the union-find method is used to transfer reachability. During the transfer process, the existence of objects from the same aerial image is detected. If it is found that an object in a certain set is reachable to two or more sets that cannot be merged for reachability, the set with a shorter reachability distance is selected to enter, and the reachability of other sets is modified to unreachable.

[0058] Figure 3 This is the display of clustering results in 3D geographic data. As shown in the figure, the dead trees circled by white circles are obtained by clustering the markers in a set of aerial images, the white box marks the sub-image x, and the black box marks the corresponding x'.

[0059] The technical concept of the present invention is: using twin neural networks to calculate image similarity; using central flow and surround flow in the neural network to improve calculation accuracy; combining the differences in longitude and latitude coordinates to improve the accuracy of reachability judgment; and obtaining specific classification of objects through reachability transfer.

[0060] The present invention realizes the clustering of objects existing in multi-view aerial images, makes the semantics of target detection results clearer, and solves the problem that image results in drone inspections cannot be mapped to actual objects.

[0061] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A clustering method for objects in drone aerial images, characterized in that: The following steps are involved: 1) According to the marked box of the object in the image, two sub-images, one large and one small, are cut out; 2) Pairing each group of subgraphs as input to the center-loop two-stream twin neural network; the center-loop two-stream twin neural network refers to a center stream and surround stream twin neural network composed of two sets of convolutional layers and pooling layers, the input of the center stream is a pair of subgraphs (x1′, x′2), and the input of the surround stream is a pair of subgraphs (x1, x2); The similarity distance calculated by the decision layer of the above-mentioned middle-loop two-stream twin neural network is expressed as: The learning objective function of the middle-loop two-stream twin neural network is: Among them, ω represents the weight of the neural network; f ω represents the output of the neural network; m represents the number of samples; y i Indicates whether the i-th sample is similar, if similar, it is 1, otherwise it is 0; β represents the similarity threshold, D ω If it exceeds β, it is considered to be dissimilar. is the regularization term; The activation function of the middle-loop two-stream twin neural network is the ReLu function; 3) The similarity distance is calculated through the output of the neural network and the longitude and latitude coordinates of the center of the corresponding sub-image. The image pairs with a similarity distance less than the threshold are considered to be reachable; 4) Transfer distance reachability to obtain a set of images whose internal elements are mutually reachable as the clustering result.

2. The method for clustering objects in drone aerial images according to claim 1, characterized in that: In step (1), a sub-image is cut out according to the object marking frame and recorded as x′, and then the marking frame is enlarged to cut out a slightly larger sub-image and recorded as x; When expanding the border to obtain sub-image x, if it touches the boundary of the original image, stop expanding the border.

3. The method for clustering objects in drone aerial images according to claim 1, characterized in that: The described mid-loop two-stream twin neural network can accept input images of different resolutions.

4. The method for clustering objects in drone aerial images according to claim 1, characterized in that: When allocating image pairs to enter the center-ring dual-stream twin neural network calculation, a pair of sub-images from the same aerial photo is considered to be unreachable and will not be calculated in the neural network; if the center longitude and latitude coordinate paradigm distance of the image pair exceeds 3 times the average positioning error of the longitude and latitude positioning, it is considered to be unreachable and will not be calculated in the neural network.

5. The method for clustering objects in drone aerial images according to claim 1, characterized in that: In step (3), the paradigm distance of the longitude and latitude coordinates of the image pair is calculated, and the distance accessibility is determined in combination with the similarity distance output by the neural network, specifically including: the longitude and latitude distance between the image pairs is expressed as: D lla =(day1-day2) 2 +(lat1-lat2) 2 (3) (lon1, lat1), (lon2, lat2) are the latitude and longitude coordinates of the image pair approximated to meters. For my country, the latitude and longitude coordinates are multiplied by 10. 5 An approximate solution can be obtained; The neural network output D ω Perform min-max normalization according to [0,β], denoted as D' ω ;D lla Then the min-max normalization is performed with 3 times the positioning error as the boundary, denoted as D' lla , according to D reachable =α1D' ω +α2D' lla Get the reachable distance, where α1, α2 are weight coefficients, satisfying α1+α2=1; β reachable is the distance reachable threshold, satisfying D reachable ≤β reachable is considered to be reachable by distance.

6. The method for clustering objects in drone aerial images according to claim 1, characterized in that: In step (4), distance reachability is transferred, and images with reachable distances belong to the same set. During the process of transferring distance reachability, it is necessary to detect the existence of objects from the same aerial image. If it is found that an object from a certain set is reachable to two or more sets that cannot be merged by reachability, the set with a shorter reachable distance is selected to enter, and the reachability to other sets is modified to be unreachable.

Citation Information

Patent Citations

  • Pedestrian tracking method based on twin neural network

    CN111814604A

  • Single target tracking method based on clustering difference and deep twin convolutional neural network

    CN113808166A