Target identification method and system under view angle of unmanned aerial vehicle
Multi-view images are acquired through multi-camera and registered and fused, which solves the problem of low recognition accuracy caused by obstruction of targets during drone flight, and achieves higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510115096.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-24
AI Technical Summary
During the flight of the drone, the target's perspective changes constantly, and the target is prone to being blocked, making it difficult to maintain high accuracy in the target recognition results.
The drone is equipped with multiple cameras to capture the target area from different angles, obtain multi-view image data, and use image registration and fusion technology to supplement the obstructed area to generate a complete target view.
Through the fusion of multi-view images, information loss caused by single-view occlusion is reduced, the accuracy and robustness of target recognition are improved, and the stability and reliability of feature points are ensured.
Smart Images

Figure CN120014495A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles, and in particular to a target recognition method and system from the perspective of an unmanned aerial vehicle. Background Art
[0002] Drones are widely used in many fields such as industry, agriculture, and environmental monitoring. Target recognition technology can help quickly locate problem areas and provide timely data support so that appropriate protective measures can be taken. It not only improves efficiency, but also changes the way traditional industries operate. With the continuous development of technology, drones have great potential for application in more fields. Target recognition technology, as one of the core applications of drones, has promoted the advancement of drone intelligence and automation, and brought innovation and change to all walks of life.
[0003] During the flight of a drone, the target's perspective keeps changing, and it is easy for the target to be partially occluded, which will make it difficult to maintain high accuracy in target recognition results. Therefore, how to fuse multi-perspective images to restore the occluded area and improve the accuracy of target recognition is a problem we need to solve. To this end, a target recognition method and system from the perspective of a drone are proposed. Summary of the invention
[0004] The purpose of the present invention is to provide a target recognition method and system from the perspective of a drone to solve the problems raised in the above background technology.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: In a first aspect, a method for identifying a target from the perspective of a drone comprises the following steps: Step 1: Use multiple cameras carried by the drone to simultaneously shoot the target area from different angles, and perform preprocessing to obtain multi-view image data; Step 2: Register the preprocessed multi-view images, apply convolutional neural network to extract image features from each view, and establish the association between different views through feature matching; Step 3: Based on the feature matching results, the multi-view images are fused, and then the fusion algorithm is used to supplement the occluded area to generate a complete target view; Step 4: Automatically detect and identify the target object by combining the fused image data; Step 5: Verify the results of target detection and recognition and analyze the accuracy of target recognition.
[0006] A further improvement of the technical solution of the present invention is that in step 1, the process of acquiring multi-view image data includes: Identify the target area to be photographed and the type of information to be obtained from different angles. Based on the shooting requirements, equip the drone with multiple cameras and adjust the angle and focal length of the camera to ensure that detailed information of the target area can be captured from different perspectives. Perform a detailed inspection of the drone and camera to ensure that they are in good working condition, including battery power, signal connection, camera clarity, etc. Plan the flight route, altitude and speed of the drone according to the terrain, weather conditions and shooting requirements of the target area, ensure that all angles and positions required for shooting are covered, ensure that the flight route avoids obstacles and no-fly zones, and comply with the flight regulations and safety regulations of the target area; Operate the drone to shoot according to the set flight plan, control the drone to fly along the predetermined flight route and altitude, and start multiple cameras for synchronous shooting; The drone transmits the multi-view image data it captures to the ground control station in real time, which is then sorted out, and photos that obviously do not meet the requirements (such as overexposure and blur) are deleted. The remaining images are then pre-processed with denoising and color correction. An image database is established to record the shooting time and geographic location information of each picture, and the acquired multi-view image data is stored in the image database for use in subsequent processing.
[0007] A further improvement of the technical solution of the present invention is that in step 2, the registration process of the multi-view images includes: Select an image with the most central perspective from the multi-view images as the reference image, use the SIFT feature detection algorithm to detect key points in the reference image and each image to be registered, and generate a 128-dimensional descriptor vector for each key point, which describes the local image features around the key point; Use the BFMatcher matching algorithm to find the corresponding key point pairs between the reference image and the image to be registered, use the RANSAC algorithm to remove the wrong matching points, and estimate the geometric transformation model, and then use the geometric transformation model to align the image to be registered to the spatial coordinate system of the reference image; Extract image features for image fusion from the registered multi-view images, including color features, texture features, shape features, spatial relationship features, and multi-view consistency features. Input the registered images into the convolutional neural network model. After passing through multiple convolutional layers and pooling layers of the convolutional neural network model, a series of feature maps are obtained. For each image from each perspective, local feature vectors are extracted from the selected feature map, and the similarity between feature vectors from different perspectives is calculated using similarity measurement. Based on the similarity measurement results, feature matching relationships between images from different perspectives are established, and matching thresholds are set to determine which pairs are considered matching pairs. Pairs above the matching threshold are considered valid matching relationships. The feature matching results between each view image and the reference image are saved, including matching point pairs and their corresponding feature descriptors.
[0008] A further improvement of the technical solution of the present invention is that the calculation formula of the similarity metric is: ; In the formula, represents the similarity measure, represents the dimension of the feature vector, Indicates The perspective image The local eigenvector Dimensional component, Indicates The perspective image The local eigenvector Dimensional component, The value range is between 0 and 1. When two feature vectors are exactly the same, the similarity reaches the maximum value of 1.
[0009] A further improvement of the technical solution of the present invention is that in step 3, the process of generating the complete target view includes: According to the calculated similarity measurement results, confirm each pair of matching points and their corresponding feature descriptors, and organize all valid matching point pairs and their corresponding coordinate information for subsequent use; Create a blank target view image whose size should be able to accommodate the content of all input images. The size of the target view image is a 16x16 pixel block. For each local area of the 16x16 pixel block, determine the corresponding position of the area in the different view images based on the information of the matching point pair, and use the weighted averaging method to perform weighted averaging on the pixel values in the overlapping area to fuse the image contents of multiple view angles into the target view. By analyzing the matching point pairs and feature descriptors, the occluded area is identified. For the detected occluded area, its size is m×n pixels. For each pixel position in the occluded area, the new pixel value after filling is calculated and filled using the image restoration algorithm based on content-aware filling. The filling algorithm is inferred and generated based on the texture, color and structure information of the surrounding pixels; The filled image is fused with the original image to ensure the natural transition and consistency of the fused area, and the fused image is post-processed including denoising, sharpening, color correction, etc., and the processed image is used as the final output to obtain a complete target view.
[0010] A further improvement of the technical solution of the present invention is that the calculation expression of the new pixel value after filling is: ; ; ; In the formula, Represents the new pixel value after filling. represents the center point, Indicated in Radius around point The actual pixel value in the range, represents the weight function, defining the importance of pixels in different directions, represents the energy difference with the center pixel, and Relative to the center point The vertical and horizontal offsets, represents the difference between two pixels, is a smaller than An integer that limits the size of the local region to be compared.
[0011] A further improvement of the technical solution of the present invention is that: in the step 4, the process of identifying the target object includes: Locate the position of the target object in the fused image, and detect feature points including texture features and shape features from the image, and then generate a descriptor for each detected feature point to describe the texture shape information of its local area; Use the sliding window method to traverse the entire image, use a fixed-size window to extract features at each position, match the extracted features with the features in the database, preset the matching threshold, and classify the content in the window to calculate the matching score to determine whether it is the target object; The detection results are post-processed and the non-maximum suppression algorithm is applied to remove the target frames of repeated detections to improve the quality of the final output. The positions of the identified target objects are marked on the image, and the positions of the identified target objects and related information are output.
[0012] A further improvement of the technical solution of the present invention is that the calculation expression of the matching score is: ; In the formula, represents the matching score, Indicates The matching score of feature points, Indicates the number of feature points extracted within the window, Represents a small constant (such as 0.01) to avoid the denominator being zero. Represents the preset matching threshold, represents the ideal upper limit of the matching score, and represents the ideal similarity between feature points. Indicates selection and The smaller one of the two ensures that the matching score does not exceed the ideal upper limit. The value range of is generally between [0,1]. Also within this range, so The maximum value of , The value range is between 0 and 1.
[0013] A further improvement of the technical solution of the present invention is that in step 5, the process of analyzing the target recognition accuracy includes: Obtain a set of real scene images containing target objects, and use bounding boxes to mark the location and size of the targets to accurately label the target objects in the images; Compare the target detection and recognition results with the target objects in the real scene image, analyze the color features, texture features, shape features and spatial relationship features of the two, clarify the matching degree of each feature, and analyze the accuracy of target recognition; Visualization tools are used to display the analysis results of target recognition accuracy, as well as the location and size of target objects.
[0014] In a second aspect, a target recognition system from the perspective of a drone is used to implement the target recognition method from the perspective of a drone, including an image data management platform, wherein the image data management platform is communicatively connected to an image acquisition module, a multi-view image registration and fusion module, a feature extraction and matching module, and a target detection and classification module, wherein electrical signals are connected between the modules; The image acquisition module uses the drone to capture multi-view images of the target object and provides high-quality original image data, laying the foundation for subsequent image processing and recognition; The multi-view image registration and fusion module detects key points in the reference image and the image to be registered using a feature detection algorithm, and aligns the images to the same coordinate system, thereby fusing the image contents of multiple views into the target view and supplementing the occluded areas; The feature extraction and matching module is used to extract color features, texture features, shape features, spatial relationship features and multi-view consistency features from the registered images, generate feature maps through convolutional neural networks, calculate the similarity of feature vectors between different viewpoints, and establish feature matching relationships between images of different viewpoints; The target detection and classification module uses a fixed-size window to traverse the entire image, extract features and match them with features in a database, and classify the content in the window to determine whether it is a target object.
[0015] Due to the adoption of the above technical solution, the present invention has the following technical advances compared with the prior art: 1. The present invention provides a target recognition method and system from the perspective of an unmanned aerial vehicle. The target area is photographed from different angles by multiple cameras carried by the unmanned aerial vehicle to obtain multi-perspective image data, which helps to eliminate the occlusion problem existing in a single perspective. Secondly, the obtained comprehensive target view is used to reduce information loss caused by single-perspective occlusion, ensure the stability and reliability of feature points, and improve the recognition accuracy and robustness of the entire system.
[0016] 2. The present invention provides a method and system for target recognition from the perspective of a drone. By registering and fusing images from multiple perspectives, the problem of information loss caused by occlusion can be effectively eliminated. The fused image not only contains more detailed information, but also has higher resolution and clarity, which helps to improve the accuracy of target detection and makes feature extraction more accurate. The complete target view generated by image fusion technology further expands the target recognition accuracy from the perspective of a drone. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic diagram of the method flow of the present invention; Figure 2 Schematic diagram of the registration workflow of multi-view images of the present invention; Figure 3 A schematic diagram of a workflow for generating a complete target view of the present invention; Figure 4 It is a schematic diagram of the functional modules of the system of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] Embodiment 1, as Figure 1 , Figure 2 As shown, the present invention provides a target recognition method from the perspective of a drone, comprising the following steps: Step 1: Use multiple cameras on the drone to simultaneously shoot the target area from different angles, perform preprocessing to obtain multi-view image data, clarify the target area to be photographed, and the type of information to be obtained from different angles, and according to the shooting requirements, install multiple cameras on the drone, adjust the angle and focal length of the camera to ensure that detailed information of the target area can be captured from different perspectives, and conduct a detailed inspection of the drone and camera to ensure that they are in good working condition, including battery power, signal connection, camera clarity, etc. According to the terrain, weather conditions and shooting requirements of the target area, plan the flight route, altitude and speed of the drone to ensure that all angles and positions that need to be shot can be covered , ensure that the flight route avoids obstacles and no-fly zones, abides by the flight regulations and safety provisions of the target area, operates the drone to perform shooting operations according to the set flight plan, controls the drone to fly according to the predetermined flight route and altitude, and starts multiple cameras for synchronous shooting. The drone transmits the captured multi-view image data to the ground control station in real time, performs preliminary sorting, deletes photos that obviously do not meet the requirements (such as overexposure, blur), and performs pre-processing operations such as denoising and color correction on the retained images, establishes an image database, records the shooting time and geographic location information of each picture, and stores the acquired multi-view image data in the image database for subsequent processing; Step 2: Register the preprocessed multi-view images, and apply convolutional neural networks to extract image features from images of each view, and establish associations between different viewpoints through feature matching. Select an image with the most central viewpoint from the multi-view images as the reference image, and use the SIFT feature detection algorithm to detect key points in the reference image and each image to be registered. For each key point, generate a 128-dimensional descriptor vector, which describes the local image features around the key point. Use the BFMatcher (Brute-Force Matcher) matching algorithm to find the corresponding key point pairs between the reference image and the image to be registered, use the RANSAC (Random Sample Consensus) algorithm to remove erroneous matching points, and estimate the geometric transformation model, and then apply The image to be registered is aligned to the spatial coordinate system of the reference image using a geometric transformation model. Image features for image fusion are extracted from the registered multi-view images, including color features, texture features, shape features, spatial relationship features, and multi-view consistency features. The registered images are input into a convolutional neural network model. After passing through multiple convolutional layers and pooling layers of the convolutional neural network model, a series of feature maps are obtained. For each view image, local feature vectors are extracted from the selected feature maps, and the similarity of feature vectors between different view angles is calculated using a similarity metric. Based on the similarity metric results, a feature matching relationship between images of different view angles is established. A matching threshold is set to determine whether a pair is considered a matching pair. A pairing above the matching threshold is considered a valid matching relationship. The feature matching results between each view image and the reference image are saved, including matching point pairs and their corresponding feature descriptors. Furthermore, the calculation formula of the similarity measure is: ; In the formula, represents the similarity measure, represents the dimension of the feature vector, Indicates The perspective image The local eigenvector Dimensional component, Indicates The perspective image The local eigenvector Dimensional component, The value range is between 0 and 1. When the two eigenvectors are exactly the same, the similarity reaches the maximum value of 1. When the two eigenvectors are completely different (that is, there is no common non-zero component), the similarity is close to 0. Step 3: Based on the feature matching results, the multi-view images are fused, and then the fusion algorithm is used to supplement the occluded area to generate a complete target view; Step 4: Automatically detect and identify the target object by combining the fused image data; Step 5: Verify the results of target detection and recognition and analyze the accuracy of target recognition.
[0021] Embodiment 2, as Figure 3 As shown, based on Example 1, the present invention provides a technical solution: Preferably, in step three, the process of generating a complete target view includes: According to the calculated similarity measurement results, confirm each pair of matching points and their corresponding feature descriptors, and organize all valid matching point pairs and their corresponding coordinate information for subsequent use, create a blank target view image, the size of which should be able to accommodate the content of all input images. The size of the target view image is a 16x16 pixel block. For each local area of the 16x16 pixel block, determine the corresponding position of the area in the image of different viewing angles based on the information of the matching point pair, and apply the weighted averaging method to perform weighted averaging on the pixel values in the overlapping area, and fuse the image content of multiple viewing angles into the target view. By analyzing the matching points For the and feature descriptors, the occluded area is identified. For the detected occluded area, its size is m×n pixels. For each pixel position in the occluded area, the new pixel value after filling is calculated, and the image restoration algorithm based on content-aware filling is used to fill it. The filling algorithm is based on the texture, color and structure information of the surrounding pixels. The filled image is fused with the original image to ensure the natural transition and consistency of the fused area. The fused image is post-processed including denoising, sharpening, color correction, etc., and the processed image is used as the final output to obtain a complete target view; Furthermore, the calculation expression of the new pixel value after filling is: ; ; ; In the formula, Represents the new pixel value after filling. represents the center point, Indicated in Radius around point The actual pixel value in the range, represents the weight function, defining the importance of pixels in different directions, represents the energy difference with the center pixel, and Relative to the center point The vertical and horizontal offsets, represents the difference between two pixels, is a smaller than An integer used to limit the size of the local area for comparison. The increase of (that is, the difference between the surrounding pixels and the central pixel increases), The value of gradually decreases, which means that the influence of pixels with large differences on the final result will be reduced accordingly, that is, the more similar the pixels are, the greater their contribution to the filling effect; In step 4, the process of identifying the target object includes: The target object is located in the fused image, and feature points including texture features and shape features are detected from the image. A descriptor is generated for each detected feature point to describe the texture and shape information of its local area. The sliding window method is used to traverse the entire image, and a fixed-size window is used to extract features at each position. The extracted features are matched with the features in the database, and a matching threshold is preset. The content in the window is classified and the matching score is calculated to determine whether it is a target object. The detection results are post-processed, and a non-maximum suppression algorithm is applied to remove the target frame of repeated detection to improve the quality of the final output. The position of the identified target object is marked on the image, and the position of the identified target object and related information are output; Furthermore, the calculation expression of the matching score is: ; In the formula, represents the matching score, Indicates The matching score of feature points, Indicates the number of feature points extracted within the window, Represents a small constant (such as 0.01) to avoid the denominator being zero. Represents the preset matching threshold, represents the ideal upper limit of the matching score, and represents the ideal similarity between feature points. Indicates selection and The smaller one of the two ensures that the matching score does not exceed the ideal upper limit. The value range of is generally between [0,1]. Also within this range, so The maximum value of , The value range is between 0 and 1. Very close to hour, Close to 1, it means that the features in the window are highly matched with the target features, and it is likely to be the target object. Much smaller than ,but It will be much smaller than 1, indicating that the match is not high and it is unlikely to be the target object. The number of feature points When it increases, if the matching scores are generally high, then Tends to keep high values, otherwise, if most matches have low scores, will decrease as the denominator increases; In step 5, the process of analyzing the target recognition accuracy includes: Obtain a set of real scene images containing target objects, use bounding boxes to mark the position and size of the targets to accurately label the target objects in the images, compare the target detection and recognition results with the target objects in the real scene images, analyze the color features, texture features, shape features, and spatial relationship features of the two, clarify the matching degree of each feature to analyze the accuracy of target recognition, and display the analysis results of target recognition accuracy, as well as the position and size of the target objects, through visualization tools.
[0022] Embodiment 3, as Figure 4 As shown, on the basis of Embodiment 1-2, the present invention further provides a target recognition system under the perspective of a drone, which is used to implement the target recognition method under the perspective of the drone, including an image data management platform, the image data management platform is communicatively connected with an image acquisition module, a multi-view image registration and fusion module, a feature extraction and matching module, and a target detection and classification module, wherein the modules are electrically connected by signals; The image acquisition module uses drones to capture multi-view images of target objects and provides high-quality raw image data, laying the foundation for subsequent image processing and recognition; The multi-view image registration and fusion module uses feature detection algorithms to detect key points in the reference image and the image to be registered, and aligns the images to the same coordinate system, thereby fusing the image content from multiple perspectives into the target view and filling in the occluded areas. The feature extraction and matching module is used to extract color features, texture features, shape features, spatial relationship features and multi-view consistency features from the registered images, generate feature maps through convolutional neural networks, calculate the similarity of feature vectors between different viewpoints, and establish feature matching relationships between images from different viewpoints; The target detection and classification module uses a fixed-size window to traverse the entire image, extract features and match them with the features in the database, and classify the content in the window to determine whether it is a target object.
[0023] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A target recognition method from the perspective of a drone, characterized in that: The following steps are involved: Step 1: Use multiple cameras carried by the drone to simultaneously shoot the target area from different angles, and perform preprocessing to obtain multi-view image data; Step 2: Register the preprocessed multi-view images, apply convolutional neural network to extract image features from each view, and establish the association between different views through feature matching; Step 3: Based on the feature matching results, the multi-view images are fused, and then the fusion algorithm is used to supplement the occluded area to generate a complete target view; Step 4: Automatically detect and identify the target object by combining the fused image data; Step 5: Verify the results of target detection and recognition and analyze the accuracy of target recognition.
2. The target recognition method from the perspective of a drone according to claim 1, characterized in that: In the step 1, the process of acquiring multi-view image data includes: Identify the target area to be photographed and the type of information to be obtained from different angles. Mount multiple cameras on the drone and adjust the angle and focal length of the cameras according to the shooting requirements. Plan the flight route, altitude and speed of the drone based on the terrain, weather conditions and shooting requirements of the target area; Operate the drone to shoot according to the set flight plan, control the drone to fly along the predetermined flight route and altitude, and start multiple cameras for synchronous shooting; The drone transmits the multi-view image data it captures to the ground control station in real time, which is then sorted out, photos that obviously do not meet the requirements are deleted, and the remaining images are pre-processed by denoising and color correction. An image database is established to record the shooting time and geographical location information of each picture, and the acquired multi-view image data is stored in the image database.
3. The target recognition method from the perspective of a drone according to claim 2, characterized in that: In the step 2, the multi-view image registration process includes: Select an image with the most central perspective from the multi-view images as the reference image, use the SIFT feature detection algorithm to detect key points in the reference image and each image to be registered, and generate a 128-dimensional descriptor vector for each key point; Use the BFMatcher matching algorithm to find the corresponding key point pairs between the reference image and the image to be registered, use the RANSAC algorithm to remove the wrong matching points, and estimate the geometric transformation model, and then use the geometric transformation model to align the image to be registered to the spatial coordinate system of the reference image; Extract image features for image fusion from the registered multi-view images, including color features, texture features, shape features, spatial relationship features, and multi-view consistency features. Input the registered images into the convolutional neural network model. After passing through multiple convolutional layers and pooling layers of the convolutional neural network model, a series of feature maps are obtained. For each image from each perspective, local feature vectors are extracted from the selected feature map, and the similarity between feature vectors from different perspectives is calculated using similarity measurement. Based on the similarity measurement results, feature matching relationships between images from different perspectives are established, and matching thresholds are set to determine which pairs are considered matching pairs. Pairs above the matching threshold are considered valid matching relationships. The feature matching results between each view image and the reference image are saved, including matching point pairs and their corresponding feature descriptors.
4. The target recognition method from the perspective of a drone according to claim 3, characterized in that: The calculation formula of the similarity metric is: ; In the formula, represents the similarity measure, represents the dimension of the feature vector, Indicates Perspective image The local eigenvector Dimensional component, Indicates The perspective image The local eigenvector Dimensional component, The value range is between 0 and 1. When two feature vectors are exactly the same, the similarity reaches the maximum value of 1.
5. The target recognition method from the perspective of a drone according to claim 4, characterized in that: In step 3, the process of generating the complete target view includes: According to the calculated similarity measurement results, confirm each pair of matching points and their corresponding feature descriptors, and organize all valid matching point pairs and their corresponding coordinate information; Create a blank target view image with a size of 16x16 pixel blocks. For each local area of the 16x16 pixel block, determine the corresponding position of the area in the images of different view angles based on the information of the matching point pairs, and use the weighted averaging method to perform weighted averaging on the pixel values in the overlapping area to fuse the image contents of multiple view angles into the target view. By analyzing the matching point pairs and feature descriptors, the occluded area is identified. For the detected occluded area, its size is m×n pixels. For each pixel position in the occluded area, the new pixel value after filling is calculated and filled using the image restoration algorithm based on content-aware filling; The filled image is fused with the original image, and the fused image is post-processed including denoising, sharpening, and color correction, and the processed image is used as the final output to obtain a complete target view.
6. The target recognition method from the perspective of a drone according to claim 5, characterized in that: The calculation expression of the new pixel value after filling is: ; ; ; In the formula, Represents the new pixel value after filling. represents the center point, Indicated in Radius around point The actual pixel value in the range, represents the weight function, represents the energy difference with the center pixel, and Relative to the center point The vertical and horizontal offsets, represents the difference between two pixels, is a smaller than An integer that limits the size of the local region to be compared.
7. The target recognition method from the perspective of a drone according to claim 6, characterized in that: In the step 4, the process of identifying the target object includes: Locate the position of the target object in the fused image, and detect feature points including texture features and shape features from the image, and then generate a descriptor for each detected feature point to describe the texture shape information of its local area; Use the sliding window method to traverse the entire image, use a fixed-size window to extract features at each position, match the extracted features with the features in the database, preset the matching threshold, and classify the content in the window to calculate the matching score to determine whether it is the target object; The detection results are post-processed, and the non-maximum suppression algorithm is applied to remove the target frames of repeated detections, and the positions of the identified target objects are marked on the image, and the positions and related information of the identified target objects are output.
8. The target recognition method from the perspective of a drone according to claim 7, characterized in that: The calculation expression of the matching score is: ; In the formula, represents the matching score, Indicates The matching score of feature points, Indicates the number of feature points extracted within the window, represents a small constant to avoid the denominator being zero. Represents the preset matching threshold, Indicates selection and The smaller one, The value range is between 0 and 1.
9. The target recognition method from the perspective of a drone according to claim 8, characterized in that: In step 5, the process of analyzing the target recognition accuracy includes: Obtain a set of real scene images containing target objects, and use bounding boxes to mark the location and size of the targets to accurately label the target objects in the images; Compare the target detection and recognition results with the target objects in the real scene image, analyze the color features, texture features, shape features and spatial relationship features of the two, clarify the matching degree of each feature, and analyze the accuracy of target recognition; Visualization tools are used to display the analysis results of target recognition accuracy, as well as the location and size of target objects.
10. A target recognition system from the perspective of a drone, used to implement the target recognition method from the perspective of a drone as described in any one of claims 1 to 9, comprising an image data management platform, characterized in that: The image data management platform is communicatively connected with an image acquisition module, a multi-view image registration and fusion module, a feature extraction and matching module, and a target detection and classification module, wherein electrical signals are connected between the modules; The image acquisition module uses a drone to capture multi-view images of the target object; The multi-view image registration and fusion module detects key points in the reference image and the image to be registered using a feature detection algorithm, and aligns the images to the same coordinate system, thereby fusing the image contents of multiple views into the target view and supplementing the occluded areas; The feature extraction and matching module is used to extract color features, texture features, shape features, spatial relationship features and multi-view consistency features from the registered images, generate feature maps through convolutional neural networks, calculate the similarity of feature vectors between different viewpoints, and establish feature matching relationships between images of different viewpoints; The target detection and classification module uses a fixed-size window to traverse the entire image, extract features and match them with features in a database, and classify the content in the window to determine whether it is a target object.
Citation Information
Patent Citations
A multi-platform multi-view target cooperative detection method
CN109949229A
Unmanned aerial vehicle image splicing method and system based on computer vision
CN114913068A
Multi-view target detection and identification method
CN118097275A
Unmanned aerial vehicle aerial image data acquisition and three-dimensional reconstruction method for single building
CN118823247A
Multi-view overlapped video fusion splicing method and system
CN119168857A
Cited By
Unmanned aerial vehicle intelligent identification method and system for shielding target
CN120259926A
Unmanned aerial vehicle dynamic detection system based on intelligent image recognition and efficient communication
CN120411834A
Cow oestrus monitoring method based on multi-camera data fusion
CN120656237A
A multi-camera data fusion-based cow estrus monitoring method
CN120656237B