A target recognition positioning method fusing navigation information

CN112232132BActive Publication Date: 2026-09-25BEIJING INST OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010988347.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-18
Publication Date
2026-09-25
Estimated Expiration
2040-09-18

AI Technical Summary

Technical Problem

[0004]但是,在实际检测过程中,训练阶段如果采用网络模型中原有的anchor设置,会导致检测速度和检测精度均不够好;如果根据训练数据集逐步减小anchor尺寸,运算量较大;通过随机实验确定训练阶段的anchor,会具有一定的偶然性和特殊性,导致检测算法泛化能力较低

Benefits of technology

[0040]本发明所具有的有益效果包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application discloses a target recognition positioning method fusing navigation information, which comprises the following steps: obtaining the maximum pixel of a target in a field of view by combining the height information given by an aircraft altimeter; and optimizing a global random anchor problem into a random anchor in a small scale range, so that the detection efficiency is improved; and the actual position of the target can be calculated by a navigation device. The target recognition positioning method fusing navigation information can realize rapid and accurate positioning of a ground target, and the detection efficiency is improved by 76.2%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and remote sensing target detection technology, and specifically to a target recognition and positioning method that integrates navigation information. Background Technology

[0002] In the field of remote sensing target detection for aircraft, with the increasing complexity of target detection scenarios, traditional image processing target detection algorithms are becoming increasingly computationally intensive and unable to meet the requirements. In recent years, machine learning has developed rapidly, and the combination of machine vision and machine learning is becoming a growing trend. Since the advent of convolutional neural networks (CNNs), the speed and accuracy of target detection have greatly improved. For example, deep learning-based algorithms like YOLO and SSD employ direct regression, using convolutional networks to directly output target location and category information, while also outputting target confidence scores. Their advantages include very fast detection speed and significantly improved accuracy. Deep learning-based target detection algorithms like R-CNN further improve detection accuracy, but due to the increased computational load, their real-time performance is not as good as YOLO and SSD algorithms. Currently, the YOLOv3 series of algorithms represents a compromise between detection speed and accuracy.

[0003] One important factor affecting the speed of convolutional neural network detection algorithms is the setting of anchor size in the convolutional network. Researchers need to set different anchor sizes for specific detection targets and detection performance, and the basis for setting anchor size varies, such as gradually reducing anchor size or determining the best anchor size through random experiments.

[0004] However, in actual detection, if the original anchor settings in the network model are used during the training phase, the detection speed and accuracy will be insufficient; if the anchor size is gradually reduced according to the training dataset, the computational load will be large; and if the anchors are determined through random experiments during the training phase, there will be a certain degree of randomness and particularity, resulting in low generalization ability of the detection algorithm.

[0005] Therefore, how to set the anchor and how to obtain the basis for setting the anchor are the key to optimizing the detection effect. There is an urgent need to provide a method to optimize the detection effect and improve the detection performance of the aircraft for remote sensing targets. Summary of the Invention

[0006] To overcome the above problems, the inventors conducted intensive research and designed a target recognition and positioning method that integrates navigation information. This method obtains the maximum number of pixels of the target in the field of view by combining the altitude information given by the aircraft altimeter, and optimizes the global random anchor problem into random anchors within a small scale, thereby improving the detection efficiency. At the same time, the actual position of the target can be calculated by the navigation device, which can realize the rapid and accurate positioning of ground targets, thus completing the present invention.

[0007] Specifically, the object of the present invention is to provide the following aspects:

[0008] Firstly, a target identification and positioning method integrating navigation information is provided, the method comprising the following steps:

[0009] Step 1: Train the object detection network;

[0010] Step 2: Obtain the image to be detected;

[0011] Step 3: Perform target detection on the image to be detected;

[0012] Step 4: Locate the target.

[0013] Step 1 includes the following sub-steps:

[0014] Step 1-1: Label the training dataset;

[0015] Steps 1-2: Construct the detection network;

[0016] Steps 1-3 involve modifying the anchor scale of the network;

[0017] Steps 1-4: Train the network until it converges.

[0018] In steps 1-3, the anchor scale of different feature layers of the network is modified according to the target scale range of the training set, preferably obtained by the following formula:

[0019]

[0020] Among them, S k S represents the ratio of the prior bounding box size of each feature layer to the original image size; max S represents the scale value of the highest-level feature layer. min This represents the scale value of the lowest feature layer, which is set according to the detection target; m represents the number of feature layers; k represents the nth feature layer.

[0021] Step 3 includes the following sub-steps:

[0022] Step 3-1: Extract features from the image to be detected to obtain multi-layer feature maps;

[0023] Step 3-2: Determine the anchor scale of the detection network;

[0024] Step 3-3: Based on the anchor scale determined above, generate multiple bounding boxes at each point in the feature map;

[0025] Steps 3-4: Obtain the category and coordinate information of the target to be detected.

[0026] Step 3-2 includes the following sub-steps:

[0027] Step 3-2-1: Determine the target imaging scale based on the camera height information and camera imaging rules;

[0028] Step 3-2-2: Determine the anchor scale in the feature maps of different layers based on the obtained target imaging scale.

[0029] In step 4, the position of the target relative to the aircraft is first obtained, and then the position of the aircraft itself is obtained according to the aircraft navigation equipment, thereby obtaining the actual position of the detected target.

[0030] Secondly, a target identification and positioning system integrating navigation information is provided, wherein the system includes an image acquisition unit, an aircraft information acquisition unit, a target detection unit, and a target positioning unit.

[0031] The image acquisition unit is used to obtain an image of the target to be detected.

[0032] The aircraft information acquisition unit is used to obtain the aircraft's altitude and position;

[0033] The target detection unit is used to obtain target category and coordinate information;

[0034] The target localization unit is used to obtain the actual position of the target.

[0035] The target detection unit includes a feature extraction subunit and an anchor scale setting subunit.

[0036] Among them, the feature extraction subunit is used to obtain multi-layer feature maps of the image to be detected;

[0037] The anchor scale setting sub-unit is used to determine the anchor scale in different layer feature maps based on the target imaging scale.

[0038] Thirdly, a computer-readable storage medium is provided, wherein a target identification and positioning program that integrates navigation information is stored, and when the program is executed by a processor, the processor performs the steps of the target identification and positioning method that integrates navigation information.

[0039] Fourthly, a computer device is provided, including a memory and a processor, wherein the memory stores a target identification and positioning program that integrates navigation information, and when the program is executed by the processor, the processor performs the steps of the target identification and positioning method that integrates navigation information.

[0040] The beneficial effects of this invention include:

[0041] (1) The target identification and positioning method that integrates navigation information provided by the present invention obtains the maximum pixel of the target in the field of view based on the altitude of the aircraft, optimizes the global random anchor problem into random anchors within a small scale, thereby improving the detection efficiency by 76.2%.

[0042] (2) The target identification and positioning method that integrates navigation information provided by the present invention can obtain the distance between the target and the camera based on the target’s line of sight angle and height information relative to the camera, and obtain the target’s precise position through the navigation device.

[0043] (3) The target identification and positioning method that integrates navigation information provided by the present invention can achieve rapid and accurate positioning of ground targets. Attached Figure Description

[0044] Figure 1 This diagram illustrates the window size of anchors in YOLOv3, where different colors represent different area sizes.

[0045] Figure 2 This diagram illustrates the labeled targets of a training dataset according to a preferred embodiment of the present invention.

[0046] Figure 3 A schematic diagram of camera imaging according to a preferred embodiment of the present invention is shown;

[0047] Figure 4 A schematic diagram of a camera coordinate system according to a preferred embodiment of the present invention is shown. Detailed Implementation

[0048] The present invention will be further described in detail below through preferred embodiments and examples. Through these descriptions, the features and advantages of the present invention will become clearer and more apparent.

[0049] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0050] The inventors discovered that the essence of anchor is the reverse of the SPP (spatial pyramid pooling) concept. SPP resizes inputs of different sizes into outputs of the same size. Therefore, the essence of anchor is to deduce inputs of different sizes from outputs of the same size.

[0051] For example, for the window size of anchors in YOLOv3, the three area dimensions are 128. 2 256 2 512 2 Then, for each area size, by taking three different aspect ratios (1:1, 1:2, 2:1), nine anchors with different area sizes can be obtained, such as... Figure 1 As shown.

[0052] Setting global anchors results in a large amount of computation, which affects detection efficiency.

[0053] Therefore, this invention employs a target detection algorithm based on convolutional neural networks, preferably optimizing the global random anchor problem into random anchors within a small scale to improve detection efficiency; at the same time, GPS location information is applied to the target positioning process to achieve accurate positioning of ground targets.

[0054] This invention provides a target identification and localization method that integrates navigation information. The method includes the steps of training a detection network and using the detection network for identification and localization.

[0055] Preferably, the method includes the following steps:

[0056] Step 1: Train the object detection network;

[0057] Step 2: Obtain the image to be detected;

[0058] Step 3: Perform target detection on the image to be detected;

[0059] Step 4: Locate the target.

[0060] The target identification and positioning method that integrates navigation information according to the present invention is further described below:

[0061] Step 1: Train the object detection network.

[0062] Step 1 includes the following sub-steps:

[0063] Step 1-1: Label the training dataset.

[0064] Deep learning-based target detection algorithms require different types of datasets for different application scenarios. In this invention, large-scale remote sensing images, such as VisDrone and DOTA datasets, are used as training datasets.

[0065] Specifically, when labeling a target, the bounding rectangle of the labeled target is as follows: Figure 2 As shown. Taking one of the annotation entries in the annotation file as an example, the content is as follows:

[0066] {"area":169,"bbox":[102,81,13,13],"category_name":"car"}

[0067] Here, the area value represents the pixel area of ​​the rectangular bounding box; the four values ​​after bbox: the first value represents the horizontal pixel coordinate of the top-left corner of the rectangle relative to the top-left corner of the image, positive to the right; the second value represents the vertical pixel coordinate of the top-left corner of the rectangle relative to the top-left corner of the image, positive downwards; the third value represents the width of the rectangle; the fourth value represents the height of the rectangle; and category_name represents the target category.

[0068] Steps 1-2: Construct the detection network.

[0069] This involves constructing a deep residual network with multiple convolutional layers, i.e., a deep convolutional network. In this invention, ResNet101 is preferably used as the base network to construct a new feature extraction network.

[0070] Steps 1-3 involve modifying the anchor scale of the network.

[0071] According to a preferred embodiment of the present invention, the anchor scale of different feature layers of the network is modified according to the target scale range of the training set, preferably obtained by the following formula:

[0072]

[0073] Among them, S k S represents the ratio of the prior bounding box size of each feature layer to the original image size; max S represents the scale value of the highest-level feature layer. min This represents the scale value of the lowest feature layer, which is set according to the detection target; m represents the number of feature layers; k represents the nth feature layer.

[0074] In this invention, a set of S is determined based on the target scale range of the training dataset. max and S min, which serves as the initial value for training the network.

[0075] Steps 1-4: Train the network until it converges.

[0076] Specifically, for the pre-trained deep residual network, the labeled training dataset is used to train the detection network according to the image category labels, and the network parameters are updated until a converged detection network is obtained.

[0077] Step 2: Obtain the image to be detected.

[0078] In this invention, the aircraft acquires an image of the target to be detected via a visual camera during flight. The aircraft can be an unmanned aerial vehicle (UAV), such as a drone, or a manned aircraft.

[0079] According to a preferred embodiment of the present invention, when the aircraft is flying, it is also necessary to obtain the aircraft's altitude and position;

[0080] Preferably, the altitude of the aircraft is obtained by an altimeter, and the position of the aircraft is obtained by a navigation device.

[0081] Step 3: Perform target detection on the image to be detected.

[0082] Step 3 includes the following sub-steps:

[0083] Step 3-1: Extract features from the image to be detected to obtain multi-layer feature maps.

[0084] In this process, the image to be detected is processed through the detection network (including the backbone network and the feature extraction network) trained in step 1 to obtain a multi-layer feature map.

[0085] Step 3-2: Determine the anchor scale of the detection network.

[0086] To address the issues of high computational load due to global anchor settings and low generalization ability caused by random experiments to determine anchor scale in existing technologies, the inventors have discovered that, based on the principle that the target has the largest image in the camera when it is located directly below the camera, the imaging pixel range can be obtained based on the actual size range of the target. This imaging pixel range can then be used as the maximum anchor size in the detection structure, significantly improving the detection speed.

[0087] Specifically, step 3-2 includes the following sub-steps:

[0088] Step 3-2-1: Determine the target imaging scale based on the camera height information and camera imaging rules.

[0089] Among these, an altimeter can be used to obtain the camera's altitude information; the camera's imaging rules are as follows: Figure 3 As shown.

[0090] Specifically: Suppose there are two points P1 and P2 on the target being detected, with coordinates P1 = [X1 Y1 Z1] in the camera coordinate system. T P2 = [X2 Y2 Z2] T ,

[0091] The coordinates of the two points in the image are p1 = [u1 v1]. T p2 = [u2 v2] T ;

[0092] The camera imaging model is shown below:

[0093]

[0094]

[0095] Among them, f x =αf,f y =βf, where f is the camera focal length in millimeters; α and β are the number of pixels per millimeter in pixels per millimeter.

[0096] The distance between the projections of points P1 and P2 onto the image is obtained by the following formula:

[0097] (pixels)

[0098] (pixels)

[0099] The target is at its largest size when it is directly below the camera. The target's imaging scale can be obtained based on the target's actual size range, the camera height obtained from the altimeter, and the imaging principle.

[0100] Step 3-2-2: Determine the anchor scale in the feature maps of different layers based on the obtained target imaging scale.

[0101] According to a preferred embodiment of the present invention, the anchor scale range of different layer feature maps is obtained based on the obtained target imaging scale using the following formula:

[0102]

[0103] Among them, S k S represents the ratio of the prior bounding box size of each feature layer to the original image size; max and S min These represent the scale values ​​of the top and bottom feature layers, respectively, and are set according to the detection target; m represents the number of feature layers; k represents the nth feature layer.

[0104] Update the S determined during training based on the obtained imaging scale of the target to be detected. max and S min This updates the anchor scale.

[0105] Step 3-3: Based on the anchor scale determined above, generate multiple bounding boxes at each point in the feature map.

[0106] In this invention, because different feature layers correspond to different receptive fields in the original image, the generated bounding boxes on different feature layers have different sizes. When generating bounding boxes, a series of concentric bounding boxes are generated at each point on the feature layer. m feature layers of different sizes are used for prediction, with the scale value of the lowest feature layer being S. min The scale value of the highest-level feature layer is S. max The other layers are obtained using the following formula:

[0107]

[0108] Using different aspect ratio values, i.e., γ (aspect ratio) = [1, 2, 3, 1 / 2, 1 / 3], the width of each anchor varies. high When γ = 1, increase the anchor size by one scale.

[0109] Steps 3-4: Obtain the category and coordinate information of the target to be detected.

[0110] In this invention, the optimal target box position is preferably output through Non-Maximum Suppression (NMS) to obtain the coordinate position of the target center in the field of view, namely the center x, y of the target box and the width and height w, h of the target box; at the same time, the target category information is output.

[0111] Step 4: Locate the target.

[0112] Step 4 includes the following sub-steps:

[0113] Step 4-1: Obtain the target's position relative to the aircraft.

[0114] In this invention, such as Figure 4 As shown, first define the camera coordinate system, the origin (x-axis), and the y-axis. Then, the pixel coordinates of the target in the camera coordinate system are (x0, y0), and the coordinates of the camera center point are (x0, y0). c ,y c ).

[0115] Preferably, ignoring the installation error between the center of the visual camera and the center of the aircraft, it can be approximated that (x c ,y c ) is the center of the drone.

[0116] According to a preferred embodiment of the present invention, the pixel coordinates of the target and the pixel coordinates of the spacecraft center are normalized, and the pixel error is obtained by the following formula:

[0117]

[0118]

[0119] In a further preferred embodiment, the pixel error in the camera coordinate system is converted into position errors in the x and y directions in the stable coordinate system of the aircraft using the following formula:

[0120]

[0121]

[0122] Among them, e x e y θ1 represents the target position deviation in the stable coordinate system; H represents the current altitude of the aircraft; θ1 and θ2 represent the field of view angles of the camera in the x and y directions, respectively, and are inherent parameters of the camera; K is an empirical coefficient, preferably...

[0123] The above steps can be used to obtain the target's directional position relative to the aircraft.

[0124] Step 4-2: Obtain the actual location of the target.

[0125] The process involves obtaining the aircraft's own position using its navigation equipment (such as GPS), calculating the actual position of the detected target, and thus completing the target positioning.

[0126] The target recognition and positioning method fused with navigation information provided by this invention calculates the maximum number of pixels of the target in the field of view by using the camera's height information obtained by the altimeter, thereby optimizing the global random anchor problem into random anchors within a small scale range, improving the detection efficiency; at the same time, the application of navigation equipment to the target positioning process enables accurate positioning of ground targets.

[0127] The present invention also provides a target identification and positioning system that integrates navigation information, the system comprising an image acquisition unit, an aircraft information acquisition unit, a target detection unit, and a target positioning unit.

[0128] The image acquisition unit is used to obtain an image of the target to be detected.

[0129] The aircraft information acquisition unit is used to obtain the aircraft's altitude and position;

[0130] The target detection unit is used to obtain target category and coordinate information;

[0131] The target localization unit is used to obtain the actual position of the target.

[0132] According to a preferred embodiment of the present invention, the target detection unit includes a feature extraction subunit and an anchor scale setting subunit.

[0133] Among them, the feature extraction subunit is used to obtain multi-layer feature maps of the image to be detected;

[0134] The anchor scale setting sub-unit is used to determine the anchor scale in different layer feature maps based on the target imaging scale.

[0135] The present invention also provides a computer-readable storage medium storing a target identification and positioning program that integrates navigation information. When the program is executed by a processor, the processor performs the steps of the target identification and positioning method that integrates navigation information.

[0136] The target identification and positioning method that integrates navigation information described in this invention can be implemented by means of software plus necessary general-purpose hardware platform. The software is stored in a computer-readable storage medium (including ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, network device, etc.) to execute the method described in this invention.

[0137] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a target identification and positioning program that integrates navigation information, and when the program is executed by the processor, the processor performs the steps of the target identification and positioning method that integrates navigation information.

[0138] Example

[0139] The present invention is further described below through specific examples; however, these examples are merely exemplary and do not constitute any limitation on the scope of protection of the present invention.

[0140] Example 1

[0141] 1. Dataset

[0142] The target recognition and localization method fused with navigation information described in this invention was evaluated using the VisDrone dataset. This dataset contains 263 video clips, 179,264 video frames, and 10,209 still images. For the target detection task, VisDrone contains 10,209 fully annotated still images that can be categorized into 10 classes. Of these, 6,471 images were used for training, 548 for validation, and 3,190 for testing. The image resolution is approximately 2000×1500 pixels.

[0143] 2. Task Description

[0144] Using the training dataset in the COCO dataset, the detection network is trained using the method described in this invention. After learning the network parameters and anchor scale modifications, the target detection and localization are achieved using the test dataset, UAV simulation, and GPS device. After testing, the performance is evaluated and compared with the method without altimeter fusion.

[0145] The experimental platform was an Nvidia TX2 computer.

[0146] 3. Results and Analysis

[0147] The comparative test results are shown in Table 1:

[0148] Table 1

[0149]

[0150]

[0151] FPS represents the detection speed, which is the number of images processed per second.

[0152] AP represents the detection accuracy, which is the area enclosed by the PR curve and the coordinate axis.

[0153] P represents precision, and R represents recall. The formulas for calculation are as follows:

[0154]

[0155]

[0156] TP means that the target is found as a positive example and is found correctly; FP means that a negative example is incorrectly detected as a positive example; TN means that a negative example is correctly found; and FN means that a positive example is incorrectly detected as a negative example.

[0157] The test results show that, compared with the method of determining the anchor scale without integrating altimeter information, the target identification and positioning method of the present invention has a slightly lower detection accuracy (AP) but a 76.2% increase in FPS, which significantly improves the detection speed while ensuring detection accuracy.

[0158] The present invention has been described in detail above with reference to specific embodiments and exemplary examples. However, these descriptions should not be construed as limiting the present invention. Those skilled in the art will understand that various equivalent substitutions, modifications, or improvements can be made to the technical solutions and embodiments of the present invention without departing from the spirit and scope of the present invention, and all such modifications and improvements fall within the scope of the present invention.

Claims

1. A target identification and positioning method integrating navigation information, characterized in that, The method includes the following steps: Step 1: Train the object detection network; Step 2: Obtain the image to be detected; Step 3: Perform target detection on the image to be detected; Step 4: Locate the target; Step 1 includes the following sub-steps: Step 1-1: Label the training dataset; Steps 1-2: Construct the detection network; Steps 1-3 involve modifying the anchor scale of the network; Steps 1-4: Train the network until it converges; In steps 1-3, the anchor scale of different feature layers of the network is modified according to the target scale range of the training set, obtained by the following formula: , in, This represents the ratio of the prior bounding box size of each feature layer to the original image size; This represents the scale value of the highest-level feature layer. This represents the scale value of the lowest feature layer, which is set according to the target being detected; m represents the number of feature layers; k represents the nth feature layer. Based on the target scale range of the training dataset, determine a set of... and , as the initial values ​​for training the network; Step 3 includes the following sub-steps: Step 3-1: Extract features from the image to be detected to obtain multi-layer feature maps; Step 3-2: Determine the anchor scale of the detection network; Step 3-3: Based on the anchor scale determined above, generate multiple bounding boxes at each point in the feature map; Steps 3-4: Obtain the category and coordinate information of the target to be detected; Step 3-2 includes the following sub-steps: Step 3-2-1: Determine the target imaging scale range based on camera height information and camera imaging rules; Step 3-2-2: Set according to the obtained target imaging scale range and Determine the anchor scale in the feature maps of different layers; In step 3-3, when generating the bounding boxes, a series of concentric bounding boxes are generated at each point on the feature layer. Prediction is performed using m feature layers of different sizes, with the scale value of the lowest feature layer being [value missing]. The scale value of the highest-level feature layer is To obtain the anchor scale in the feature maps of different layers, the target box uses different aspect ratios; In step 4, the position of the target relative to the aircraft is first obtained, and then the position of the aircraft itself is obtained according to the aircraft's navigation equipment, thereby obtaining the actual position of the detected target.

2. A computer-readable storage medium, characterized in that, A target identification and positioning program storing fused navigation information, when executed by a processor, causes the processor to perform the steps of the target identification and positioning method with fused navigation information as described in claim 1.

Citation Information

Patent Citations

  • Target locating method of miniature drone full-strapdown down looking camera

    CN107727079A

  • Unmanned-aerial-vehicle low-altitude-target accurate detection identification method

    CN108681718A

  • A target detection method in a vehicle-mounted environment

    CN109740463A

  • Ground target geographic coordinate positioning method based on unmanned aerial vehicle visual system

    CN111178148A