Object Detection Method, Device, Storage Medium and UAV

By segmenting and resizing images to match the resolution of multiple analysis units on drones, the method enhances target detection accuracy and success rate by minimizing information loss and enabling parallel processing.

CN110796104BActive Publication Date: 2025-07-15AUTEL ROBOTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911060737.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-01
Publication Date
2025-07-15
Estimated Expiration
2039-11-01

AI Technical Summary

Technical Problem

In the existing drone object detection scheme, due to the low resolution of the input image of the image processing chip, the high-resolution image captured by the gimbal camera loses some information during the size change, making it difficult to detect small targets, especially at high flight altitudes.

Method used

At least two image analysis units are integrated in the drone, and size changes are made by segmenting the images captured by the gimbal camera, so that the resolution of the sub-image matches the image analysis unit, and parallel analysis is performed to reduce information loss and improve detection accuracy.

Benefits of technology

Through segmentation and size change operations, image information loss is reduced, and the accuracy and success rate of object detection is improved, especially the detection effect of small objects is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110796104B_ABST
    Figure CN110796104B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention disclose a target detection method, apparatus, storage medium, and unmanned aerial vehicle. The method includes: obtaining a first image captured by a pan-tilt camera, performing segmentation processing on the first image to obtain at least two segmented images, performing a size change operation on the at least two segmented images to obtain at least two sub-images, where the resolutions of the at least two sub-images match the resolutions corresponding to at least two image analysis units, inputting the at least two sub-images into the at least two image analysis units, and determining a target detection result according to the analysis results of the at least two image analysis units. By adopting the above technical solution, the present invention can reduce the loss amount of image information and improve the accuracy and success rate of target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of unmanned aerial vehicles, and in particular, to a target detection method, apparatus, storage medium, and unmanned aerial vehicle. Background Art

[0002] Unmanned Aerial Vehicles (UAVs) refer to unpiloted aircrafts that are operated through radio remote control devices and independent program control devices, or are completely or intermittently autonomously operated by on-board computers. Compared with piloted aircrafts, UAVs have the advantages of small size, low cost, low environmental requirements, and strong survivability, and are often more suitable for dangerous or harsh environment tasks. With the rapid development of the UAV manufacturing industry, UAV systems are widely used in fields such as smart city management and intelligent traffic monitoring. Among them, target detection is a basic but challenging functional requirement in UAV systems, and is closely related to applications such as infrastructure inspection, urban perception, map reconstruction, and traffic control. These applications have promoted the development of UAV-based online monitoring systems, which can perform various tasks, such as inspection of on-site facilities and detection of violations, identification of unhealthy crops, and acquisition of map data.

[0003] Generally, an image processing chip is integrated on a UAV. An image analysis unit in the image processing chip analyzes and processes the images captured by a gimbal camera on the UAV to achieve target detection. Currently, in order to ensure the real-time performance of detection, the resolution of the input images of the image processing chip is generally low, while the resolution of the gimbal camera is very high. It is necessary to perform a size change (resize) operation on the images captured by the gimbal camera and then hand them over to the image processing chip for processing. This causes some image information to be lost during the image processing process, making the detection of some small targets very difficult or even impossible. Therefore, the existing UAV target detection solutions need to be improved. Summary of the Invention

[0004] Embodiments of the present invention provide a target detection method, apparatus, storage medium, and device, which can optimize the existing target detection solutions.

[0005] In a first aspect, an embodiment of the present invention provides a target detection method, which is applied to a UAV. At least two image analysis units are integrated in the UAV. The method includes:

[0006] Obtain a first image captured by a gimbal camera;

[0007] Perform segmentation processing on the first image to obtain at least two segmented images;

[0008] Perform a size change operation on the at least two segmented images to obtain at least two sub-images, where the resolution of the at least two sub-images matches the resolution corresponding to the at least two image analysis units;

[0009] Input the at least two sub-images into the at least two image analysis units, and determine the target detection result according to the analysis results of the at least two image analysis units.

[0010] In a second aspect, an embodiment of the present invention provides a target detection device, which is applied to a drone. At least two image analysis units are integrated in the drone. The device includes:

[0011] An image acquisition module, configured to acquire a first image captured by a pan-tilt camera;

[0012] An image segmentation module, configured to perform segmentation processing on the first image to obtain at least two segmented images;

[0013] A size change module, configured to perform a size change operation on the at least two segmented images to obtain at least two sub-images, where the resolution of the at least two sub-images matches the resolution corresponding to the at least two image analysis units;

[0014] A target detection module, configured to input the at least two sub-images into the at least two image analysis units, and determine the target detection result according to the analysis results of the at least two image analysis units.

[0015] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the target detection method provided by the embodiment of the present invention.

[0016] In a fourth aspect, an embodiment of the present invention provides a drone, including a memory, at least two image analysis units, a processor, and a computer program stored on the memory and executable on the processor. The processor implements the target detection method provided by the embodiment of the present invention when executing the computer program.

[0017] The object detection solution provided in the embodiments of the present invention is applied to a drone. At least two image analysis units are integrated in the drone. The first image captured by the pan-tilt camera is obtained, segmented, and the size of at least two segmented images is changed. The resolutions of at least two sub-images obtained are matched with the resolutions corresponding to at least two image analysis units. At least two sub-images are input into at least two image analysis units, and the object detection result is determined according to the analysis results of at least two image analysis units. By adopting the above technical solution, after the original image captured by the pan-tilt camera is segmented and the size is changed, the loss of image information can be reduced. Subsequently, it is analyzed by at least two image analysis units, which can improve the accuracy and success rate of object detection. Description of the Drawings

[0018] Figure 1 It is a schematic flowchart of an object detection method provided in Embodiment 1 of the present invention;

[0019] Figure 2 It is a schematic diagram of an image captured by a pan-tilt camera provided in Embodiment 1 of the present invention;

[0020] Figure 3 It is a schematic diagram of an image after size change processing provided in Embodiment 1 of the present invention;

[0021] Figure 4 It is a schematic diagram of the left image after segmentation provided in Embodiment 1 of the present invention;

[0022] Figure 5 It is a schematic diagram of the right image after segmentation provided in Embodiment 1 of the present invention;

[0023] Figure 6 It is a schematic diagram of the left image after size change provided in Embodiment 1 of the present invention;

[0024] Figure 7 It is a schematic diagram of the right image after size change provided in Embodiment 1 of the present invention;

[0025] Figure 8 It is a schematic flowchart of an object detection method provided in Embodiment 2 of the present invention;

[0026] Figure 9 It is a schematic flowchart of an object detection method provided in Embodiment 3 of the present invention;

[0027] Figure 10 It is a schematic diagram of the left image including position information provided in Embodiment 3 of the present invention;

[0028] Figure 11 It is a schematic diagram of the right image including position information provided in Embodiment 3 of the present invention;

[0029] Figure 12 A schematic diagram of analysis result fusion provided in Embodiment III of the present invention;

[0030] Figure 13 A structural block diagram of a target detection device provided in Embodiment IV of the present invention;

[0031] Figure 14 A structural block diagram of a drone provided in Embodiment VI of the present invention. Detailed implementation manners

[0032] The technical solution of the present invention will be further described below in conjunction with the drawings and through specific implementation manners. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings rather than all structures.

[0033] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0034] Embodiment I

[0035] Figure 1 A schematic flowchart of a target detection method provided in Embodiment I of the present invention. This method can be executed by a target detection device, which can be implemented by software and / or hardware and is generally integrated in a drone. As Figure 1 shown, this method includes:

[0036] Step 101, obtain a first image captured by a pan-tilt camera.

[0037] In the embodiment of the present invention, the pan-tilt camera can be integrated inside the drone or externally mounted on the drone, and is connected to the drone by wired or wireless means. The pan-tilt camera can perform real-time image acquisition during the flight of the drone and can obtain the images captured by the pan-tilt camera in real time or at a preset frequency. The first image can be an image captured at any time, and the embodiment of the present invention does not make any limitation.

[0038] In an embodiment of the present invention, at least two image analysis units are integrated in the drone. The specific type of the image analysis unit is not limited. For example, it can be a forward inference engine (Neural Network Inference Engine, NNIE) that is accelerated based on a neural network. The image analysis unit can be integrated in an image processing chip, such as the HI3559C chip, which carries two forward inference engines specifically for accelerating neural networks. The two forward inference engines can independently process tasks such as detection and classification.

[0039] Step 102: Perform segmentation processing on the first image to obtain at least two segmented images.

[0040] With the rapid development of camera technology, the resolution of current gimbal cameras can reach a very high level, such as 8K, and the ratio of the captured images is generally 4:3 or 16:9. To ensure the real-time performance of detection, the resolution of the input image of the image analysis unit is generally relatively low, such as 512*512. In this way, when performing a resize operation on the image captured by the gimbal camera, a large amount of image information will be lost, resulting in a decrease in the accuracy of target detection. It is particularly difficult to detect small targets. When the drone flies at a high altitude, the targets on the ground are smaller in the image and are even more difficult to detect.

[0041] Figure 2 It is a schematic diagram of an image captured by a gimbal camera provided in Embodiment 1 of the present invention. The ratio of this image is 16:9, and the resolution is 1920*1080. Figure 3 It is a schematic diagram of an image after size change processing provided in Embodiment 1 of the present invention. After resize, Figure 2 the image in Figure 3 becomes a 1:1 image in Figure 2 with a resolution of 512*512. In this way,

[0042] In this step, the first image can be segmented according to a preset rule to obtain at least two segmented images. Among them, the preset rule can include the specific number of segments and the segmentation method, etc., such as average segmentation or segmentation according to a ratio, etc., and also such as left-right structure segmentation, up-down structure segmentation, grid segmentation, and nine-grid segmentation, etc. The specific number of segments and the segmentation method can be determined according to the number of image analysis units and other reference factors, and are not limited in the embodiments of the present invention.

[0043] In the embodiments of the present invention, in order to strengthen the comparison effect with the prior art, still taking the image in Figure 2 as an example. Figure 4 It is a schematic diagram of the left image after segmentation provided in Embodiment 1 of the present invention. Figure 5Schematic diagram of a right - hand image after segmentation provided in Embodiment 1 of the present invention. Refer to Figure 4 and Figure 5 , the image in Figure 2 can be evenly segmented into left - and - right structures to obtain two segmented images with an image ratio of 8:9.

[0044] Step 103: Perform a size - change operation on the at least two segmented images to obtain at least two sub - images, and the resolution of the at least two sub - images matches the resolution corresponding to the at least two image analysis units.

[0045] In the embodiment of the present invention, the specific manner of performing the size - change operation on the segmented image is not limited. The purpose of this size - change operation is to make the resolution of the at least two sub - images match the resolution corresponding to the at least two image analysis units. Exemplarily, generally, the resolutions corresponding to the at least two image analysis units are equal. In this way, the resolutions of the at least two sub - images are the same as this resolution.

[0046] Exemplarily, the size - change operation may include a size - change operation using an interpolation algorithm. The interpolation algorithm can be, for example, the nearest - neighbor method, the bilinear method, the bicubic method, an algorithm based on pixel - region relationships, and the Lanczos interpolation method, etc.

[0047] Figure 6 Schematic diagram of a left - hand image after size - change provided in Embodiment 1 of the present invention. Figure 7 Schematic diagram of a right - hand image after size - change provided in Embodiment 1 of the present invention. Refer to Figure 6 and Figure 7 , perform size - change operations on the images in Figure 5 and Figure 6 respectively to obtain two images with an image ratio of 1:1 and a resolution of 512*512.

[0048] Step 104: Input the at least two sub - images into the at least two image analysis units, and determine the target detection result according to the analysis results of the at least two image analysis units.

[0049] By adopting the solution of the embodiment of the present invention, the total area of the image input to the image analysis unit becomes larger. Referring to the above example, it is equivalent to doubling the target in the image. If the number of segmented images is more, the magnification factor is also larger, reducing the difficulty of target detection and effectively improving the accuracy and success rate of target detection.

[0050] The working timing of the at least two image analysis units is not specifically limited in the embodiments of the present invention. Optionally, after inputting the at least two sub-images into the at least two image analysis units, the method further includes: controlling the at least two image analysis units to analyze and process the received sub-images in parallel. The advantage of such a setting is that while ensuring the real-time detection, the accuracy and success rate of target detection are improved.

[0051] Exemplarily, the content included in the analysis result of the image analysis unit is related to the specific type, model, and function of the image analysis unit, etc., which are not limited in the embodiments of the present invention. Generally, the analysis result may include the position information and type information of the analyzed target, etc. The target can be a target object or a target person, etc., which can be a pre-specified target or an automatically recognized target, and can be set according to actual needs. The analysis results of the at least two image analysis units can be integrated to determine the final target detection result.

[0052] The target detection method provided in the embodiments of the present invention is applied to a drone. At least two image analysis units are integrated in the drone. The drone acquires a first image captured by a gimbal camera, performs segmentation processing on the first image, performs a size change operation on at least two segmented images, and the resolutions of the at least two obtained sub-images match the resolutions corresponding to the at least two image analysis units. The at least two sub-images are input into the at least two image analysis units, and the target detection result is determined according to the analysis results of the at least two image analysis units. By adopting the above technical solution, after the original image captured by the gimbal camera is segmented and then the size change operation is performed, the loss amount of image information can be reduced, and then it is analyzed by at least two image analysis units, which can improve the accuracy and success rate of target detection.

[0053] Embodiment 2

[0054] Figure 8 FIG. 13 is a schematic flowchart of a target detection method provided in Embodiment 2 of the present invention. This method is optimized based on the above embodiment, and the specific process of determining the target detection result according to the analysis results of the at least two image analysis units is refined.

[0055] Exemplarily, determining the target detection result according to the analysis results of the at least two image analysis units includes: acquiring the analysis results of the at least two image analysis units; performing fusion processing on at least two analysis results to obtain the target detection result. The advantage of such a setting is that during the segmentation processing, it is possible that the target is at the segmentation position, resulting in the target being detected in adjacent images. In view of this situation, the analysis results can be fused to fuse the analysis results of the segmented target.

[0056] Further, the analysis results include the type information and position information of the analyzed targets. The fusion process for at least two analysis results includes: sequentially denoting each two adjacent sub-images as a current sub-image pair, where the current sub-image pair includes a first sub-image and a second sub-image, and performing the following operations on the current sub-image pair: determining a first target in the first sub-image and a second target in the second sub-image, and determining whether the first target and the second target correspond to the same target according to the first position information and first type information corresponding to the first target, and the second position information and second type information corresponding to the second target. If so, fusing the first target and the second target into the same target. The advantage of this setting is that it can determine whether there are images corresponding to the targets segmented in each two adjacent sub-images one by one according to the position information and type information, and can accurately identify the targets segmented in the image.

[0057] Specifically, the method includes the following steps:

[0058] Step 201, obtain a first image captured by a pan-tilt camera.

[0059] Step 202, perform segmentation processing on the first image to obtain at least two segmented images.

[0060] Step 203, perform a size change operation on at least two segmented images to obtain at least two sub-images, and the resolutions of the at least two sub-images match the resolutions corresponding to at least two image analysis units.

[0061] Step 204, input at least two sub-images into the at least two image analysis units, and control the at least two image analysis units to perform analysis processing on the received sub-images in parallel.

[0062] Step 205, obtain the analysis results of the at least two image analysis units.

[0063] Step 206, sequentially denote each two adjacent sub-images as a current sub-image pair, and perform the following operations on the current sub-image pair: determine a first target in the first sub-image and a second target in the second sub-image, and determine whether the first target and the second target correspond to the same target according to the first position information and first type information corresponding to the first target, and the second position information and second type information corresponding to the second target. If so, fusing the first target and the second target into the same target.

[0064] Among them, the analysis results include the type information and position information of the analyzed targets, and the current sub-image pair includes a first sub-image and a second sub-image. The first target can be any target analyzed in the first sub-image, and the second target can be any target analyzed in the second sub-image.

[0065] Optionally, to reduce the computational amount, the first target and the second target can be preliminarily screened, and the target closer to the segmentation boundary is determined as the first target or the second target. Exemplarily, taking the first target as an example, the boundary coinciding with the second sub-image in the first sub-image is denoted as the segmentation boundary, the alternative targets analyzed in the first sub-image are obtained, and for each alternative target, the fifth distance between the current alternative target and the segmentation boundary is judged. If the fifth distance is less than the third preset threshold, the current alternative target is determined as the first target. Similarly, the second target can be determined with reference to the above content. The third preset threshold can be set according to the size of the sub-image. For example, it can be a preset ratio of the side length perpendicular to the segmentation boundary, and the preset ratio can be 10% for example.

[0066] Exemplarily, the position information in the analysis result can include a coordinate range, and the coordinate range can form a certain shape, such as a circle, an ellipse or a rectangle, etc., or a shape matching the target appearance.

[0067] Optionally, the position information includes the coordinates of a rectangular frame, and the image corresponding to the target analyzed is included within the rectangular frame. The rectangular frame corresponding to the first target is denoted as the first rectangular frame, and the rectangular frame corresponding to the second target is denoted as the second rectangular frame. The type information can be determined by the specific capabilities of the image analysis unit. For example, it can be analyzed whether it is a moving object or a still object, an animal or a person, and the specific category can also be analyzed, such as a vehicle, a house, etc., and more detailed categories can also be analyzed, such as a sedan, a bus, a fire truck, an ambulance, and so on.

[0068] Exemplarily, the same rule is adopted for numbering the four sides of the first rectangle and the second rectangle, and then based on the distance between the sides and the type information corresponding to the two targets, it is judged whether the two rectangles correspond to the same target. For the case of adjacent on the left and right, it can be judged whether the distance between the two rectangular frames in the horizontal direction is close enough and whether the degree of deviation from the same horizontal line in the vertical direction is small enough; for the case of adjacent on the top and bottom, it can be judged whether the distance between the two rectangular frames in the vertical direction is close enough and whether the degree of deviation from the same vertical line in the horizontal direction is small enough. If the above conditions are met and the type information of the first target and the second target is the same, the first target and the second target can be considered as the same target.

[0069] Specifically, determining whether the first target and the second target are the same target according to the first position information and the first type information corresponding to the first target, and the second position information and the second type information corresponding to the second target may include:

[0070] Calculate the first distance between the first boundary of the first rectangular box and the third boundary of the second rectangular box, the second distance between the third boundary of the first rectangular box and the first boundary of the second rectangular box, the third distance between the second boundary of the first rectangular box and the fourth boundary of the second rectangular box, and the fourth distance between the fourth boundary of the first rectangular box and the second boundary of the second rectangular box, based on the coordinates of the first rectangular box and the coordinates of the second rectangular box. Herein, the first boundary and the third boundary in each rectangular box are parallel, and the second boundary and the fourth boundary are parallel. When the first sub-image and the second sub-image are adjacent left and right, the first boundary in each rectangular box is the left boundary; when the first sub-image and the second sub-image are adjacent up and down, the first boundary in each rectangle is the upper boundary;

[0071] Calculate the first ratio of the smaller value to the larger value among the first distance and the second distance; and calculate the second ratio of the smaller value to the larger value among the third distance and the fourth distance;

[0072] When the first ratio is less than the first preset threshold, the second ratio is greater than the second preset threshold, and the first type information and the second type information are the same, determine whether the first target and the second target correspond to the same target, where the first preset threshold is less than the second preset threshold.

[0073] Herein, the first preset threshold and the second preset threshold can be set according to actual requirements. For example, the first preset threshold is 0.1 and the second preset threshold is 0.6.

[0074] Optionally, the step of fusing the first target and the second target into the same target includes: determining a target rectangular box according to the coordinates of the first rectangular box and the coordinates of the second rectangular box, where the target rectangular box contains both the first rectangular box and the second rectangular box; determining the target rectangular box and the first type information as the analysis result corresponding to the fused target. The advantage of such a setting is that when it is determined that the first target and the second target are the same target, the target rectangular box that contains both the first rectangular box and the second rectangular box can be determined as the final position information of the target, avoiding errors in the number of targets in the final target detection result.

[0075] Step 207: Determine the result after the fusion process as the target detection result.

[0076] The object detection method provided by the embodiment of the present invention, after image segmentation and size change, is handed over to at least two image analysis units for parallel processing. After obtaining the preliminary analysis results of each image analysis unit, fully considering the situation where the same object is segmented into two sub-images, the analysis results of the judged segmented objects are fused, and finally an accurate object detection result is obtained.

[0077] Based on the above embodiment, after determining the object detection result, it may further include stitching at least two sub-images containing the object detection result and outputting them to the user device. The user device may be a mobile terminal or a computer and other devices for the user to view the object detection result.

[0078] Embodiment III

[0079] Figure 9 It is a flowchart of an object detection method provided by Embodiment III of the present invention. This method takes the example of two image analysis units integrated in a drone to perform left-right average segmentation on an image. Specifically, the method includes the following steps:

[0080] Step 301, obtain the first image captured by the gimbal camera.

[0081] For the convenience of description, the image in Figure 2 is taken as the first image for the following description. The image ratio is 16:9, and the resolution is 1920*1080.

[0082] Step 302, perform left-right average segmentation processing on the first image to obtain two segmented images.

[0083] Perform left-right structure average segmentation on the image in Figure 2 to obtain two segmented images with an image ratio of 8:9, as shown in Figure 4 and Figure 5 shown.

[0084] Step 303, perform size change operation on the two segmented images to obtain two sub-images, and the resolutions of the two sub-images match the resolutions corresponding to the two image analysis units.

[0085] Exemplarily, the two image analysis units are two forward inference engines dedicated to accelerating neural networks carried on the HI3559C chip, and the resolution supported by each inference engine is 512*512. Referring to Figure 6 and Figure 7 , respectively perform size change operations on the images in Figure 5 and Figure 6 to obtain two images with an image ratio of 1:1 and a resolution of 512*512.

[0086] Step 304: Input the two sub-images into two image analysis units, and control the two image analysis units to analyze and process the received sub-images in parallel.

[0087] Step 305: Obtain the analysis results of the two image analysis units.

[0088] Exemplarily, the two inferencers will respectively give the position information (BBox) and class (Class) information of the target obtained by analysis in the respective sub-images. Among them, the position information is represented by the coordinates of a rectangular box. When a target coexists in the left and right sub-images, such as the car in the figure, which appears in both Figure 6 and Figure 7 , then the two inferencers will also respectively give the position information of the car in the sub-images they are responsible for. Figure 10 FIG. Figure 11 is a schematic diagram of the left image containing position information provided by Embodiment 3 of the present invention. Figure 10 and Figure 11 FIG.

[0089] Step 306: Determine the first target in the first sub-image and the second target in the second sub-image. According to the first position information and first type information corresponding to the first target, and the second position information and second type information corresponding to the second target, determine whether the first target and the second target correspond to the same target. If so, fuse the first target and the second target into the same target.

[0090] Figure 12 FIG.

[0091] is a schematic diagram of analysis result fusion provided by Embodiment 3 of the present invention. As shown in the figure, the left rectangular box Left BBox represents the first rectangular box of the first target, and the right rectangular box Right BBox represents the second rectangular box of the second target.

[0092] W = Rrigth – Lleft w = Rleft – Lrigth

[0093] H = Rbottom – Ltop h = Lbottom - Rtop

[0094] When w / W < 0.1 and h / H > 0.6, and the class information is the same, it can be considered that the two BBoxes on the left and right are the same target. At this time, the fused BBox is (Ltop, Rbottom, Lleft, Rrigth). Among them, 0.1 is the first preset threshold, and 0.6 is the second preset threshold.

[0095] The above is just one case. Due to the possible differences in the size and relative position relationship of the rectangular boxes, a general formula can be expressed as:

[0096] W = Rrigth – Lleft w = Rleft – Lrigth

[0097] H = max(Rbottom, Lbottom) – min(Ltop, Rtop)

[0098] h = min(Lbottom, Rbottom) – max(Rtop, Ltop)

[0099] When w / W < 0.1 and h / H > 0.6, and the class information is the same, it can be considered that the two BBoxes on the left and right are the same target. At this time, the fused BBox is (min(Ltop, Rtop), max(Lbottom, Rbottom), Lleft, Rrigth).

[0100] Step 307: Determine the result after the fusion process as the target detection result.

[0101] The target detection method provided by the embodiments of the present invention performs an average segmentation of the left and right structures on the original image collected by the pan-tilt camera, and performs a size change operation, and hands it over to two inference engines for parallel processing. After obtaining the preliminary analysis results of the two inference engines, it fully considers the situation where the same target is segmented into two sub-images, and fuses the analysis results of the segmented targets judged, and finally obtains an accurate target detection result. It can be seen from the above example that the detection resolution is expanded from 512*512 to 1024*512, which can improve the detection success rate of small targets. The two inference engines perform parallel operations, and the time consumption is the same as that of a single inference engine at 512*512. Resizing an 8:9 image to 1:1 retains more information in the image compared to resizing a 16:9 image to 1:1, further improving the accuracy and success rate of target detection.

[0102] Example 4

[0103] Figure 13The block diagram of a target detection device provided in the fourth embodiment of the present invention. The device can be implemented by software and / or hardware, and is generally integrated in a drone. It can perform target detection by executing a target detection method. Among them, at least two image analysis units are integrated in the drone. As Figure 13 shown, the device includes:

[0104] An image acquisition module 401, configured to acquire a first image captured by a pan-tilt camera;

[0105] An image segmentation module 402, configured to perform segmentation processing on the first image to obtain at least two segmented images;

[0106] A size change module 403, configured to perform a size change operation on the at least two segmented images to obtain at least two sub-images, and the resolution of the at least two sub-images matches the resolution corresponding to the at least two image analysis units;

[0107] A target detection module 404, configured to input the at least two sub-images into the at least two image analysis units, and determine a target detection result according to the analysis results of the at least two image analysis units.

[0108] The target detection device provided in the embodiment of the present invention is applied to a drone. At least two image analysis units are integrated in the drone. A first image captured by a pan-tilt camera is acquired, the first image is segmented, a size change operation is performed on at least two segmented images, the resolution of the at least two obtained sub-images matches the resolution corresponding to the at least two image analysis units, the at least two sub-images are input into the at least two image analysis units, and a target detection result is determined according to the analysis results of the at least two image analysis units. By adopting the above technical solution, after the original image captured by the pan-tilt camera is segmented, when performing a size change operation, the loss of image information can be reduced, and then it is handed over to at least two image analysis units for analysis, which can improve the accuracy and success rate of target detection.

[0109] Optionally, determining the target detection result according to the analysis results of the at least two image analysis units includes:

[0110] Acquiring the analysis results of the at least two image analysis units;

[0111] Performing fusion processing on at least two analysis results to obtain a target detection result.

[0112] Optionally, the analysis results include the type information and position information of the analyzed target, and the performing fusion processing on at least two analysis results includes:

[0113] Successively record every two adjacent sub-images as the current sub-image pair. The current sub-image pair includes a first sub-image and a second sub-image. Perform the following operations on the current sub-image pair:

[0114] Determine a first target in the first sub-image and a second target in the second sub-image;

[0115] According to the first position information and first type information corresponding to the first target, and the second position information and second type information corresponding to the second target, determine whether the first target and the second target correspond to the same target. If so, fuse the first target and the second target into the same target.

[0116] Optionally, the position information includes the coordinates of a rectangular box, and the image corresponding to the analyzed target is contained within the rectangular box. The rectangular box corresponding to the first target is denoted as the first rectangular box, and the rectangular box corresponding to the second target is denoted as the second rectangular box;

[0117] The determining whether the first target and the second target are the same target according to the first position information and first type information corresponding to the first target, and the second position information and second type information corresponding to the second target includes:

[0118] According to the coordinates of the first rectangular box and the coordinates of the second rectangular box, calculate the first distance between the first boundary of the first rectangular box and the third boundary of the second rectangular box, the second distance between the third boundary of the first rectangular box and the first boundary of the second rectangular box, the third distance between the second boundary of the first rectangular box and the fourth boundary of the second rectangular box, and the fourth distance between the fourth boundary of the first rectangular box and the second boundary of the second rectangular box. Among them, the first boundary and the third boundary in each rectangular box are parallel, and the second boundary and the fourth boundary are parallel. When the first sub-image and the second sub-image are adjacent left and right, the first boundary in each rectangular box is the left boundary. When the first sub-image and the second sub-image are adjacent up and down, the first boundary in each rectangular box is the upper boundary;

[0119] Calculate the first ratio of the smaller distance to the larger distance among the first distance and the second distance; and calculate the second ratio of the smaller distance to the larger distance among the third distance and the fourth distance;

[0120] When the first ratio is less than a first preset threshold, the second ratio is greater than a second preset threshold, and the first type information and the second type information are the same, determine whether the first target and the second target are the same target, where the first preset threshold is less than the second preset threshold.

[0121] Optionally, the fusion of the first target and the second target into the same target includes:

[0122] Determine a target rectangular box according to the coordinates of the first rectangular box and the coordinates of the second rectangle, where the target rectangular box includes both the first rectangular box and the second rectangular box;

[0123] Determine the analysis result corresponding to the fused target by using the target rectangular box and the first type of information.

[0124] Optionally, denote the boundary where the first sub-image coincides with the second sub-image as the segmentation boundary. The determination of the first target in the first sub-image includes:

[0125] Obtain alternative targets analyzed from the first sub-image;

[0126] For each alternative target, determine the fifth distance between the current alternative target and the segmentation boundary. If the fifth distance is less than the third preset threshold, determine the current alternative target as the first target.

[0127] Optionally, the at least two image analysis units include at least two forward inference engines NNIEs accelerated based on neural networks.

[0128] Embodiment 5

[0129] The embodiment of the present invention further provides a storage medium including computer-executable instructions, where the computer-executable instructions are used to execute a target detection method when executed by a computer processor. The method includes:

[0130] Obtain a first image captured by a pan-tilt camera;

[0131] Perform segmentation processing on the first image to obtain at least two segmented images;

[0132] Perform a size change operation on the at least two segmented images to obtain at least two sub-images, where the resolutions of the at least two sub-images match the resolutions corresponding to the at least two image analysis units;

[0133] Input the at least two sub-images into the at least two image analysis units, and determine the target detection result according to the analysis results of the at least two image analysis units.

[0134] Storage medium - Any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media such as CD-ROMs, floppy disks or magnetic tape devices; computer system memory or random access memory such as DRAM, DDRRAM, SRAM, EDORAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. The storage medium may also include other types of memory or combinations thereof. Additionally, the storage medium may be located in a first computer system in which the program is executed, or may be located in a different second computer system that is connected to the first computer system via a network (such as the Internet). The second computer system may provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media that may reside in different locations (such as in different computer systems connected via a network). The storage medium may store program instructions (such as embodied as a computer program) executable by one or more processors.

[0135] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present invention, the computer-executable instructions are not limited to the object detection operations described above, and may also execute related operations in the object detection method provided by any embodiment of the present invention.

[0136] Embodiment Six

[0137] An embodiment of the present invention provides a drone in which the object detection device provided by the embodiment of the present invention can be integrated. Figure 14 The block diagram of a drone provided for Embodiment Six of the present invention. The drone 500 may include: a memory 501, a processor 502, and at least two image analysis units 503 (only one is shown in the figure). A computer program stored on the memory 501 and executable by the processor, when the processor 502 executes the computer program, implements the object detection method as described in the embodiment of the present invention. The method may include:

[0138] Obtain a first image captured by a pan-tilt camera;

[0139] Perform segmentation processing on the first image to obtain at least two segmented images;

[0140] Perform a size change operation on the at least two segmented images to obtain at least two sub-images, and the resolution of the at least two sub-images matches the resolution corresponding to the at least two image analysis units;

[0141] Input the at least two sub-images into the at least two image analysis units, and determine the target detection result according to the analysis results of the at least two image analysis units.

[0142] The computer device provided by the embodiments of the present invention can reduce the loss of image information when performing size change operations after segmenting the original image captured by the pan-tilt camera, and then hand it over to at least two image analysis units for analysis, which can improve the accuracy and success rate of target detection.

[0143] The target detection device, storage medium, and computer device provided in the above embodiments can execute the target detection method provided in any embodiment of the present invention, and have the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in the above embodiments, reference can be made to the target detection method provided in any embodiment of the present invention.

[0144] Note that the above is only the preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A target detection method, characterized in that, Applied to a drone, at least two image analysis units are integrated in the drone, and the method includes: Obtain a first image captured by a gimbal camera; Perform segmentation processing on the first image to obtain at least two segmented images; Perform a size change operation on the at least two segmented images to obtain at least two sub-images, and the resolutions of the at least two sub-images match the resolutions corresponding to the at least two image analysis units; Input the at least two sub-images into the at least two image analysis units, and determine a target detection result according to the analysis results of the at least two image analysis units; The determining the target detection result according to the analysis results of the at least two image analysis units includes: Obtain the analysis results of the at least two image analysis units; Perform fusion processing on at least two analysis results to obtain a target detection result; The analysis results include type information and position information of the analyzed targets, and the performing fusion processing on at least two analysis results includes: Sequentially record every two adjacent sub-images as a current sub-image pair, the current sub-image pair includes a first sub-image and a second sub-image, and perform the following operations on the current sub-image pair: Determine a first target in the first sub-image and a second target in the second sub-image; According to the first position information and first type information corresponding to the first target, and the second position information and second type information corresponding to the second target, determine whether the first target and the second target correspond to the same target. If so, fuse the first target and the second target into the same target; The position information includes the coordinates of a rectangular frame, and the image corresponding to the analyzed target is included in the rectangular frame; the rectangular frame corresponding to the first target is denoted as the first rectangular frame, and the rectangular frame corresponding to the second target is denoted as the second rectangular frame; The determining whether the first target and the second target are the same target according to the first position information and first type information corresponding to the first target, and the second position information and second type information corresponding to the second target includes: According to the coordinates of the first rectangular frame and the coordinates of the second rectangular frame, calculate a first distance between a first boundary of the first rectangular frame and a third boundary of the second rectangular frame, a second distance between a third boundary of the first rectangular frame and a first boundary of the second rectangular frame, a third distance between a second boundary of the first rectangular frame and a fourth boundary of the second rectangular frame, and a fourth distance between a fourth boundary of the first rectangular frame and a second boundary of the second rectangular frame. Among them, the first boundary and the third boundary in each rectangular frame are parallel, the second boundary and the fourth boundary are parallel. When the first sub-image and the second sub-image are adjacent left and right, the first boundary in each rectangular frame is the left boundary. When the first sub-image and the second sub-image are adjacent up and down, the first boundary in each rectangular frame is the upper boundary; Calculate the first ratio of the smaller distance to the larger distance among the first distance and the second distance; and calculate the second ratio of the smaller distance to the larger distance among the third distance and the fourth distance; When the first ratio is less than a first preset threshold, the second ratio is greater than a second preset threshold, and the first type information and the second type information are the same, determine whether the first target and the second target correspond to the same target, where the first preset threshold is less than the second preset threshold.

2. The method according to claim 1, characterized in that, The fusing the first target and the second target into the same target includes: Determine a target rectangular frame according to the coordinates of the first rectangular frame and the coordinates of the second rectangle, where the target rectangular frame includes both the first rectangular frame and the second rectangle; Determine the analysis result corresponding to the fused target by using the target rectangular frame and the first type information.

3. The method according to claim 1, wherein Denote the boundary where the first sub-image coincides with the second sub-image as the segmentation boundary. The determining the first target in the first sub-image includes: Obtain the alternative targets analyzed from the first sub-image; For each alternative target, judge the fifth distance between the current alternative target and the segmentation boundary. If the fifth distance is less than a third preset threshold, determine the current alternative target as the first target.

4. The method according to any one of claims 1-3, characterized in that, The at least two image analysis units include at least two forward inference engines NNIE accelerated by a neural network.

5. A target detection device, characterized in that, Applied to a drone, where at least two image analysis units are integrated in the drone, the device includes: An image acquisition module, configured to acquire a first image captured by a gimbal camera; An image segmentation module, configured to perform segmentation processing on the first image to obtain at least two segmented images; A size change module, configured to perform a size change operation on the at least two segmented images to obtain at least two sub-images, where the resolutions of the at least two sub-images match the resolutions corresponding to the at least two image analysis units; A target detection module, configured to input the at least two sub-images into the at least two image analysis units, and determine a target detection result according to the analysis results of the at least two image analysis units; The determining the target detection result according to the analysis results of the at least two image analysis units includes: Obtain the analysis results of the at least two image analysis units; Perform fusion processing on at least two analysis results to obtain a target detection result; The analysis results include the type information and position information of the analyzed targets. The performing fusion processing on at least two analysis results includes: Successively denote each two adjacent sub-images as a current sub-image pair. The current sub-image pair includes a first sub-image and a second sub-image. Perform the following operations on the current sub-image pair: Determine a first target in the first sub-image and a second target in the second sub-image; Determine whether the first target and the second target correspond to the same target according to the first position information and the first type information corresponding to the first target, and the second position information and the second type information corresponding to the second target. If so, fuse the first target and the second target into the same target; The position information includes the coordinates of a rectangular box, and the image corresponding to the analyzed target is included within the rectangular box; the rectangular box corresponding to the first target is denoted as the first rectangular box, and the rectangular box corresponding to the second target is denoted as the second rectangular box; The determining whether the first target and the second target are the same target according to the first position information and the first type information corresponding to the first target, and the second position information and the second type information corresponding to the second target includes: According to the coordinates of the first rectangular box and the coordinates of the second rectangular box, calculate the first distance between the first boundary of the first rectangular box and the third boundary of the second rectangular box, the second distance between the third boundary of the first rectangular box and the first boundary of the second rectangular box, the third distance between the second boundary of the first rectangular box and the fourth boundary of the second rectangular box, and the fourth distance between the fourth boundary of the first rectangular box and the second boundary of the second rectangular box. Among them, the first boundary and the third boundary in each rectangular box are parallel, and the second boundary and the fourth boundary are parallel. When the first sub-image and the second sub-image are adjacent left and right, the first boundary in each rectangular box is the left boundary; when the first sub-image and the second sub-image are adjacent up and down, the first boundary in each rectangular box is the upper boundary; Calculate the first ratio of the smaller distance to the larger distance among the first distance and the second distance; and calculate the second ratio of the smaller distance to the larger distance among the third distance and the fourth distance; When the first ratio is less than the first preset threshold, the second ratio is greater than the second preset threshold, and the first type information and the second type information are the same, determine whether the first target and the second target are the same target, where the first preset threshold is less than the second preset threshold.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1-4.

7. A drone, comprising a memory, at least two image analysis units, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Multiprocessor-embedded image acquisition and processing method and device

    CN102510448A

  • Traffic flow monitoring method based on unmanned aerial vehicle, intelligent system and data set

    CN108831161A