Monocular camera-based target detection method, apparatus and device, and medium

By using a monocular camera to acquire multi-view images in object detection, perform candidate box detection and non-maximum suppression processing, and merge and processed the processed candidate boxes, the problem of low target detection efficiency and accuracy in the prior art is solved, and a more efficient and accurate target detection effect is achieved.

CN119942073APending Publication Date: 2025-05-06ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510008651.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has low efficiency and accuracy in the target detection process, especially when the detection target is dense, it is prone to false detection and missed detection problems.

Method used

By acquiring the original image acquired by the monocular camera at a preset focal length and converting it into a telephoto image and a wide-angle image, the two are respectively detected by the preset candidate box detection model. Then, the detected candidate boxes are subjected to a non-maximum suppression process, and the processed candidate boxes are obtained, and they are mapped back to the original image for merging. Finally, the merged candidate boxes are subjected to a non-maximum suppression process to obtain the detection target.

Benefits of technology

By performing multiple non-maximum suppression processing on the detection box from different perspectives, the efficiency and accuracy of target detection are improved, and the difficulty of deployment on low-computing platforms is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942073A_ABST
    Figure CN119942073A_ABST
Patent Text Reader

Abstract

The invention relates to a target detection method and device based on a monocular camera, equipment and a medium. Obtaining an original image collected by the monocular camera under a preset focal length, and converting the original image into a long-focus image and a wide-angle image; respectively carrying out candidate frame detection on the long-focus image and the wide-angle image by utilizing a preset candidate frame detection model to obtain a first candidate frame and a second candidate frame; performing non-maximum suppression processing on the first candidate frame and the second candidate frame to obtain a processed first candidate frame and a processed second candidate frame; respectively mapping the processed first candidate frame and the processed second candidate frame to the original image, and merging the mapped candidate frames to obtain a merged candidate frame in the original image; and performing non-maximum suppression processing on the merged candidate box to obtain a detection target. In this way, multiple times of non-maximum suppression processing are performed on the detection frame under different visual angles, the target detection efficiency and accuracy are improved, and mass production deployment on a low-computing-power platform is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular to a method, device, equipment and medium for target detection based on a monocular camera. Background Art

[0002] As one of the core tasks of image processing, object detection aims to accurately identify and locate specific objects in images. However, due to the diversity and complexity of objects in images and the limitations of the detection algorithm itself, a large number of redundant detection frames are often generated during the object detection process, which seriously affects the accuracy and efficiency of detection.

[0003] In order to improve the accuracy and efficiency of target detection, the related technology adopts non-maximum suppression (NMS) technology, which only post-processes the candidate boxes detected from the image under a single perspective to reduce false detection and missed detection, and reduces the number of detection boxes to improve target detection efficiency.

[0004] However, directly using NMS technology to post-process the candidate boxes detected from the image has low efficiency and accuracy. Therefore, how to use non-maximum suppression processing technology for target detection to improve the efficiency and accuracy of target detection is a technical problem that needs to be solved urgently. Summary of the invention

[0005] In order to solve the above technical problems, the present disclosure provides a target detection method, device, equipment and medium based on a monocular camera.

[0006] In a first aspect, the present disclosure provides a target detection method based on a monocular camera, the method comprising:

[0007] Acquire an original image captured by a monocular camera at a preset focal length, and convert the original image into a telephoto image and a wide-angle image;

[0008] Using a preset candidate frame detection model, perform candidate frame detection on the telephoto image and the wide-angle image respectively, and obtain a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image respectively;

[0009] Performing non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image, respectively, to obtain a processed first candidate frame and a processed second candidate frame;

[0010] Mapping the processed first candidate frame and the processed second candidate frame onto the original image respectively, and merging the processed first candidate frame and the processed second candidate frame mapped onto the original image to obtain a merged candidate frame in the original image;

[0011] Non-maximum suppression processing is performed on the merged candidate box in the original image to obtain the detection target in the original image.

[0012] In a second aspect, the present disclosure provides a target detection device based on a monocular camera, the device comprising:

[0013] An image acquisition module is used to acquire an original image captured by a monocular camera at a preset focal length, and convert the original image into a telephoto image and a wide-angle image;

[0014] a candidate frame detection module, configured to perform candidate frame detection on the telephoto image and the wide-angle image respectively by using a preset candidate frame detection model, and obtain a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image respectively;

[0015] A first non-maximum suppression processing module is used to perform non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image, respectively, to obtain a processed first candidate frame and a processed second candidate frame;

[0016] A candidate box mapping module, used to map the processed first candidate box and the processed second candidate box onto the original image respectively;

[0017] a candidate frame merging module, configured to merge the processed first candidate frame and the processed second candidate frame mapped onto the original image to obtain a merged candidate frame in the original image;

[0018] The second non-maximum suppression processing module is used to perform non-maximum suppression processing on the merged candidate box in the original image to obtain the detection target in the original image.

[0019] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the device comprising:

[0020] one or more processors;

[0021] a storage device for storing one or more programs,

[0022] When one or more programs are executed by one or more processors, the one or more processors implement the method provided in the first aspect.

[0023] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method provided in the first aspect is implemented.

[0024] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:

[0025] The present invention discloses a method, apparatus, device and medium for target detection based on a monocular camera, which obtains an original image captured by a monocular camera at a preset focal length, and converts the original image into a telephoto image and a wide-angle image; uses a preset candidate frame detection model to perform candidate frame detection on the telephoto image and the wide-angle image, respectively, and obtains a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image, respectively; performs non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image, respectively, to obtain a processed first candidate frame and a processed second candidate frame; maps the processed first candidate frame and the processed second candidate frame to the original image, respectively, and merges the processed first candidate frame and the processed second candidate frame mapped to the original image to obtain a merged candidate frame in the original image; performs non-maximum suppression processing on the merged candidate frame in the original image to obtain a detection target in the original image. Therefore, first, non-maximum suppression processing is performed on the candidate frames obtained from the telephoto image and the wide-angle image, and then the processed candidate frames in the telephoto image and the processed candidate frames in the wide-angle image are mapped back to the original images and merged, and non-maximum suppression processing is performed again on the merged candidate frames. In this way, multiple non-maximum suppression processing is performed on the detection frames at different viewing angles, which improves the efficiency and accuracy of target detection and increases the possibility of deployment and mass production on low computing power platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0028] Figure 1 A schematic diagram of a flow chart of a target detection method based on a monocular camera provided in an embodiment of the present disclosure;

[0029] Figure 2 A schematic diagram of the structure of a target detection device based on a monocular camera provided in an embodiment of the present disclosure;

[0030] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0032] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0033] In the related art, the method of post-processing the candidate boxes detected from the image under a single perspective using the NMS technology is: first, the intersection over union (IoU) between each candidate box and the candidate box with the highest score is obtained, and then the candidate box with the highest IoU with the candidate box with the highest score is deleted to achieve post-processing of the candidate boxes, thereby eliminating redundant detection results.

[0034] However, the main problem with this method is that when the detection targets in the scene are dense, candidate boxes that originally belonged to different detection targets may be mistakenly deleted due to a high intersection-over-union ratio, resulting in false detection or missed detection, and the efficiency and accuracy of target detection cannot be guaranteed.

[0035] In order to solve the above problems, the following Figure 1 The target detection method based on a monocular camera provided in the embodiment of the present disclosure is described. In the embodiment of the present disclosure, the target detection method based on a monocular camera can be executed by an electronic device or a server. Among them, the electronic device may include a device with a communication function such as a tablet computer, a desktop computer, a laptop computer, etc., and may also include a device simulated by a virtual machine or a simulator. The server may be a cloud server or a server cluster.

[0036] Figure 1 A flow chart of a target detection method based on a monocular camera provided in an embodiment of the present disclosure is shown.

[0037] like Figure 1 As shown, the target detection method based on a monocular camera may include the following steps.

[0038] S110, acquiring an original image captured by a monocular camera at a preset focal length, and converting the original image into a telephoto image and a wide-angle image.

[0039] In this embodiment, when target detection is required for a target scene, a monocular camera is used to capture the original image of the target scene at a preset focal length, and the captured original image is sent to an electronic device. The electronic device then performs target detection after converting the original image into a telephoto image and a wide-angle image by performing multi-perspective transformation on the original image.

[0040] The target scene refers to any scene where target detection needs to be performed, for example, the target scene is the driving scene around the vehicle, the tourist activity scene at a tourist attraction, etc.

[0041] The preset focal length is a fixed focal length used to capture images, and the focal length can be pre-configured.

[0042] The original image refers to an image directly captured by a monocular camera at a preset focal length without any focal length adjustment or image information adjustment.

[0043] The telephoto image and the wide-angle image can both be determined by scaling and cropping the original image.

[0044] Specifically, the specific implementation method of "converting the original image into a telephoto image and a wide-angle image" in S110 includes but is not limited to the following method: using a preset first scaling factor to scale the original image to obtain a first scaled image, and using a preset second scaling factor to scale the original image to obtain a second scaled image, wherein the first scaling factor and the second scaling factor are adjustment ratios of the length and width of the original image, the first scaling factor is greater than 1, and the second scaling factor is less than 1; based on a preset first cropping starting point coordinate, the first scaled image is cropped to obtain a telephoto image, and based on a preset second cropping starting point coordinate, the second scaled image is cropped to obtain a wide-angle image.

[0045] Optionally, the telephoto image may be determined as follows:

[0046] H1=S1*H0-y1,W1=S1*W0-x1

[0047] Among them, H0 is the height of the original image, W0 is the width of the original image, the original image is represented as (H0, W0), S1 is the preset first scaling factor, x1 is the horizontal coordinate of the preset first cropping starting point, and y1 is the vertical coordinate of the preset first cropping starting point.

[0048] Optionally, the wide-angle image can be determined as follows:

[0049] H2=S2*H0-y2,W2=S2*W0-x2

[0050] Wherein, S2 is a preset second scaling factor, x2 is a preset horizontal coordinate of a second cropping starting point, and y2 is a preset vertical coordinate of the second cropping starting point.

[0051] S120 , using a preset candidate frame detection model, performing candidate frame detection on the telephoto image and the wide-angle image respectively, and obtaining a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image respectively.

[0052] In this embodiment, the electronic device inputs the telephoto image and the wide-angle image into a preset candidate frame detection model to implement candidate frame detection under multiple viewing angles, thereby obtaining a first candidate frame in the telephoto image and a second candidate frame in the wide-angle image.

[0053] Specifically, the specific implementation method of S120 includes but is not limited to the following methods: based on the feature extraction network in the preset candidate frame detection model, feature extraction is performed on the telephoto image and the wide-angle image respectively to obtain a feature map of the telephoto image and a feature map of the wide-angle image; based on the multi-scale feature fusion network in the preset candidate frame detection model, multi-feature fusion is performed on the feature map of the telephoto image and the feature map of the wide-angle image respectively to obtain multi-scale fusion features of the telephoto image and multi-scale fusion features of the wide-angle image; based on the prediction head in the preset candidate frame detection model, voxel prediction is performed on the multi-scale fusion features of the telephoto image and the multi-scale fusion features of the wide-angle image respectively to obtain a first candidate frame and a second candidate frame.

[0054] Optionally, the preset candidate box detection model may be a deep learning model for three-dimensional object detection (Fully Convolutional One-Stage Object Detection in 3D, FCOS-3D), or other neural network models with candidate box detection function.

[0055] It is understandable that, by using the preset candidate frame detection model, not only can the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image be output, but also the detection information of the candidate frames can be determined. Specifically, the first candidate frame corresponds to the first detection information, and the second candidate frame corresponds to the second detection information.

[0056] Optionally, both the first detection information and the second detection information include: two-dimensional information, three-dimensional information, size information of the detection target, posture of the detection target, and category of the detection target.

[0057] S130, performing non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image, respectively, to obtain a processed first candidate frame and a processed second candidate frame.

[0058] In order to improve the efficiency and accuracy of target detection, the electronic device can perform non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image, respectively, to achieve non-maximum suppression processing on the candidate frames under multi-viewing angles, and obtain the processed first candidate frame and the processed second candidate frame as the non-maximum suppression processing results under multi-viewing angles.

[0059] Specifically, the electronic device can perform non-maximum suppression processing on the first candidate frame and the second candidate frame in the wide-angle image based on the two-dimensional information of the first candidate frame and the two-dimensional information of the second candidate frame, respectively, to achieve two-dimensional non-maximum suppression processing and obtain a processed first candidate frame and a processed second candidate frame.

[0060] In this way, by performing two-dimensional non-maximum suppression processing on the candidate boxes under multiple perspectives, the impact of multi-perspective detection errors is effectively reduced.

[0061] S140, mapping the processed first candidate box and the processed second candidate box to the original image respectively, and merging the processed first candidate box and the processed second candidate box mapped to the original image to obtain a merged candidate box in the original image.

[0062] It can be understood that through the process of non-maximum suppression processing under multi-view, redundant candidate boxes can be suppressed to a certain extent. Therefore, after the processed first candidate box and the processed second candidate box are respectively mapped to the original image, the number of candidate boxes mapped to the original image is greatly reduced compared to the number of candidate boxes detected directly from the original image.

[0063] Specifically, the specific implementation method of "mapping the processed first candidate box and the processed second candidate box to the original image respectively" in S140 includes but is not limited to the following method: based on the two-dimensional information contained in the first detection information, the first mapping relationship between the telephoto image and the original image, mapping the processed first candidate box to the original image; based on the second mapping relationship between the two-dimensional information contained in the second detection information, the wide-angle image and the original image, mapping the processed second candidate box to the original image.

[0064] The first mapping relationship may be determined according to a first scaling factor between the original image and the telephoto image and a preset first cropping starting point coordinate. The second mapping relationship may be determined according to a second scaling factor between the original image and the wide-angle image and a preset second cropping starting point coordinate.

[0065] Optionally, the first candidate box processed and mapped to the original image can be represented as follows:

[0066] x'i=(xi+x1) / S1,y'i=(yi+y1) / S1

[0067] Among them, (x'i, y'i) is the coordinate data of the first candidate box after processing mapped to the original image, (xi, yi) is the coordinate data of the first candidate box after processing, S1 is the preset first scaling factor, x1 is the horizontal coordinate of the preset first cropping starting point, and y1 is the vertical coordinate of the preset first cropping starting point.

[0068] Optionally, the second candidate box processed and mapped to the original image can be represented as follows:

[0069] x'j=(xj+x2) / S2, y'j=(yj+y2) / S2

[0070] Among them, (x'j, y'j) is the coordinate data of the second candidate box after processing mapped to the original image, (xj, yj) is the coordinate data of the second candidate box after processing, S2 is the preset second scaling factor, x2 is the horizontal coordinate of the preset second cropping starting point, and y2 is the vertical coordinate of the preset second cropping starting point.

[0071] Among them, after the processed first candidate frame is mapped to the original image, it corresponds to the third candidate frame, the third candidate frame corresponds to the third detection information, and the third detection information is obtained by mapping the first detection information; after the processed second candidate frame is mapped to the original image, it corresponds to the fourth candidate frame, the fourth candidate frame corresponds to the fourth detection information, and the fourth detection information is obtained by mapping the second detection information. Then the third detection information and the fourth detection information both include: two-dimensional information, three-dimensional information, size information of the detection target, posture of the detection target, and category of the detection target.

[0072] Specifically, the specific implementation method of "merging the processed first candidate box and the processed second candidate box mapped on the original image to obtain a merged candidate box in the original image" in S140 includes but is not limited to the following method: merging the processed first candidate box and the processed second candidate box mapped on the original image according to the third detection information and the fourth detection information to obtain a merged candidate box in the original image.

[0073] Specifically, the electronic device can merge the processed first candidate box and the processed second candidate box mapped to the original image based on the two-dimensional information, three-dimensional information, size information of the detection target, posture of the detection target and category of the detection target respectively included in the third detection information and the fourth detection information to obtain a merged candidate box in the original image.

[0074] In this way, by mapping the candidate frames processed under two different perspectives, the telephoto image and the wide-angle image, onto the original image, and merging the processed first candidate frame and the processed second candidate frame mapped onto the original image, the purpose of reducing the number of candidate frames can be achieved again, and at the same time, it is convenient for more accurate target detection.

[0075] S150, performing non-maximum suppression processing on the merged candidate box in the original image to obtain the detection target in the original image.

[0076] In order to further improve the accuracy and efficiency of target detection, the electronic device can perform non-maximum suppression processing on the merged candidate box in the original image to determine the detection target in the original image.

[0077] Specifically, the specific implementation method of S150 includes but is not limited to the following methods: performing non-maximum suppression processing on the merged candidate box in the original image according to the category and two-dimensional information of the detection target contained in the third detection information, and the category and two-dimensional information of the detection target contained in the fourth detection information, to obtain the detection box to be selected in the original image; converting the original image from an initial perspective to a bird's-eye view; in the bird's-eye view, based on the size and three-dimensional information of the detection target contained in the third detection information, and the size and three-dimensional information of the detection target contained in the fourth detection information, performing non-maximum suppression processing on the detection box to be selected in the original image to obtain the detection target in the original image.

[0078] Specifically, the electronic device can first perform non-maximum suppression processing on the merged candidate box in the original image at the initial viewing angle of the original image based on the category and two-dimensional information of the detection target contained in the two detection information to obtain the detection box to be selected in the original image to achieve two-dimensional non-maximum suppression processing, and then perform non-maximum suppression processing on the detection box to be selected at the bird's-eye view (BEV) based on the size and three-dimensional information of the detection target contained in the two detection information to achieve three-dimensional non-maximum suppression processing, and finally obtain the detection target in the original image.

[0079] It is understandable that since the above-mentioned target detection method only uses a monocular camera, compared with the target detection method based on multi-camera or laser-collected images, it reduces the hardware cost and computing power, and increases the possibility of deployment and mass production on low-computing power platforms.

[0080] The present invention discloses a target detection method based on a monocular camera. The method comprises the following steps: obtaining an original image captured by a monocular camera at a preset focal length, and converting the original image into a telephoto image and a wide-angle image; performing candidate frame detection on the telephoto image and the wide-angle image respectively using a preset candidate frame detection model, and obtaining a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image respectively; performing non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image respectively to obtain a processed first candidate frame and a processed second candidate frame; mapping the processed first candidate frame and the processed second candidate frame respectively to the original image, and merging the processed first candidate frame and the processed second candidate frame mapped to the original image to obtain a merged candidate frame in the original image; performing non-maximum suppression processing on the merged candidate frame in the original image to obtain a detection target in the original image. Therefore, first, non-maximum suppression processing is performed on the candidate frames obtained from the telephoto image and the wide-angle image, and then the processed candidate frames in the telephoto image and the processed candidate frames in the wide-angle image are mapped back to the original images and merged, and non-maximum suppression processing is performed again on the merged candidate frames. In this way, multiple non-maximum suppression processing is performed on the detection frames at different viewing angles, which improves the efficiency and accuracy of target detection and increases the possibility of deployment and mass production on low computing power platforms.

[0081] The present disclosure also provides a monocular camera-based target detection device for implementing the above-mentioned monocular camera-based target detection method. Figure 2 In the disclosed embodiment, the target detection device based on a monocular camera may be an electronic device or a server. The electronic device may include a tablet computer, a desktop computer, a laptop computer, or other devices with communication functions, or may include a virtual machine or a device simulated by a simulator. The server may be a cloud server or a server cluster.

[0082] Figure 2 A schematic structural diagram of a target detection device based on a monocular camera provided in an embodiment of the present disclosure is shown.

[0083] like Figure 2 As shown, the target detection device 200 based on a monocular camera may include:

[0084] The image acquisition module 210 is used to acquire the original image captured by the monocular camera at a preset focal length, and convert the original image into a telephoto image and a wide-angle image;

[0085] A candidate frame detection module 220 is used to perform candidate frame detection on the telephoto image and the wide-angle image respectively using a preset candidate frame detection model, and obtain a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image respectively;

[0086] A first non-maximum suppression processing module 230 is used to perform non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image, respectively, to obtain a processed first candidate frame and a processed second candidate frame;

[0087] A candidate box mapping module 240, configured to map the processed first candidate box and the processed second candidate box onto the original image respectively;

[0088] A candidate frame merging module 250 is used to merge the processed first candidate frame and the processed second candidate frame mapped onto the original image to obtain a merged candidate frame in the original image;

[0089] The second non-maximum suppression processing module 260 is used to perform non-maximum suppression processing on the merged candidate box in the original image to obtain the detection target in the original image.

[0090] An object detection device based on a monocular camera according to an embodiment of the present disclosure obtains an original image captured by a monocular camera at a preset focal length, and converts the original image into a telephoto image and a wide-angle image; performs candidate frame detection on the telephoto image and the wide-angle image respectively using a preset candidate frame detection model, and obtains a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image respectively; performs non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image respectively to obtain a processed first candidate frame and a processed second candidate frame; maps the processed first candidate frame and the processed second candidate frame to the original image respectively, and merges the processed first candidate frame and the processed second candidate frame mapped to the original image to obtain a merged candidate frame in the original image; performs non-maximum suppression processing on the merged candidate frame in the original image to obtain a detection object in the original image. Therefore, first, non-maximum suppression processing is performed on the candidate frames obtained from the telephoto image and the wide-angle image, and then the processed candidate frames in the telephoto image and the processed candidate frames in the wide-angle image are mapped back to the original images and merged, and non-maximum suppression processing is performed again on the merged candidate frames. In this way, multiple non-maximum suppression processing is performed on the detection frames at different viewing angles, which improves the efficiency and accuracy of target detection and increases the possibility of deployment and mass production on low computing power platforms.

[0091] In some embodiments of the present disclosure, the image acquisition module 210 includes:

[0092] a scaling unit, configured to scale the original image using a preset first scaling factor to obtain a first scaled image, and scale the original image using a preset second scaling factor to obtain a second scaled image, wherein the first scaling factor and the second scaling factor are adjustment ratios of the length and width of the original image, the first scaling factor is greater than 1, and the second scaling factor is less than 1;

[0093] The cropping unit is used to crop the first zoomed image based on the preset first cropping starting point coordinates to obtain the telephoto image, and to crop the second zoomed image based on the preset second cropping starting point coordinates to obtain the wide-angle image.

[0094] In some embodiments of the present disclosure, the candidate box detection module 220 includes:

[0095] a feature extraction unit, configured to extract features from the telephoto image and the wide-angle image respectively based on a feature extraction network in the preset candidate frame detection model, so as to obtain a feature map of the telephoto image and a feature map of the wide-angle image;

[0096] A multi-feature fusion unit, configured to perform multi-feature fusion on the feature map of the telephoto image and the feature map of the wide-angle image based on the multi-scale feature fusion network in the preset candidate frame detection model, so as to obtain multi-scale fusion features of the telephoto image and multi-scale fusion features of the wide-angle image;

[0097] A prediction unit is used to perform voxel prediction on the multi-scale fusion features of the telephoto image and the multi-scale fusion features of the wide-angle image based on the prediction head in the preset candidate frame detection model to obtain the first candidate frame and the second candidate frame.

[0098] In some embodiments of the present disclosure, the first candidate box corresponds to first detection information, and the second candidate box corresponds to second detection information; after the processed first candidate box is mapped to the original image, it corresponds to a third candidate box, and the third candidate box corresponds to third detection information, and the third detection information is obtained by mapping the first detection information; after the processed second candidate box is mapped to the original image, it corresponds to a fourth candidate box, and the fourth candidate box corresponds to fourth detection information, and the fourth detection information is obtained by mapping the second detection information.

[0099] In some embodiments of the present disclosure, the candidate box mapping module 240 includes:

[0100] a first mapping unit, configured to map the processed first candidate frame onto the original image based on the two-dimensional information included in the first detection information and a first mapping relationship between the telephoto image and the original image;

[0101] The second mapping unit is used to map the processed second candidate frame onto the original image based on the two-dimensional information included in the second detection information and a second mapping relationship between the wide-angle image and the original image.

[0102] In some embodiments of the present disclosure, the candidate frame merging module 250 is specifically used to:

[0103] The processed first candidate box and the processed second candidate box mapped onto the original image are merged according to the third detection information and the fourth detection information to obtain a merged candidate box in the original image.

[0104] In some embodiments of the present disclosure, the second non-maximum suppression processing module 260 includes:

[0105] a two-dimensional non-maximum suppression processing unit, configured to perform non-maximum suppression processing on the merged candidate frame in the original image according to the category and two-dimensional information of the detection target contained in the third detection information and the category and two-dimensional information of the detection target contained in the fourth detection information, so as to obtain a detection frame to be selected in the original image;

[0106] A perspective conversion unit, used to convert the original image from an initial perspective to a bird's-eye view;

[0107] A three-dimensional non-maximum suppression processing unit is used to perform non-maximum suppression processing on the selected detection box in the original image under the bird's-eye view, based on the size and three-dimensional information of the detection target contained in the third detection information, and the size and three-dimensional information of the detection target contained in the fourth detection information, to obtain the detection target in the original image.

[0108] It should be noted that Figure 2 The monocular camera-based target detection device 200 shown can perform Figure 1 The various steps in the method embodiment shown in the figure are implemented Figure 1 The various processes and effects in the method embodiment shown are not described in detail here.

[0109] Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure is shown.

[0110] like Figure 3 As shown, the electronic device may include a processor 301 and a memory 302 storing computer program instructions.

[0111] Specifically, the processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0112] The memory 302 may include a large capacity memory for information or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory 302 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 302 may be inside or outside the integrated gateway device. In a particular embodiment, the memory 302 is a non-volatile solid-state memory. In a particular embodiment, the memory 302 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (Electrically Erasable Programmable ROM, EEPROM), an electrically rewritable ROM (EAROM) or a flash memory, or a combination of two or more of these.

[0113] The processor 301 reads and executes the computer program instructions stored in the memory 302 to perform the steps of the target detection method based on a monocular camera provided in the embodiment of the present disclosure.

[0114] In one example, the electronic device may further include a transceiver 303 and a bus 304. Figure 3 As shown, the processor 301, the memory 302 and the transceiver 303 are connected via a bus 304 and communicate with each other.

[0115] The bus 304 includes hardware, software, or both. For example, but not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a Memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 304 may include one or more buses. Although embodiments of the present application describe and illustrate a particular bus, the present application contemplates any suitable bus or interconnect.

[0116] The following is an embodiment of a computer-readable storage medium provided in an embodiment of the present disclosure. The computer-readable storage medium and the target detection method based on a monocular camera in the above-mentioned embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the computer-readable storage medium, reference can be made to the embodiment of the target detection method based on a monocular camera in the above-mentioned embodiments.

[0117] This embodiment provides a storage medium containing computer executable instructions. When the computer executable instructions are executed by a computer processor, they are used to perform a target detection method based on a monocular camera. The method includes:

[0118] Acquire an original image captured by a monocular camera at a preset focal length, and convert the original image into a telephoto image and a wide-angle image;

[0119] Using a preset candidate frame detection model, perform candidate frame detection on the telephoto image and the wide-angle image respectively, and obtain a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image respectively;

[0120] Performing non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image, respectively, to obtain a processed first candidate frame and a processed second candidate frame;

[0121] Mapping the processed first candidate frame and the processed second candidate frame onto the original image respectively, and merging the processed first candidate frame and the processed second candidate frame mapped onto the original image to obtain a merged candidate frame in the original image;

[0122] Non-maximum suppression processing is performed on the merged candidate box in the original image to obtain the detection target in the original image.

[0123] Of course, the storage medium containing computer executable instructions provided in the embodiments of the present disclosure is not limited to the above method operations, and the computer executable instructions can also execute related operations in the target detection method based on a monocular camera provided in any embodiment of the present disclosure.

[0124] Through the above description of the implementation methods, technicians in the relevant field can clearly understand that the present disclosure can be implemented with the help of software and necessary general hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions to enable a computer cloud platform (which can be a personal computer, server, or network cloud platform, etc.) to execute the target detection method based on a monocular camera provided in each embodiment of the present disclosure.

[0125] Note that the above are only preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present disclosure. Therefore, although the present disclosure is described in more detail through the above embodiments, the present disclosure is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present disclosure, and the scope of the present disclosure is determined by the scope of the appended claims.

Claims

1. A target detection method based on a monocular camera, characterized in that: include: Acquire an original image captured by a monocular camera at a preset focal length, and convert the original image into a telephoto image and a wide-angle image; Using a preset candidate frame detection model, perform candidate frame detection on the telephoto image and the wide-angle image respectively, and obtain a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image respectively; Performing non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image, respectively, to obtain a processed first candidate frame and a processed second candidate frame; Mapping the processed first candidate frame and the processed second candidate frame onto the original image respectively, and merging the processed first candidate frame and the processed second candidate frame mapped onto the original image to obtain a merged candidate frame in the original image; Non-maximum suppression processing is performed on the merged candidate box in the original image to obtain the detection target in the original image.

2. The method according to claim 1, characterized in that The converting the original image into a telephoto image and a wide-angle image comprises: Using a preset first scaling factor, scaling the original image to obtain a first scaling image, and using a preset second scaling factor, scaling the original image to obtain a second scaling image, wherein the first scaling factor and the second scaling factor are adjustment ratios of the length and width of the original image, the first scaling factor is greater than 1, and the second scaling factor is less than 1; Based on the preset first cropping starting point coordinates, the first zoomed image is cropped to obtain the telephoto image, and based on the preset second cropping starting point coordinates, the second zoomed image is cropped to obtain the wide-angle image.

3. The method according to claim 1, characterized in that The method of using a preset candidate frame detection model to perform candidate frame detection on the telephoto image and the wide-angle image, respectively, and obtaining a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image, respectively, includes: Based on the feature extraction network in the preset candidate frame detection model, feature extraction is performed on the telephoto image and the wide-angle image respectively to obtain a feature map of the telephoto image and a feature map of the wide-angle image; Based on the multi-scale feature fusion network in the preset candidate frame detection model, multi-feature fusion is performed on the feature map of the telephoto image and the feature map of the wide-angle image to obtain multi-scale fusion features of the telephoto image and multi-scale fusion features of the wide-angle image; Based on the prediction head in the preset candidate box detection model, voxel prediction is performed on the multi-scale fusion features of the telephoto image and the multi-scale fusion features of the wide-angle image to obtain the first candidate box and the second candidate box.

4. The method according to claim 1, characterized in that: The first candidate box corresponds to the first detection information, and the second candidate box corresponds to the second detection information; after the processed first candidate box is mapped to the original image, it corresponds to the third candidate box, and the third candidate box corresponds to the third detection information, and the third detection information is obtained by mapping the first detection information; after the processed second candidate box is mapped to the original image, it corresponds to the fourth candidate box, and the fourth candidate box corresponds to the fourth detection information, and the fourth detection information is obtained by mapping the second detection information.

5. The method according to claim 4, characterized in that The processing of mapping the processed first candidate frame and the processed second candidate frame onto the original image respectively includes: Mapping the processed first candidate frame onto the original image based on the two-dimensional information included in the first detection information and a first mapping relationship between the telephoto image and the original image; Based on the two-dimensional information included in the second detection information and a second mapping relationship between the wide-angle image and the original image, the processed second candidate frame is mapped onto the original image.

6. The method according to claim 4, characterized in that The step of merging the processed first candidate frame and the processed second candidate frame mapped onto the original image to obtain a merged candidate frame in the original image includes: The processed first candidate box and the processed second candidate box mapped onto the original image are merged according to the third detection information and the fourth detection information to obtain a merged candidate box in the original image.

7. The method according to claim 4, characterized in that The performing non-maximum suppression processing on the merged candidate box in the original image to obtain the detection target in the original image includes: According to the category and two-dimensional information of the detection target contained in the third detection information and the category and two-dimensional information of the detection target contained in the fourth detection information, non-maximum suppression processing is performed on the merged candidate frame in the original image to obtain a detection frame to be selected in the original image; Converting the original image from an initial perspective to a bird's-eye view; Under the bird's-eye view, based on the size and three-dimensional information of the detection target contained in the third detection information, and the size and three-dimensional information of the detection target contained in the fourth detection information, non-maximum suppression processing is performed on the selected detection box in the original image to obtain the detection target in the original image.

8. A target detection device based on a monocular camera, characterized in that: include: An image acquisition module is used to acquire an original image captured by a monocular camera at a preset focal length, and convert the original image into a telephoto image and a wide-angle image; a candidate frame detection module, configured to perform candidate frame detection on the telephoto image and the wide-angle image respectively by using a preset candidate frame detection model, and obtain a first candidate frame and a second candidate frame from the telephoto image and the wide-angle image respectively; A first non-maximum suppression processing module is used to perform non-maximum suppression processing on the first candidate frame in the telephoto image and the second candidate frame in the wide-angle image, respectively, to obtain a processed first candidate frame and a processed second candidate frame; A candidate box mapping module, used to map the processed first candidate box and the processed second candidate box onto the original image respectively; a candidate frame merging module, configured to merge the processed first candidate frame and the processed second candidate frame mapped onto the original image to obtain a merged candidate frame in the original image; The second non-maximum suppression processing module is used to perform non-maximum suppression processing on the merged candidate box in the original image to obtain the detection target in the original image.

9. An electronic device, characterized in that: include: processor; A memory for storing executable instructions; The processor is used to read the executable instructions from the memory and execute the executable instructions to implement the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the method according to any one of claims 1 to 7.