A Target Detection Method, Device, and Storage Medium for Expressways
By splitting the image into multiple ROI areas and combining the detection results, the problems of large and long-term calculations in the prior art are solved, and fast and accurate object detection is achieved.
Patent Information
- Application Number
- CN202111574940.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-21
AI Technical Summary
The prior art methods of measuring large amounts of calculations and time-consuming in object detection, and compressing the calculation amount will lead to performance losses, making it difficult to quickly and accurately complete object detection under limited resources.
By splitting the original image into multiple ROI areas, using scaling and rotation matrix processing, the calculation amount of neural network is reduced, and the detection results are combined through the NMS algorithm to improve detection speed and accuracy.
It effectively reduces the calculation amount of the target detection model, improves the detection speed and accuracy, and reduces environmental mis-checking, and improves the utilization efficiency of computing resources.
Smart Images

Figure CN114387579B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method, a device and a storage medium for detecting an object on an expressway, and belongs to the technical field of driving assistance. Background Art
[0002] As the process of electric and intelligent vehicles continues to deepen, people's expectations are rising, software has changed the attributes of cars, and the continuous evolution of autonomous driving has gradually liberated the driver's attention until it is completely free of human intervention. The realization of autonomous driving mainly relies on various sensors to perceive the environment around the vehicle, and to assess the danger through processing technologies such as target detection, recognition and tracking. Among them, target detection technology is the foundation of autonomous driving technology. Through the camera, image data of the road ahead is obtained in real time, and the category and location of various targets on the image are preliminarily determined, providing information for more accurate identification and control of subsequent vehicles. Target detection is often the longest module in the various modules of autonomous driving. How to quickly and accurately complete the detection of various targets with limited resources has become the top priority to ensure timeliness.
[0003] In the task of detecting road targets, it is usually necessary to directly send the entire image (1080P or 2K) into the neural network for forward transmission, and target detection requires the use of a neural network with a large amount of computation to ensure performance, so it takes a long time to process a frame of image. Due to the particularity of vehicle-mounted equipment and the timeliness requirements of ADAS functions, it is necessary to more reasonably allocate the already scarce computing resources, so we ensure the detection performance while compressing its computation as much as possible to reduce time consumption. Generally, there are two types of computational compression operations for detection models: (1) reducing the output data, extracting part of the original image or directly scaling the image size for detection, and (2) compressing the neural network structure, using a model with a smaller amount of computation to fit the results of a large model. Although these two methods can compress the computational complexity of the target detection model and reduce time consumption, they will both bring the side effect of performance loss. Summary of the invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and to provide a method, device and storage medium for detecting targets on a freeway, which can reduce the computational complexity of the target detection model and improve the detection accuracy.
[0005] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0006] In a first aspect, the present invention provides a method for detecting an object on an expressway, comprising:
[0007] Collect the original image of the road through the on-board ADAS forward camera;
[0008] Obtain multiple levels of ROI areas based on the original image;
[0009] Process the original image according to the ROI regions at multiple levels to obtain multiple input images;
[0010] Input the multiple input images into the target detection model respectively to obtain multiple target detection results;
[0011] Merge the multiple target detection results according to the NMS algorithm to obtain the final target detection result.
[0012] Optionally, the number of levels of the ROI regions is 2 - 3 levels.
[0013] Optionally, the number of levels of the ROI regions is 2 levels;
[0014] The acquisition of the first - layer ROI region includes:
[0015] Scale the original image according to the scaling ratio k to obtain the first - layer ROI region; the width and height of the first - layer ROI region are W / k and H / k;
[0016] The minimum value k of the scaling ratio min is:
[0017] k min = Max(W, H) / w max
[0018] The minimum value of the width of the target that can be detected by the first - layer ROI region target detection model is k min *r;
[0019] where, W and H are the width and height of the original image, w max and w min are the maximum and minimum values of the width of the target that can be detected by the target detection model;
[0020] The acquisition of the second - layer ROI region includes:
[0021] Determine the distance d between the target with a width less than k min *r to be detected and the camera in the world coordinate, the width W d and height H d of the region to be detected, and obtain the four vertex coordinates of the region to be detected as:
[0022]
[0023] where, h0 is the height of the camera;
[0024] Determine the rotation matrix R according to the pitch angle α and yaw angle β of the camera:
[0025]
[0026] The four vertex coordinates are corrected according to the rotation matrix R, and the corrected four vertex coordinates are transformed into image coordinates by using the principle of pinhole imaging. The upper left vertex and the lower right vertex in the image are found to obtain the second-layer ROI region, and the width and height of the second-layer ROI region are w and h;
[0027] Calculate the total area Sum of the input regions according to the width and height of the first-layer ROI region and the second-layer ROI region:
[0028] Sum = W / k × H / k + w × h
[0029] Starting from k = k min as the starting point, gradually increase the scaling ratio k to minimize the total area Sum of the region, and obtain the final scaling ratio k, thereby determining the final first-layer ROI region.
[0030] In a second aspect, the present invention provides a target detection device for a highway, and the device includes:
[0031] An image acquisition module, configured to acquire the original image of the road through an in-vehicle ADAS front camera;
[0032] An ROI region acquisition module, configured to acquire multiple levels of ROI regions based on the original image;
[0033] An image processing module, configured to process the original image according to multiple levels of ROI regions to obtain multiple input images;
[0034] A target detection module, configured to input the multiple input images into a target detection model respectively to obtain multiple target detection results;
[0035] A result generation module, configured to merge the multiple target detection results according to the NMS algorithm to obtain the final target detection result.
[0036] In a third aspect, the present invention provides a target detection device for a highway, including a processor and a storage medium;
[0037] The storage medium is used to store instructions;
[0038] The processor is configured to operate according to the instructions to execute the steps of the method according to any one of the above.
[0039] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the above are implemented.
[0040] Compared with the prior art, the beneficial effects achieved by the present invention:
[0041] The present invention provides a method, apparatus and storage medium for target detection on a fast road. By setting ROI regions, the original image is split into multiple ROI region images and then sent into a target detection model for target detection, which can not only effectively reduce the overall computational amount, improve the detection speed, but also increase the coverage range of the target detection scale and reduce some environmental misdetections. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flowchart of a method for target detection on a fast road provided by an embodiment of the present invention;
[0043] Figure 2 is a schematic diagram of processing the original image by two layers of ROI regions provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.
[0045] Embodiment 1:
[0046] As Figure 1 shown, an embodiment of the present invention provides a method for target detection on a fast road, including the following steps:
[0047] 1. Collect the original image of the road through an in-vehicle ADAS front camera.
[0048] 2. Obtain multiple levels of ROI regions based on the original image;
[0049] Due to the use of deep neural networks, the object scale range that the target detection model can detect is relatively wide, and the resource overhead caused by multiple calls to the neural network cannot be underestimated. Therefore, usually only two or three layers of ROI region designs are adopted;
[0050] The following is an example of the method for obtaining two layers of ROI regions:
[0051] 2.1. The acquisition of the first layer of ROI region includes:
[0052] Scale the original image according to the scaling ratio k to obtain the first layer of ROI region; the width and height of the first layer of ROI region are W / k and H / k;
[0053] The minimum value k of the scaling ratio min is:
[0054] k min = Max(W, H) / w max
[0055] The minimum value of the width of the object that can be detected by the first-layer ROI region target detection model is k min *r;
[0056] where W and H are the width and height of the original image, and w max and w min are the maximum and minimum values of the width of the object that can be detected by the target detection model;
[0057] 2.1. The acquisition of the second-layer ROI region includes:
[0058] Determine the distance d between the object with a width less than k min *r to be detected and the camera in the world coordinates. The width W d and height H d of the region to be detected are obtained, and the four vertex coordinates of the region to be detected are:
[0059]
[0060] where h0 is the height of the camera;
[0061] Determine the rotation matrix R according to the pitch angle α and yaw angle β of the camera:
[0062]
[0063] Correct the four vertex coordinates according to the rotation matrix R, and transform the corrected four vertex coordinates into image coordinates using the principle of pinhole imaging. Find the upper left vertex and lower right vertex in the image to obtain the second-layer ROI region. The width and height of the second-layer ROI region are w and h;
[0064] 2.3. Calculate the total area Sum of the input regions according to the widths and heights of the first-layer ROI region and the second-layer ROI region:
[0065] Sum = W / k × H / k + w × h
[0066] Starting from k = k min as the starting point, gradually increase the scaling ratio k to minimize the total area Sum of the region, and obtain the final scaling ratio k, thereby determining the final first-layer ROI region.
[0067] 3. Process the original image according to multiple levels of ROI regions to obtain multiple input images; as Figure 2 shown, the first-layer ROI region is the image obtained by scaling the original image by the scaling ratio k, and the second-layer ROI region is the image for the object with a width less than k min *r.
[0068] 4. Input multiple input images into the target detection model respectively to obtain multiple target detection results. The target detection model can adopt an existing conventional neural network detection model.
[0069] 5. Merge multiple target detection results according to the NMS algorithm to obtain the final target detection result.
[0070] Embodiment 2:
[0071] The present invention provides a target detection device for a fast road, and the device includes:
[0072] An image acquisition module, configured to acquire the original image of the road through an in-vehicle ADAS front camera;
[0073] An ROI region acquisition module, configured to acquire multiple levels of ROI regions based on the original image;
[0074] An image processing module, configured to process the original image according to multiple levels of ROI regions to obtain multiple input images;
[0075] A target detection module, configured to input multiple input images into the target detection model respectively to obtain multiple target detection results;
[0076] A result generation module, configured to merge multiple target detection results according to the NMS algorithm to obtain the final target detection result.
[0077] Embodiment 3:
[0078] Based on Embodiment 1, an embodiment of the present invention provides a target detection device for a fast road, including a processor and a storage medium;
[0079] The storage medium is used to store instructions;
[0080] The processor is configured to operate according to the instructions to execute the steps of the above method.
[0081] Embodiment 4:
[0082] Based on Embodiment 1, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0083] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0084] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0085] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0087] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A target detection method for a rapid road, characterized in that, Including: Collecting the original image of the road through the in-vehicle ADAS forward camera; Obtaining ROI regions at multiple levels based on the original image; Processing the original image according to the ROI regions at multiple levels to obtain multiple input images; Inputting the multiple input images into the target detection model respectively to obtain multiple target detection results; Merging the multiple target detection results according to the NMS algorithm to obtain the final target detection result; Wherein, the number of levels of the ROI region is 2; The obtaining of the first-layer ROI region includes: Scaling the original image according to the scaling ratio k to obtain the first-layer ROI region; the width and height of the first-layer ROI region are W / k and H / k; The minimum value k of the scaling ratio min is as follows: k min = Max(W, H) / w max The minimum value of the width of the target that can be detected by the first-layer ROI region target detection model is k min *r; where W and H are the width and height of the original image, and w max and w min are the maximum and minimum values of the width of the object that can be detected by the object detection model; The obtaining of the second-layer ROI region includes: Determine the distance d between the target with a width less than k to be detected and the camera in the world coordinates, the width W of the area to be detected min *r, and the height H d of the area to be detected, and obtain the coordinates of the four vertices of the area to be detected as follows: d Wherein, h0 is the height of the camera; Determining the rotation matrix R according to the pitch angle α and yaw angle β of the camera: Correcting the four vertex coordinates according to the rotation matrix R, and transforming the corrected four vertex coordinates into image coordinates by using the principle of pinhole imaging, and finding the upper left vertex and the lower right vertex in the image to obtain the second-layer ROI region, the width and height of the second-layer ROI region are w and h; Calculating the total area Sum of the input regions according to the width and height of the first-layer ROI region and the second-layer ROI region: Sum = W / k × H / k + w × h Starting from k = k min gradually increase the scaling ratio k so that the total area Sum is minimized, and obtain the final scaling ratio k, thereby determining the final first-layer ROI region.
2. An object detection device for a rapid road, characterized in that, The device includes: An image acquisition module, configured to collect the original image of the road through the in-vehicle ADAS forward camera; An ROI region acquisition module, configured to obtain ROI regions at multiple levels based on the original image; An image processing module, configured to process the original image according to the ROI regions at multiple levels to obtain multiple input images; A target detection module, configured to input the multiple input images into the target detection model respectively to obtain multiple target detection results; A result generation module, configured to merge the multiple target detection results according to the NMS algorithm to obtain the final target detection result; Wherein, the number of levels of the ROI region is 2; The obtaining of the first-layer ROI region includes: Scaling the original image according to the scaling ratio k to obtain the first-layer ROI region; the width and height of the first-layer ROI region are W / k and H / k; The minimum value k of the scaling ratio min is k min = Max(W, H) / w max The minimum value of the width of the target that can be detected by the first-layer ROI region target detection model is k min *r; where W and H are the width and height of the original image, and w max and w min are the maximum and minimum values of the width of the objects that can be detected by the object detection model; The obtaining of the second-layer ROI region includes: Determine the distance d between the target with a width less than k to be detected and the camera in the world coordinates, the width W of the area to be detected min *r, and the height H d of the area to be detected, and obtain the coordinates of the four vertices of the area to be detected as follows: d Wherein, h0 is the height of the camera; Determining the rotation matrix R according to the pitch angle α and yaw angle β of the camera: Correcting the four vertex coordinates according to the rotation matrix R, and transforming the corrected four vertex coordinates into image coordinates by using the principle of pinhole imaging, and finding the upper left vertex and the lower right vertex in the image to obtain the second-layer ROI region, the width and height of the second-layer ROI region are w and h; Calculating the total area Sum of the input regions according to the width and height of the first-layer ROI region and the second-layer ROI region: Sum = W / k × H / k + w × h Starting from k = k min Gradually increase the scaling ratio k so that the total area Sum is minimized, and obtain the final scaling ratio k, thereby determining the final first-layer ROI region.
3. An object detection device for a rapid road, characterized in that, Including a processor and a storage medium; The storage medium is used for storing instructions; The processor is used for operating according to the instructions to execute the steps of the method according to claim 1.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to claim 1 are implemented.
Citation Information
Patent Citations
Method, apparatus and computer readable media for object detection
CN112840347A