Monocular distance measurement method, device, equipment, medium, and vehicle for intelligent vehicle driving
By using a monocular camera and a convolutional neural network model in an intelligent driving system for distance measurement, combined with clustering and geometric calculation models, the problems of high computational complexity, many application limitations and high cost of traditional binocular vision systems are solved, and high precision, wide application and low-cost distance measurement effects are achieved.
Patent Information
- Application Number
- JP2024535539
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-28
- Filing Date
- 2023-07-20
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2043-07-20
AI Technical Summary
In the prior art, the distance measurement method based on binocular vision systems has problems such as high computational complexity, many application limitations and high cost, and it is difficult to ensure measurement accuracy, wide application and low cost implementation at the same time.
A single-eye camera is used to measure intelligent driving distances, read the image information of the video stream, and use a convolutional neural network model to perform vehicle instance segmentation and depth map generation. Combining the clustering calculation model and the geometric distance measurement model, the minimum distance of the image point is calculated and fused to obtain the accurate distance from the single-eye camera to the target object.
It realizes high-precision, wide application and low-cost monocular distance measurement, which reduces computational complexity and cost compared with traditional binocular vision system methods, while improving measurement accuracy and flexibility.
Smart Images

Figure 2025515240000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to the technical field of vehicle intelligent driving, and in particular to a monocular distance measurement method, device, equipment, medium, and vehicle for vehicle intelligent driving. [Background technology]
[0002] Advanced driver assistance systems (ADAS) use sensors installed in the vehicle to constantly sense the surrounding environment, acquire data, and recognize, detect, and track static and dynamic objects, while also combining map data from the navigator to perform systematic calculations and analysis to allow the driver to be aware of possible dangers in advance, effectively improving the comfort and safety of driving.Here, measuring the distance to the vehicle ahead is one of the basic needs of advanced driver assistance systems, and is often used for functions such as the vehicle's adaptive cruise control (ACC) and automatic emergency braking (AEB), so there is a high demand for the accuracy of the calculation results when calculating the distance to the vehicle ahead.
[0003] Conventional distance measurement methods based on binocular vision systems mainly estimate the distance to an obstacle in a scene based on calculating the disparity between two cameras. Such methods require a large amount of calculations, and while ensuring accuracy, they also require high computing power of the processing unit. They also require the two cameras to be horizontally coaxial, the optical axes to be parallel, the acquisition to be synchronous, and the exposure parameters to be consistent, which greatly limits their practical application. In addition, the cost of providing two cameras and a processing unit with high computing power is high. Summary of the Invention [Problem to be solved by the invention]
[0004] The technical problem that the present invention aims to solve is that the distance measurement methods in the prior art cannot simultaneously take into account measurement accuracy, application limitations, and costs. In order to solve this technical problem, the present invention provides a monocular distance measurement method for intelligent vehicle driving, which has high measurement accuracy, no limitations in use, and low implementation costs. [Means for solving the problem]
[0005] The technical means adopted by the present invention to solve the technical problem are as follows: Step S1 of reading image information for each frame from a video stream captured by an on-board monocular camera; Step S2: inputting the image information into a first convolutional neural network model for detection, and obtaining a vehicle instance segmentation map and a depth map including at least one object detection frame; S3: extracting a rectangular frame corresponding to each of the object detection frames, dividing the vehicle instance segmentation map by the rectangular frame, obtaining a region of interest map corresponding to the object, and inputting the region of interest map into a second convolutional neural network model for detection, to obtain information of the object, including object category, actual height, and height of the object in the image; Step S4: extracting a depth ROI (region of interest) map using the object detection frame and the depth map, calculating a minimum distance of pixel points in the depth ROI map using a clustering calculation model, and acquiring coordinates s(u,v) of the pixel points, where the minimum distance of the pixel points is a first distance D1 from the vehicle-mounted monocular camera to the object in real space; A step S5 of calculating a second distance D2 from the vehicle-mounted monocular camera to the object in real space by linking the information of the object with the coordinates s(u,v) of the pixel point using a geometric distance measurement model; and step S6 of performing a fusion processing calculation for the first distance D1 and the second distance D2 to obtain a final distance from the vehicle-mounted monocular camera to the object in real space.fruit, Specifically, the calculation by the clustering calculation model in step S4 is Suppose there are n pixel points in the depth ROI map, and in the initial calculation, each pixel point is set to its own cluster S, and the depth value corresponding to the pixel point is set to dn in step S41; Step S42 of calculating the distance between two clusters and combining the two closest clusters into one cluster; Step S43 of repeating step S42 until the distance between every two clusters becomes greater than a threshold T; Step S44: Counting the number of all pixel points in the two clusters, obtaining a cluster SMax with the largest number of pixel points, reading a depth value corresponding to each pixel point in the cluster SMax, obtaining a minimum depth value dmin in the cluster SMax and a pixel point corresponding to the minimum depth value, and setting the minimum depth value dmin as the minimum distance of the pixel points; When the number of the minimum depth value dmin is 1, obtain the coordinates s(u,v) of the pixel point corresponding to the minimum depth value dmin; Alternatively, when the number of the minimum depth values dmin is multiple, there are multiple coordinates of pixel points, and the coordinates of the intersection point m in the pixel coordinate system o-uv of the optical axis of the vehicle-mounted monocular camera are (u0, v0), the Euclidean distance from each of the pixel points to the intersection point m is calculated, and when the Euclidean distance is a minimum value, the coordinates s(u, v) of the pixel point corresponding to the minimum depth value dmin are obtained; Specifically, in step S6, the standard deviation σ1 of the first distance D1 and the standard deviation σ2 of the second distance D2 are obtained, and a distance fusion standard deviation calculation formula, i.e.,
number
number
[0006] More specifically, extracting a rectangular frame corresponding to each of the object detection frames in step S3 is specifically Step S31 of creating a pixel coordinate system o-uv; The upper boundary point (u t ,v t ), lower boundary point (u d ,v d ), left boundary point (u l ,v l ) and the right boundary point (u r ,v r ) in step S32; A rectangular frame of the object detection frame is drawn using the coordinates of the upper boundary point, the lower boundary point, the left boundary point, and the right boundary point, and h=v d -v t and step S33 of setting the height of the object in the image.
[0007] More specifically, the calculation by the geometric distance measurement model in step S5 is Step S51: creating a camera coordinate system O-XYZ with the on-vehicle monocular camera as the origin, and setting the direction perpendicular to the pixel coordinate system o-uv plane as the Z-axis direction of the camera coordinate system; In the pixel coordinate system o-uv, the pixel point coordinates s(u,v) and the intersection point m(u0,v0) of the optical axis in the pixel coordinate system o-uv are known, the actual height of the object is H, the height of the object in the image is h, S(x,y,z) is the closest point of the object in the camera coordinate system, the SMN surface is parallel to the imaging plane, the intersection point of the optical axis and the SMN surface is M, and the M coordinate in the camera coordinate system is (0,0,z), and the following calculation formula, i.e.
number
[0008] More specifically, a Kalman filter algorithm is used to process the distance D3 from the vehicle-mounted monocular camera to the object in the final real space.
[0009] A monocular distance measurement device for intelligent driving of a vehicle using the above-mentioned monocular distance measurement method for intelligent driving of a vehicle, A first acquisition module for reading image information frame by frame from a video stream captured by an on-board monocular camera; a first detection module for inputting the image information into a first convolutional neural network model for detection, and obtaining a vehicle instance segmentation map and a depth map including a detection frame of at least one object; a second detection module for extracting a rectangular frame corresponding to each of the object detection frames, dividing the vehicle instance segmentation map by the rectangular frame, obtaining a region of interest map corresponding to the object, inputting the region of interest map into a second convolutional neural network model for detection, and obtaining information of the object, including object category, actual height, and height of the object in the image; A first calculation module extracts a depth ROI map using the object detection frame and the depth map, calculates a minimum distance of pixel points in the depth ROI map using a clustering calculation model, and obtains coordinates s(u,v) of the pixel points, where the minimum distance of the pixel points is a first distance D1 from the vehicle-mounted monocular camera to the object in real space; A second calculation module that calculates a second distance D2 from the vehicle-mounted monocular camera to the object in real space by linking the information of the object and the coordinates s(u,v) of the pixel point using a geometric distance measurement model; A monocular distance measurement device for intelligent vehicle driving is equipped with a fusion calculation module that performs fusion processing on the first distance D1 and the second distance D2 and obtains the final distance from the vehicle-mounted monocular camera to the object in real space.
[0010] A processor; a memory adapted to store executable commands; The processor is a computing device used to read the executable commands from the memory, execute the executable commands, and realize the above-mentioned monocular distance measurement method for vehicle intelligent driving.
[0011] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor realizes the above-mentioned monocular distance measurement method for intelligent vehicle driving.
[0012] The vehicle is equipped with the above-mentioned monocular distance measurement device for vehicle intelligent driving. Effect of the Invention
[0013] (1) Compared to the prior art, which estimates the distance to an obstacle in a scene based on calculating the disparity between two cameras, the present application is not limited in use and has low implementation costs thanks to the video stream captured by the on-board monocular camera. (2) After processing the original image information, the first distance is calculated using a clustering algorithm model, which can effectively eliminate external interference, and the accuracy of the first distance result is high; by combining the acquired actual data and calculating the second distance using a geometric distance measurement model, the accuracy of the second distance result is high; and as a result of performing the fusion calculation of the first distance and the second distance, the distance to the vehicle in front in the real space is obtained. Compared with the monocular distance measurement of the prior art, the present invention has high measurement accuracy, and the requirements for the computing power of the processing unit are relatively low, which further reduces the implementation cost.
[0014] The invention will now be further described with reference to the accompanying drawings and examples. [Brief description of the drawings]
[0015] [Figure 1] 1 is a flowchart of a first embodiment of the present invention. [Diagram 2] FIG. 2 is a schematic diagram of image information according to the first embodiment of the present invention. [Diagram 3] FIG. 2 is a schematic diagram of a vehicle instance segmentation map according to the first embodiment of the present invention; [Figure 4] FIG. 2 is a schematic diagram of a depth map according to the first embodiment of the present invention; [Diagram 5] FIG. 2 is a schematic diagram of an object detection frame in a pixel coordinate system according to the first embodiment of the present invention. [Figure 6] 1 is a partial depth ROI map according to the first embodiment of the present invention; [Figure 7] FIG. 2 is a schematic diagram of the calculation of the distance between two clusters in the first embodiment of the present invention; [Figure 8]FIG. 2 is a schematic diagram illustrating the principle of a geometric distance measurement model in the first embodiment of the present invention. [Figure 9] FIG. 11 is a schematic diagram of a configuration according to a second embodiment of the present invention. [Figure 10] FIG. 11 is a schematic diagram of a hardware configuration according to a third embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] The present invention will now be described in more detail with reference to the accompanying drawings, all of which are simplified schematic diagrams and only show the configuration related to the present invention, as they are used to briefly explain the basic configuration of the present invention.
[0017] In describing the present invention, it is to be understood that the orientations or positional relationships indicated by terms such as "center," "longitudinal," "lateral," "length," "width," "thickness," "up," "down," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," "circumferential," and the like, are based on the drawings and are intended merely to facilitate and simplify the description of the present invention, and are not to be understood as limiting the present invention, as they do not expressly or imply that the devices or elements described necessarily have, are configured, or are operated in a particular orientation. It is to be understood that a feature defined as "first" or "second" may expressly or imply one such feature or more such features. In describing the present invention, unless otherwise specified, "plurality" means two or more than two.
[0018] In the description of the present invention, it should be explained that unless otherwise clearly specified or limited, the terms "attach", "couple" and "connect" should be understood in a broad sense, for example, fixedly connected, detachably connected, integrally connected, mechanically connected, electrically connected, directly connected, or indirectly connected via an intermediate medium, and the insides of two elements may be communicated. Those skilled in the art can understand the specific meaning of the above terms in the present invention according to the specific situation.
[0019] Example 1: The example of the present application provides a monocular distance measurement method for intelligent driving of a vehicle, including the following steps S1 to S6.
[0020] Step S1: Compared with the conventional technology of reading image information frame by frame from a video stream captured by an on-board monocular camera, and estimating the distance to an obstacle in the scene based on calculating the disparity between the two cameras, as shown in FIG. 2, the present application has no limitations in use, a wide range of application, and low implementation costs due to the video stream captured by the on-board monocular camera provided.
[0021] Step S2: Input the image information into a first convolutional neural network model for detection, and obtain a vehicle instance segmentation map and a depth map including at least one object detection frame.
[0022] The first convolutional neural network uses, but is not limited to, a SOLOv2 segmentation network configuration to obtain the vehicle instance segmentation map. The first convolutional neural network is provided with a depth map prediction branch to obtain the depth map, and the first convolutional neural network has high accuracy and high speed of realizing image segmentation. When there is only one vehicle in the forward road situation, there is only one object detection frame in the vehicle instance segmentation map, and when there are two vehicles in the forward road situation, there are two object detection frames in the vehicle instance segmentation map, and the number of object detection frames in the vehicle instance segmentation map is the same as the number of vehicles in the forward road situation, as shown in FIG. 3. The depth is for representing the distance of each point in the forward road situation scene to the vehicle-mounted monocular camera, that is, each pixel value in the depth map represents the distance of a point in the forward road situation scene to the vehicle-mounted monocular camera. A depth map is an image formed by the vertical depth value of an object instead of the grayscale of a grayscale image, and the length and width of the depth map are the same as the length and width of the image information. Figure 4 is a schematic diagram of the depth map of this embodiment, in which the deeper the color, the closer the distance to the vehicle, and the lighter the color, the farther the distance to the vehicle.
[0023] Step S3: Extract a rectangular frame corresponding to each object detection frame, divide the vehicle instance segmentation map by the rectangular frame, obtain a region of interest map corresponding to the object, and input the region of interest map into a second convolutional neural network model for detection to obtain information of the object, including the object category, actual height, and height in the image of the object.
[0024] In this embodiment, as shown in FIG. 5, extracting rectangular frames corresponding to each object detection frame in step S3 specifically includes the following steps S31 to S33. S31: Create a pixel coordinate system o-uv. S32: Obtain the upper boundary point (ut, vt), the lower boundary point (ud, vd), the left boundary point (ul, vl), and the right boundary point (ur, vr) of the object detection frame, and the object detection frame is a polygon but is not limited to this. S33: A rectangular frame of the object detection frame is drawn based on the coordinates of the upper boundary point, lower boundary point, left boundary point, and right boundary point, and h=vd-vt is set as the height of the object in the image.
[0025] In this embodiment, the vehicle instance segmentation map is segmented by a rectangular frame drawn based on an image segmentation algorithm, a region of interest map (ROI) corresponding to the object is obtained, and the type of the object is obtained by detection using a second convolutional neural network model, and a ShuffleNetv2 network configuration is used as the second convolutional neural network, but is not limited to this, which improves the accuracy of object classification and has a fast processing speed, and the object categories include, but are not limited to, passenger cars, SUVs, or buses, and further obtains the actual height corresponding to the object according to the object category, for example, the actual height of a passenger car is 1.5m, the actual height of an SUV is 1.7m, and the height of a bus is 3m, thereby further improving the accuracy of the calculation results.
[0026] Step S4: Extract a depth ROI map according to the object detection frame and the depth map, calculate the minimum distance of pixel points in the depth ROI map according to the clustering calculation model, obtain the coordinates s(u,v) of the pixel points corresponding to the minimum distance of the pixel points, and the minimum distance of the pixel points is the first distance D1 from the vehicle-mounted monocular camera to the object in the real space. Specifically, extract a depth ROI map according to the object detection frame and the depth map, and fuse the coordinates of the object detection frame in the vehicle instance segmentation map with the depth map to obtain a depth ROI (region of interest) map.
[0027] In this embodiment, the calculation using the clustering calculation model in step S4 includes the following steps S41 to S44. S41: Assume that there are n pixel points in the depth ROI map. During initial calculation, each pixel point is assigned to its own cluster S, and the depth value corresponding to the pixel point is dn. S42: Calculate the distance between every two clusters, and combine the two closest clusters into one cluster. Specifically, as shown in FIG. 7, the two clusters are SP and SQ, and cluster SP has AP pixel points and cluster SQ has AQ pixel points. The difference between the depth value dp1 corresponding to pixel point AP1 and the depth value dQ1 corresponding to pixel point AQ1 is calculated, and the result dPQ1 is set as the distance between pixel point AP1 and pixel point AQ1, and the distance between the two clusters L(SP, SQ) = min(dpQ) is obtained. S43: Repeat step S42 until the distance between the two clusters becomes greater than the threshold T. S44: Calculate the number of all pixel points in the two clusters, obtain the cluster SMax with the largest number of pixel points, read the depth value corresponding to each pixel point in the cluster SMax, obtain the minimum depth value dmin in the cluster SMax and the pixel point corresponding to the minimum depth value, and determine the minimum depth value dmin as the minimum distance between the pixel points. Here, when the number of the minimum depth value dmin is 1, the coordinate s(u,v) of the pixel point corresponding to the minimum depth value dmin is obtained; or when the number of the minimum depth value dmin is multiple, there are multiple coordinates of the pixel point, and the coordinate of the intersection point m in the pixel coordinate system o-uv of the optical axis of the vehicle-mounted monocular camera is (u0,v0), the Euclidean distance from each pixel point to the intersection point m is calculated, and when the Euclidean distance is the minimum value, the coordinate s(u,v) of the pixel point corresponding to the minimum depth value dmin is obtained. After performing a fusion process of the coordinate of the detection frame of the object in the vehicle instance segmentation map and the depth map, the minimum distance of the pixel point is obtained, and this minimum distance of the pixel point is the shortest distance from the forward vehicle to the host vehicle, and the coordinate of this pixel point in the pixel coordinates is obtained by the minimum distance of the pixel point, which facilitates the calculation of the subsequent steps, reduces the calculation deviation, and improves the accuracy of the calculation result.
[0028] When the vehicle-mounted monocular camera captures an image, the object will be shielded by a nearby object due to the perspective relationship of the position, so the contour of the nearby object will be relatively complete, but the contour of the distant object may be missing. The depth can be used to obtain the context of the object, and the distance from the vehicle-mounted monocular camera to the distant object is calculated. The case of being shielded in the depth map is specifically described as an example. Figure 6 is a 16*16 partial depth ROI map, where the pixels in area A are the depth of the nearby object, the pixels in area B are the depth of the distant object, and the black pixels are noise. When calculating the distance to the distant object, it is necessary to avoid the calculation interference caused by noise and nearby objects. The clustering calculation model of this embodiment is used to obtain three clusters, namely, cluster SA in area A, cluster SB in area B, and abnormal cluster SY, and the number of pixels in each cluster is statistically calculated. As a result, the cluster SA has the most pixels, and the minimum depth value in cluster SA is the minimum distance, that is, the first distance from the vehicle-mounted monocular camera to the object in real space. The clustering calculation model can eliminate the interference caused by nearby objects and noise (including nearby and distant noise) on the calculation of the distance to the object, improving the accuracy and stability of the calculation results. Furthermore, accurate calculations can be performed regardless of the vehicle's orientation, position, or whether it is obstructed, making it applicable in a wide range of applications.
[0029] Step S5: The object information is linked to the pixel point coordinates s(u, v) to calculate a second distance D2 from the vehicle-mounted monocular camera to the object in real space using a geometric distance measurement model.
[0030] In this embodiment, as shown in FIG. 8, the calculation by the geometric distance measurement model in step S5 is as follows: Step S51: creating a camera coordinate system O-XYZ with the on-vehicle monocular camera as the origin, and setting the direction perpendicular to the pixel coordinate system o-uv plane as the Z-axis direction of the camera coordinate system; In the pixel coordinate system o-uv, the pixel point coordinate s(u,v) and the intersection point m(u0,v0) of the optical axis in the pixel coordinate system o-uv are known, the actual height of the object is H, the height of the object in the image is h, S(x,y,z) is the closest point of the object in the camera coordinate system, the SMN plane is parallel to the imaging plane, the intersection point of the optical axis and the SMN plane is M, and the M coordinate in the camera coordinate system is (0,0,z), and the following calculation formula is used, i.e.
number
[0031] Step S6: A fusion process calculation is performed on the first distance D1 and the second distance D2 to obtain the final distance from the vehicle-mounted monocular camera to the object in real space.
[0032] Specifically, the standard deviation σ1 of the first distance D1 and the standard deviation σ2 of the second distance D2 are obtained, and the distance fusion standard deviation is calculated by the following formula:
number
number
[0033] In this embodiment, a Kalman filter algorithm 3 is further used to process the final distance D from the vehicle-mounted monocular camera to the object in real space, and the distance obtained after processing further improves the measurement accuracy of the distance measurement method.
[0034] In this embodiment, before executing the distance measurement method of the present invention, a threshold value T is calculated and the calculation result is stored. Specifically, in the experimental scene environment, a nearby object may occlude a distant object due to the perspective relationship of the positions of the objects. This scene environment is a special situation in this example. When the objects are occluded, a high-precision laser rangefinder is used to measure the distances from the two objects to the vehicle, which are L1 and L2 respectively. The difference between the two distances Ld = |L1-L2| is obtained. In this embodiment, the distance difference Ld in 1,000 sets of different occlusion situations is statistically calculated, and the smallest Ld(min) result is used as the threshold value T, thereby improving the accuracy of the calculation results of the clustering calculation model and further improving the measurement accuracy of the distance measurement method.
[0035] In this embodiment, the greater the distance to the object, the greater the distance measurement error. In the experimental scene environment, a high-precision laser range finder is used to measure the actual distance DZS to the object, and 1000 sets of image information samples are obtained in steps S1 to S5 of this embodiment, and the first distance D1 and the second distance D2 are calculated to obtain 1000 sets of [DZS, D1, D2] data. X1 and X2 are the relative distance ratios of D1 and D2, respectively, and the calculation formula for the relative distance ratio is:
number
number
[0036] According to the monocular distance measurement method for vehicle intelligent driving of the present invention, compared with the prior art that estimates the distance to an obstacle in a scene based on the calculation of the parallax between two cameras, the present invention is not limited in use by the video stream captured by the on-board monocular camera, and the implementation cost is low. After processing the original image information, the first distance is calculated by the clustering algorithm model, so that the external interference can be effectively eliminated, and the accuracy of the first distance result is high; the second distance is calculated by combining the acquired actual data with the geometric distance measurement model, so that the accuracy of the second distance result is high; and the distance to the front vehicle in the real space is obtained as a result of the fusion calculation of the first distance and the second distance; compared with the prior art monocular distance measurement, the present invention has high measurement accuracy, and the requirements for the computing power of the processing unit are relatively low, which further reduces the implementation cost.
[0037] Example 2: The embodiment of the present application is a monocular distance measurement device for intelligent driving of a vehicle using the above-mentioned monocular distance measurement method for intelligent driving of a vehicle, as shown in FIG. A first acquisition module 200 for reading image information frame by frame from a video stream captured by an on-board monocular camera; A first detection module 201 inputs image information into a first convolutional neural network model for detection, and obtains a vehicle instance segmentation map including at least one object detection frame and a depth map; a second detection module 202 for extracting a rectangular frame corresponding to each object detection frame, dividing the vehicle instance segmentation map by the rectangular frame, obtaining a region of interest map corresponding to the object, inputting the region of interest map into a second convolutional neural network model for detection, and obtaining information of the object, including the object category, actual height, and height in the image of the object; A first calculation module 203 extracts a depth ROI map using the object detection frame and the depth map, calculates the minimum distance of pixel points in the depth ROI map using a clustering calculation model, and obtains the coordinates s(u,v) of the pixel points, where the minimum distance of the pixel points is a first distance D1 from the vehicle-mounted monocular camera to the object in real space; a second calculation module 204 that calculates a second distance D2 from the vehicle-mounted monocular camera to the object in real space by linking the information of the object and the coordinates s(u,v) of the pixel point using a geometric distance measurement model; A monocular distance measurement device for intelligent vehicle driving is provided, which is equipped with a fusion calculation module 205 that performs fusion processing on the first distance D1 and the second distance D2 to obtain the final distance from the vehicle-mounted monocular camera to the object in real space.
[0038] The first acquisition module 200 and the second detection module 202 are both connected to the first detection module 201, the first calculation module 203 is connected to the first detection module 201, the first calculation module and the second detection module are both connected to the second calculation module 204, and the first calculation module 203 and the second calculation module 204 are both connected to the fusion calculation module 205.
[0039] In the device provided in the above embodiment, when realizing its functions, only the above-mentioned functional modules are described as an example, and in actual application, the above functions may be assigned to different functional modules to be completed as necessary, that is, it should be described that the internal configuration of the device is divided into different functional modules to complete all or part of the above-described functions. In addition, the device provided in the above embodiment belongs to the same concept as the method embodiment, and the details of the specific realization process may be referred to the method embodiment, and are omitted here.
[0040] Example 3: An embodiment of the present application provides a computer device comprising a processor and a memory storing at least one command or at least one program segment, the at least one command or the at least one program segment being loaded and executed by the processor to realize the monocular distance measurement method provided in the above method embodiment.
[0041] FIG. 10 shows a schematic diagram of a hardware configuration of an apparatus for implementing the monocular distance measurement method provided in the embodiment of the present application, and the apparatus may be for constituting or may include the apparatus or system provided in the embodiment of the present application. As shown in FIG. 10, the computer device 10 may include one or more processors 1002 (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1004 for storing data, and a transmission device 1006 for communication functions. Other components may include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power source, and / or a camera. It will be understood by those skilled in the art that the configuration shown in FIG. 10 is merely schematic and does not limit the configuration of the electronic device. For example, the computer device 10 may include more or less components than those shown in FIG. 10, or may have a different arrangement than that shown in FIG. 10.
[0042] It should be noted that the one or more processors and / or other data processing circuits may be generally referred to herein as "data processing circuitry". The data processing circuitry may be implemented in whole or in part as software, hardware, firmware, or any other combination. It should be noted that the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any one of the other elements in the computing device 10 (or mobile device). For example, as described in the embodiments of the present application, the data processing circuitry may be a processor to control (e.g., select the path of the variable resistance terminals connected to the interface).
[0043] The memory 1004 may be for storing software programs and modules for operating the software, for example, in a program command / data storage device corresponding to the monocular distance measurement method in the embodiment of the present application, the processor executes the software programs and modules stored in the memory 1004 to perform various function applications and data processing, i.e., to realize the above-mentioned method. The memory 1004 may include high-speed random access memory, or may include one or more non-volatile memories, such as magnetic storage devices, flash memories, and other non-volatile solid-state memories. In some examples, the memory 1004 may further include memories that are remotely provided with respect to the processor, and these remote memories may be connected to the computer device 10 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, a corporate intranet, a local area network, a mobile communication network, and combinations thereof.
[0044] The transmission device 1006 is for transmitting and receiving data via a network. A specific example of the network may include a wireless network provided by a communication provider of the computer device 10. In one example, the transmission device 1006 includes a network interface controller (NIC) that can connect to other network devices via a base station and communicate with the Internet. In one example, the transmission device 1006 may be a radio frequency (RF) module for communicating with the Internet wirelessly.
[0045] The display may be, for example, a touch screen liquid crystal display (LCD) that allows a user to interact with the user interface of the computing device 10 (or mobile device).
[0046] Example 4: An embodiment of the present application further provides a computer-readable storage medium that can be installed on a server and stores at least one command or at least one program segment for realizing the monocular distance measurement method in the method embodiment, wherein the at least one command or the at least one program segment is loaded and executed by the processor to realize the monocular distance measurement method provided in the above method embodiment.
[0047] Optionally, in this embodiment, the storage medium may be located in at least one of a plurality of network servers of a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as a USB memory, a read-only memory (ROM), a random access memory (RAM), a removable hard disk, a magnetic disk, or an optical disk.
[0048] Example 5: An embodiment of the present invention further provides a computer program product or a computer program including computer commands, the computer commands being stored in a computer readable storage medium. A processor of a computing device reads the computer commands from the computer readable storage medium, and the processor executes the computer commands, thereby causing the computing device to perform the monocular distance measurement methods provided in the various alternative embodiments described above.
[0049] Embodiment 6: The embodiment of the present invention further provides a vehicle including the above-mentioned monocular distance measurement device for vehicle intelligent driving.
[0050] It should be noted that the order of the above-mentioned embodiments of the present application is merely for illustrative purposes and does not represent the superiority or inferiority of the embodiments. Also, specific embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. The actions or steps recited in the claims may be performed in a different order than the embodiments and still achieve the desired results. Also, the steps depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. Multitasking and parallel processing are also possible or may contribute to some embodiments.
[0051] Each embodiment in the present application will be described step by step, and the same or similar parts between each embodiment may be mutually referred to, and each embodiment will be described focusing on the differences with other embodiments. As for the embodiments of the device, the apparatus and the storage medium, they are basically similar to the method embodiments, so the description will be relatively simple, and the relevant parts may be referred to the description of the method embodiments.
[0052] As can be understood by those skilled in the art, all or part of the steps in the above embodiments can be realized by hardware, or by giving commands to relevant hardware through a program, and the program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0053] The above preferred embodiments of the present invention are merely illustrative, and it is understood that those skilled in the art may make various changes or modifications to the above description without departing from the technical concept of the present invention. The technical scope of the present invention is not limited to the contents of the specification, but should be determined by the claims. [Explanation of symbols]
[0054] 20 Monocular distance measurement device 200 First Acquisition Module 201 First detection module 202 Second detection module 203 First Calculation Module 204 Second Calculation Module 205 Fusion Computing Module 10 Computer Equipment 1002 Processor 1004 Memory 1006 Transmission Equipment
Claims
1. Step S1 of reading image information for each frame from a video stream captured by an on-vehicle monocular camera; S2 inputting the image information into a first convolutional neural network model for detection, and obtaining a vehicle instance segmentation map and a depth map including at least one object detection frame; S3: extracting a rectangular frame corresponding to each of the object detection frames, dividing the vehicle instance segmentation map by the rectangular frame, obtaining a region of interest map corresponding to the object, and inputting the region of interest map into a second convolutional neural network model for detection, to obtain information of the object, including object category, actual height, and height of the object in the image; Step S4: extracting a depth ROI map using the detection frame of the object and the depth map, calculating a minimum distance of pixel points in the depth ROI map using a clustering calculation model, and acquiring coordinates s(u, v) of the pixel points, where the minimum distance of the pixel points is a first distance D1 from the vehicle-mounted monocular camera to the object in real space; A step S5 of calculating a second distance D2 from the vehicle-mounted monocular camera to the object in real space by linking the information of the object and the coordinates s (u, v) of the pixel point using a geometric distance measurement model; A monocular distance measurement method for intelligent vehicle driving, comprising: a step S6 of performing a fusion processing calculation on the first distance D1 and the second distance D2 to obtain the final distance from the vehicle-mounted monocular camera to the object in real space.
2. Specifically, extracting rectangular frames corresponding to the object detection frames in step S3 includes the following steps: Step S31 of creating a pixel coordinate system o-uv; A step S32 of acquiring an upper boundary point (ut, vt), a lower boundary point (ud, vd), a left boundary point (ul, vl), and a right boundary point (ur, vr) in the detection frame of the object; The monocular distance measurement method for intelligent driving of a vehicle as described in claim 1, characterized in that it includes a step S33 of drawing a rectangular frame of the object detection frame using the coordinates of the upper boundary point, the lower boundary point, the left boundary point, and the right boundary point, and setting h = vd - vt to the height of the object in the image.
3. The calculation by the clustering calculation model in step S4 is Suppose there are n pixel points in the depth ROI map, and in the initial calculation, each pixel point is assigned to a cluster S, and a depth value corresponding to the pixel point is assigned to dn in step S41; Step S42: calculating the distance between two clusters and combining the two closest clusters into one cluster; Step S43 of repeating step S42 until the distance between two clusters is greater than a threshold T; and a step S44 of calculating the number of all pixel points in the two clusters, respectively, to obtain a cluster SMax having the largest number of pixel points, reading a depth value corresponding to each pixel point in the cluster SMax, obtaining a minimum depth value dmin in the cluster SMax and a pixel point corresponding to the minimum depth value, and setting the minimum depth value dmin as a minimum distance between the pixel points. When the number of the minimum depth value dmin is 1, obtain the coordinates s(u, v) of the pixel point corresponding to the minimum depth value dmin; Alternatively, when the number of the minimum depth values dmin is multiple, there are multiple coordinates of pixel points, the coordinates of the intersection point m in the pixel coordinate system o-uv of the optical axis of the vehicle-mounted monocular camera are (u0, v0), the Euclidean distance from each of the pixel points to the intersection point m is calculated, and when the Euclidean distance is at its minimum value, the coordinate s(u, v) of the pixel point corresponding to the minimum depth value dmin is obtained.This is the monocular distance measurement method for intelligent driving of a vehicle as described in claim 2.
4. The calculation using the geometric distance measurement model in step S5 is Step S51: creating a camera coordinate system O-XYZ with the on-vehicle monocular camera as the origin, and setting the direction perpendicular to the pixel coordinate system o-uv plane as the Z-axis direction of the camera coordinate system; In the pixel coordinate system o-uv, the coordinates s (u, v) of the pixel point and the intersection m (u0, v0) of the optical axis in the pixel coordinate system o-uv are known, the actual height of the object is H, the height of the object in the image is h, S (x, y, z) is the closest point of the object in the camera coordinate system, the SMN surface is parallel to the imaging plane, the intersection point of the optical axis and the SMN surface is M, and the M coordinate in the camera coordinate system is (0, 0, z), and the following calculation formula, i.e. [0097] and a step S52 of calculating a second distance D2 from the vehicle-mounted monocular camera to the object in real space by Here, dz is the length distance of OM, dx is the distance from S to the OYZ plane, dy is the distance from S to the OXZ plane, fx and fy are internal parameters of the vehicle-mounted monocular camera, △u is the difference in distance between the pixel point s and the intersection point m in the u-axis direction, and △v is the difference in distance between the pixel point s and the intersection point m in the v-axis direction.The monocular distance measurement method for intelligent driving of a vehicle as described in claim 3.
5. Specifically, in step S6, a standard deviation σ1 of the first distance D1 and a standard deviation σ2 of the second distance D2 are obtained, and a distance fusion standard deviation is calculated using the following formula: [0010] Calculating the distance fusion standard deviation by The following calculation formula, i.e. ##EQU00011## The monocular distance measurement method for intelligent driving of a vehicle as described in claim 4, further comprising a step of calculating a distance D3 from the vehicle-mounted monocular camera to the object in the final real space based on the distance fusion standard deviation by using the above-mentioned method.
6. The monocular distance measurement method for intelligent driving of a vehicle as described in claim 5, further comprising processing a distance D3 from the vehicle-mounted monocular camera to the object in the final real space using a Kalman filter algorithm.
7. A monocular distance measurement device for intelligent driving of a vehicle using the monocular distance measurement method for intelligent driving of a vehicle according to any one of claims 1 to 6, A first acquisition module (200) for reading image information frame by frame from a video stream captured by an on-board monocular camera; a first detection module (201) for inputting the image information into a first convolutional neural network model for detection, and obtaining a vehicle instance segmentation map including at least one object detection frame and a depth map; a second detection module (202) for extracting a rectangular frame corresponding to each of the object detection frames, segmenting the vehicle instance segmentation map by the rectangular frame, obtaining a region of interest map corresponding to the object, inputting the region of interest map into a second convolutional neural network model for detection, and obtaining information of the object, including object category, actual height, and height of the object in the image; A first calculation module (203) extracts a depth ROI map using the detection frame of the object and the depth map, calculates a minimum distance of pixel points in the depth ROI map using a clustering calculation model, and obtains coordinates s (u, v) of the pixel points, where the minimum distance of the pixel points is a first distance D1 from the vehicle-mounted monocular camera to the object in real space; a second calculation module (204) that calculates a second distance D2 from the vehicle-mounted monocular camera to the object in real space by linking the information of the object and the coordinates s (u, v) of the pixel point using a geometric distance measurement model; A monocular distance measurement device for intelligent vehicle driving, characterized in that it comprises a fusion calculation module (205) that performs fusion processing on the first distance D1 and the second distance D2 to obtain the final distance from the vehicle-mounted monocular camera to the object in real space.
8. A processor; a memory adapted to store executable commands; The processor is a computer device used to read the executable commands from the memory, execute the executable commands, and realize the monocular distance measurement method for vehicle intelligent driving described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, the computer program being executed by a processor, the processor realizing the monocular distance measurement method for intelligent vehicle driving according to any one of claims 1 to 6.
10. A vehicle comprising the monocular distance measuring device for intelligent vehicle driving according to claim 7.
Citation Information
Patent Citations
Object detection device and object detection method and program
JP2019008460A
Object recognition system, advanced driver assistance system, and program
JP2021056620A
Network architecture for monocular depth estimation and object detection
JP2022142789A