Target detection method and device, computer device and storage medium
By combining two-dimensional and three-dimensional image acquisition devices to extract features, the problem of insufficient resolution of lidar was solved, and high-precision target detection was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA AUTOMOTIVE INNOVATION CORP
- Filing Date
- 2023-08-10
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, the angular resolution of lidar is relatively small, resulting in low resolution of point cloud data for distant targets, which affects the accuracy of target classification, recognition, and location detection.
By combining two-dimensional and three-dimensional image acquisition devices, two-dimensional and three-dimensional features of the target to be detected are extracted respectively. The target information, including type information and location information, is determined using the two-dimensional and three-dimensional features.
It improves the accuracy and efficiency of target detection, especially in the identification and location detection of distant targets, enhancing both detection accuracy and computational efficiency.
Smart Images

Figure CN116994225B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a target detection method, apparatus, computer device, and storage medium. Background Technology
[0002] With the continuous development of automotive intelligence and related semiconductor industries, technologies such as forward environmental perception and high-precision positioning are currently used to detect the driving environment ahead of the vehicle in real time. Existing technologies use a combination of monocular CMOS cameras and LiDAR to identify and detect obstacles in front of the vehicle, providing accurate information on the type and relative coordinates of the target obstacles.
[0003] However, due to the small angular resolution of lidar, the generated point cloud data has low resolution. When measuring obstacles at a distance, the projection points of the target are relatively sparse. Therefore, the existing technology has poor accuracy in classifying, identifying and detecting the location of distant targets. Summary of the Invention
[0004] Therefore, it is necessary to provide a target detection method, apparatus, computer equipment, and storage medium that can improve detection accuracy in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a target detection method, which includes:
[0006] Acquire the first image and the second image;
[0007] Extract the two-dimensional features of the target to be detected in the first image, and extract the three-dimensional features of the target to be detected in the second image;
[0008] Based on two-dimensional and / or three-dimensional features, determine the target information of the target to be detected.
[0009] In one embodiment, the target information includes type information. The target information of the target to be detected is determined based on two-dimensional and three-dimensional features, including:
[0010] Based on the two-dimensional features, type detection is performed on the target to be detected, and the first type detection result is obtained;
[0011] Based on the three-dimensional features, type detection is performed on the target to be detected to obtain the second type detection result;
[0012] Based on the results of the first type of detection and the second type of detection, the type information of the target to be detected is determined.
[0013] In one embodiment, the first type of detection result includes a first type and a first confidence value that the target to be detected belongs to the first type, and the second type of detection result includes a second type and a second confidence value that the target to be detected belongs to the second type;
[0014] Based on the results of the first type of detection and the second type of detection, determine the type information of the target to be detected, including:
[0015] If the first type and the second type are the same, then the first type or the second type shall be taken as the target type to which the target to be detected belongs;
[0016] If the first type and the second type are different, then the target type to which the target to be detected belongs is selected from the first type and the second type based on the first confidence value and the second confidence value.
[0017] In one embodiment, the target information includes location information; determining the target information of the target to be detected based on two-dimensional or three-dimensional features includes:
[0018] If the target to be detected is determined to be a static target, then the position information of the target to be detected is determined based on the two-dimensional features of the target;
[0019] If the target to be detected is determined to be a dynamic target, then the position information of the target to be detected is determined based on the three-dimensional features of the target.
[0020] In one embodiment, the target information of the target to be detected is determined based on two-dimensional features and three-dimensional features, including:
[0021] Based on two-dimensional and three-dimensional features, real-time localization and map building (SLAM) are performed to obtain an environmental map of the data collection area; the location information of the target to be detected is marked in the environmental map.
[0022] The environmental map of the collected area is matched with the actual map to obtain the location information of the target to be detected in the actual map.
[0023] In one embodiment, the three-dimensional features of the target to be detected include the three-dimensional modeling result and depth features of the target to be detected;
[0024] Extracting two-dimensional features of the target to be detected in the first image and extracting three-dimensional features of the target to be detected in the second image, including:
[0025] Two-dimensional scene semantic segmentation is performed on the first image to obtain the two-dimensional bounding box of the target to be detected in the first image, and three-dimensional scene semantic segmentation is performed on the second image to obtain the three-dimensional bounding box of the target to be detected in the second image.
[0026] Feature extraction is performed on the image region within the two-dimensional bounding box to obtain the two-dimensional features of the target to be detected;
[0027] The depth features of the 3D bounding box are obtained, and the image region within the 3D bounding box is modeled in 3D to obtain the 3D modeling result of the target to be detected.
[0028] In one embodiment, the method further includes:
[0029] Acquire multiple first images obtained by a two-dimensional image acquisition device at different times;
[0030] Extract the two-dimensional features of the target to be detected from each first image;
[0031] Based on the similarity between the two-dimensional features, each target to be detected in each first image is tracked.
[0032] In one embodiment, the method further includes:
[0033] Acquire multiple second images at different times using a 3D image acquisition device;
[0034] Perform 3D modeling of the target to be detected in each second image;
[0035] Based on the 3D modeling results, each target to be detected in each second image is tracked.
[0036] In one embodiment, the first image and the second image are acquired simultaneously from the same area by a two-dimensional image acquisition device and a three-dimensional image acquisition device, respectively; the two-dimensional image acquisition device is a complementary metal-oxide-semiconductor (CMOS) camera, and the three-dimensional image acquisition device is a single-photon avalanche diode (SPAD) camera.
[0037] Secondly, this application also provides a target detection device, which includes:
[0038] The acquisition module is used to acquire the first image and the second image;
[0039] The extraction module is used to extract the two-dimensional features of the target to be detected in the first image and the three-dimensional features of the target to be detected in the second image.
[0040] The determination module is used to determine the target information of the target to be detected based on two-dimensional and / or three-dimensional features.
[0041] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, characterized in that the processor executes the computer program to implement the steps of the method in the first aspect described above.
[0042] Fourthly, this application also provides a computer-readable storage medium. This computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.
[0043] Fifthly, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the method in the first aspect described above.
[0044] The aforementioned target detection method, apparatus, computer equipment, and storage medium acquire a first image and a second image; extract two-dimensional features of the target to be detected from the first image and three-dimensional features of the target to be detected from the second image; and determine the target information of the target to be detected based on the two-dimensional features and / or three-dimensional features. This application can directly determine the target information of the target to be detected based on two-dimensional or three-dimensional features to reduce the computational load of target detection and improve detection efficiency. Alternatively, it can combine two-dimensional and three-dimensional features to determine the target information of the target to be detected, thereby improving the accuracy of target detection. Attached Figure Description
[0045] Figure 1 This is a diagram illustrating an application scenario of the target detection method in one embodiment;
[0046] Figure 2 This is a flowchart illustrating a target detection method in one embodiment;
[0047] Figure 3 This is a schematic diagram of the process for type detection of the target to be detected in one embodiment;
[0048] Figure 4 This is a schematic diagram of the process for position detection of the target to be detected in one embodiment;
[0049] Figure 5 This is a flowchart illustrating the type detection process for the target to be detected in another embodiment;
[0050] Figure 6 This is a flowchart illustrating the extraction of two-dimensional and three-dimensional features in one embodiment;
[0051] Figure 7 This is a schematic diagram of the process of tracking the target to be detected in one embodiment;
[0052] Figure 8 This is a schematic diagram of the process for tracking the target to be detected in another embodiment;
[0053] Figure 9 This is a flowchart illustrating the target detection method in another embodiment;
[0054] Figure 10 This is a structural block diagram of a target detection device in one embodiment;
[0055] Figure 11 This is a structural block diagram of the target detection device in another embodiment;
[0056] Figure 12 This is a structural block diagram of the target detection device in another embodiment;
[0057] Figure 13 This is a structural block diagram of the target detection device in another embodiment;
[0058] Figure 14 This is a structural block diagram of the target detection device in another embodiment;
[0059] Figure 15 This is a structural block diagram of the target detection device in another embodiment;
[0060] Figure 16 This is a structural block diagram of a computer device for target detection in one embodiment. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0062] The target detection method provided in this application embodiment can be applied to, for example... Figure 1 The application environment shown is as follows. Among them, the vehicle with intelligent driving function is equipped with an intelligent driving domain controller. The main functional modules of the intelligent driving domain controller include a forward driving perception module 101, a perception data calculation and fusion SoC (System on Chip) module 102, a planning and control MCU (Microcontroller Unit) module 103, an internal communication module 104, and an external communication module 105, etc.
[0063] The forward-facing perception module 101 captures a first image and a second image at the same time using a CMOS (Complementary Metal Oxide Semiconductor) forward-facing camera and a SPAD (Single Photon Avalanche Diode) forward-facing camera, respectively. The first image and the second image are then transmitted to the perception data calculation and fusion SoC module 102, which extracts the two-dimensional features of the target to be detected in the first image and the three-dimensional features of the target to be detected in the second image. Based on the two-dimensional features and / or the three-dimensional features, the target information of the target to be detected is determined. Furthermore, the perception data calculation and fusion SoC module 102 transmits the target information of the target to be detected to the planning and control MCU module 103 through the internal communication module 104.
[0064] Specifically, the forward-facing perception module 101 includes a SPAD forward-facing camera, a CMOS forward-facing camera, and a millimeter-wave radar, used to provide accurate forward-facing perception data for the vehicle in different scenarios and environments. The perception data calculation and fusion SoC module 102 is mainly responsible for the calculation and processing of perception data and the fusion of perception results, used to provide the vehicle with the type, static characteristic parameters, and dynamic characteristic parameters of target obstacles, and to continuously track the status of each target within the perception field of view. The planning and control MCU module 103 receives the perception result data provided by the perception data calculation and fusion SoC module 102 through the internal communication module 104, and then calculates reasonable behavior decisions and path planning, and sends the decision planning results to the external execution unit through the external communication module 105 for the lateral and longitudinal control of the vehicle. At the same time, it sends environmental perception data and vehicle motion attitude change data to the planning and control MCU module 103 in real time to form a complete PID closed-loop control system, and dynamically adjusts the vehicle attitude in real time.
[0065] In addition, the main functional modules of the intelligent driving domain controller may also include a driving lateral perception module and a surround-view parking perception module. The driving lateral perception module includes at least four side-view cameras to provide the vehicle with blind spot vision data on both sides and in front and behind. The surround-view parking perception module includes at least four surround-view cameras, four long-range ultrasonic radars, and eight short-range ultrasonic radars for parking space detection and environmental perception.
[0066] In one embodiment, such as Figure 2 As shown, a target detection method is provided, including the following steps:
[0067] S201, acquire the first image and the second image.
[0068] The first image and the second image were acquired at the same time from the same area using a two-dimensional image acquisition device and a three-dimensional image acquisition device, respectively.
[0069] Two-dimensional and three-dimensional image acquisition devices can be installed on moving vehicles to acquire images of the road environment ahead using the same frame rate and trigger signal. The real-time images are then transmitted to the intelligent driving domain controller via LVDS (Low-Voltage Differential Signaling) interface to assist in intelligent driving of the vehicle.
[0070] For example, a two-dimensional image acquisition device and a three-dimensional image acquisition device can be installed on the uppermost central axis of the inner side of the vehicle's windshield to form a forward-facing dual-redundant visual perception system. Optionally, the two-dimensional image acquisition device can be a CMOS camera, and the three-dimensional image acquisition device can be a SPAD camera. There are no special requirements for the installation position and angle of the CMOS camera and the SPAD camera, thus reducing the installation size and facilitating miniaturization.
[0071] Each pixel of a CMOS camera measures the amount of light arriving at the pixel within a given time period, while a SPAD camera measures each photon arriving at the pixel. A SPAD camera places a diode within each pixel; each diode, upon receiving an incident photon, can convert that photon into charge carriers through an "avalanche effect," generating a large electrical pulse signal. This ability to generate an avalanche multiplication effect from a single photon provides higher sensitivity and greater distance measurement accuracy during image capture.
[0072] When the processor in the intelligent driving domain controller has computing power, it can directly acquire the first and second images with the same timestamp transmitted to the intelligent driving domain controller and perform target detection.
[0073] When the processor in the intelligent driving domain controller lacks sufficient computing power, the domain controller can upload a first image and a second image with the same timestamp to a terminal or server with computing power. The terminal or server can then acquire the first and second images and perform target detection. The server can be a standalone server or a server cluster consisting of multiple servers.
[0074] S202, extract the two-dimensional features of the target to be detected in the first image, and extract the three-dimensional features of the target to be detected in the second image.
[0075] For the first image, the image region corresponding to the target to be detected in the first image is segmented, and the image features of the image region corresponding to the target to be detected are extracted to obtain the two-dimensional features of the target to be detected. The target to be detected is a target in the road environment in front of the vehicle, such as cones, vehicles, traffic signs, traffic lights, and lane lines.
[0076] Optionally, two-dimensional features of the target to be detected can be extracted using a CNN (Convolutional Neural Networks) algorithm. These two-dimensional features can be used to characterize the target's visual and motion information, such as its size, color, and whether it is in motion or stationary.
[0077] For the second image, the image region corresponding to the target to be detected in the second image is segmented, and the image features and depth features of the image region corresponding to the target to be detected are extracted. Then, the target to be detected is 3D modeled using the image features and depth features to obtain the three-dimensional features of the target to be detected.
[0078] Optionally, when using a SPAD camera to capture a second image, the depth features of the target to be detected in the second image can be directly acquired. Simultaneously, a relatively simple algorithm can be used to extract the image features of the target. Then, the depth features and image features are combined to create a 3D model of the target, yielding its three-dimensional features. The three-dimensional features of the target can include its depth features and the 3D modeling results. The depth features can characterize the relative distance between the target and the vehicle where the 3D image acquisition device is located. The 3D modeling results can characterize the target's size, color, shape, and more detailed information such as its motion or stationary state.
[0079] Compared to two-dimensional features, three-dimensional features have higher accuracy, but the computational cost required to extract them is also greater.
[0080] It is understood that the first or second image generally contains multiple targets to be detected. This embodiment uses the extraction of two-dimensional and three-dimensional features of one target as an example for illustration, and does not mean that only the two-dimensional and three-dimensional features of one target are extracted. In practical applications, it is often necessary to extract the two-dimensional features of all targets to be detected in the first image and the three-dimensional features of all targets to be detected in the second image. The methods used are similar and will not be described in detail here.
[0081] S203, determine the target information of the target to be detected based on the two-dimensional features and / or the three-dimensional features.
[0082] Specifically, the target information of the target can be determined based on the two-dimensional features of the target, the three-dimensional features of the target, or a combination of the two-dimensional and three-dimensional features of the target.
[0083] The target information for the target to be detected includes the target's type and location information.
[0084] Regarding type information, if the probability of the target being detected being a cone calculated based on two-dimensional or three-dimensional features is 95%, which exceeds the preset threshold, then the type information of the target being detected can be directly determined to be a cone.
[0085] The probability of the target being detected being a vehicle is 80% based on two-dimensional features, and the probability of the target being detected being a sign is 85% based on three-dimensional features. The target being detected may be a sign containing vehicle information. Generally, the type information of the target being detected determined based on three-dimensional features is taken as the standard. Alternatively, the type information of the target being detected can be determined by comparing the probabilities calculated based on two-dimensional features and the probabilities calculated based on three-dimensional features.
[0086] Regarding location information, if the type of the target to be detected is determined to be a traffic cone, since traffic cones are generally stationary, the relative distance between a normally moving vehicle and a traffic cone in the road environment ahead should change linearly. The location information of the target to be detected can be determined directly based on two-dimensional features. For example, the location information of the target to be detected relative to the moving vehicle can be determined directly based on the changes in the two-dimensional features of the target to be detected in the first image of adjacent frames.
[0087] When the type of the target to be detected is determined to be a vehicle, since the vehicle may be stationary or in motion, the relative distance between a normally moving vehicle and stationary and moving vehicles in the road environment ahead will not be consistent. Therefore, it is necessary to determine the position information of the target to be detected based on three-dimensional features. For example, the position information of the target to be detected relative to the moving vehicle can be determined based on the changes in the three-dimensional features of the target to be detected in the second image of adjacent frames.
[0088] Optionally, the two-dimensional and three-dimensional features of the target to be detected are fused, and an environmental map containing the target to be detected is generated based on the fusion result. Then, based on the position information of the target to be detected in the environmental map, a more accurate position information of the target to be detected relative to the moving vehicle can be determined.
[0089] The above scheme acquires a first image and a second image; wherein the first image and the second image are acquired simultaneously from the same area by a two-dimensional image acquisition device and a three-dimensional image acquisition device, respectively; two-dimensional features of the target to be detected are extracted from the first image, and three-dimensional features of the target to be detected are extracted from the second image; target information of the target to be detected is determined based on the two-dimensional features and / or three-dimensional features. The three-dimensional image acquisition device used in this embodiment has high-precision imaging resolution and ToF (Time of Flight) ranging characteristics. Compared with lidar, it not only has higher target detection accuracy but also can achieve automatic iterative upgrades based on algorithms. Furthermore, the three-dimensional image acquisition device has high photosensitivity, enabling accurate depth imaging even in completely dark environments. Simultaneously, the installation position and angle of the two-dimensional and three-dimensional image acquisition devices have no special requirements, allowing for miniaturization by reducing installation size. In addition, directly determining the target information of the target to be detected based on two-dimensional or three-dimensional features can reduce the computational load of target detection and improve detection efficiency; combining two-dimensional and three-dimensional features to determine the target information of the target to be detected can improve the accuracy of target detection.
[0090] To accurately determine the type information of the target to be detected, based on the above embodiments, in one embodiment, the type information of the target to be detected can be determined according to two-dimensional features and three-dimensional features, thereby refining the above S203. For example... Figure 3 As shown, the specific steps may include:
[0091] S301, Based on the two-dimensional features, perform type detection on the target to be detected to obtain the first type detection result.
[0092] The two-dimensional features of the target to be detected are used as input to a neural network model, such as a CNN model, to perform type detection on the target and obtain the first type detection result as output. The first type detection result includes the first type and a first confidence value that the target belongs to the first type.
[0093] Understandably, the first-type detection results output by a neural network model typically include multiple first-types, along with a first-confidence value corresponding to each first-type. For example, the first-type detection results might include: traffic cones (with a first-confidence value of 80%), bollards (with a first-confidence value of 15%), and food waste (with a first-confidence value of 5%).
[0094] S302, Based on the three-dimensional features, perform type detection on the target to be detected to obtain the second type detection result.
[0095] The three-dimensional features of the target to be detected are used as input to a neural network model to perform type detection on the target, resulting in a second type detection result. This second type detection result includes the second type and a second confidence value indicating that the target belongs to the second type.
[0096] It is understandable that the second type of detection result output by the neural network model can also contain multiple second types, as well as a second confidence value corresponding to each second type.
[0097] S303, Based on the results of the first type of detection and the second type of detection, determine the type information of the target to be detected.
[0098] The first type of detection result is determined based on the two-dimensional features of the target, while the second type of detection result is determined based on the three-dimensional features of the target. Compared to two-dimensional features, the depth features contained in the three-dimensional features can be used to characterize the relative distance between the target and the vehicle equipped with the three-dimensional image acquisition device. Furthermore, the 3D modeling results in the three-dimensional features contain the shape information of the target in three-dimensional space. Therefore, the accuracy of the second type of detection result should be higher than that of the first type of detection result.
[0099] The first preset threshold corresponding to the first confidence value can be slightly higher than the second preset threshold corresponding to the second confidence value.
[0100] If the first confidence value of the first type is not higher than the first preset threshold, and the second confidence value of the second type is higher than the second preset threshold, then the second type is taken as the target type to which the target to be detected belongs.
[0101] If the first confidence value of the first type in the first type of detection result is higher than the first preset threshold, and the second confidence value of the second type in the second type of detection result is higher than the second preset threshold, the first confidence value and the second confidence value are compared, and then the type information of the target to be detected is determined from the first type and the second type based on the comparison result.
[0102] In an optional embodiment, if the first type and the second type are the same, then the first type or the second type is taken as the target type to which the target to be detected belongs; if the first type and the second type are different, then the target type to which the target to be detected belongs is selected from the first type and the second type according to the first confidence value and the second confidence value.
[0103] For example, in the first type of detection results, the first type includes cones, and the first confidence value corresponding to the cones is 80%. In the second type of detection results, the second type also includes cones, and the second confidence value corresponding to the cones is 95%. When the first type and the second type contain the same type of information, and the confidence value of the same type of information is high, the first type or the second type, i.e., the cones, is directly determined as the target type to which the target to be detected belongs.
[0104] In the first type of detection results, the first type includes vehicles, and the first confidence value corresponding to vehicles is 80%. In the second type of detection results, the second type includes signs, and the second confidence value corresponding to signs is 85%. Therefore, the type information corresponding to the larger value between the first and second confidence values, i.e., the signs, is determined as the target type to which the target to be detected belongs.
[0105] This embodiment compares and fuses the results of type detection based on two-dimensional features and the results of type detection based on three-dimensional features to obtain the final type detection result. This result can match the characteristics of various obstacles in the road ahead of the moving vehicle, effectively improving the accuracy of type detection of the target.
[0106] To accurately determine the location information of the target to be detected, based on the above embodiments, in one embodiment, the location information of the target to be detected can be determined according to two-dimensional features or three-dimensional features, thereby refining the above-described S203. For example... Figure 4 As shown, the specific steps may include:
[0107] S401, extract the two-dimensional features of the target to be detected in the first image, and extract the three-dimensional features of the target to be detected in the second image.
[0108] Two-dimensional features of the target to be detected in the first image are extracted using a CNN algorithm. Meanwhile, the depth features of the target to be detected in the second image are directly obtained using a SPAD camera. The image features of the target to be detected are extracted using a relatively simple algorithm. Then, the depth features and image features are combined to perform 3D modeling of the target to be detected, so as to obtain the three-dimensional features of the target to be detected.
[0109] S402, determine whether the target to be detected is a static target based on two-dimensional and / or three-dimensional features.
[0110] If yes, then execute S403; otherwise, execute S404.
[0111] When the target type of the target to be detected is determined based on its two-dimensional and three-dimensional features, the target type can be used to determine whether the target is a static or dynamic target.
[0112] For example, if the target to be detected is determined to be a cone based on its two-dimensional and three-dimensional features, it can be directly determined that the target to be detected is a static target.
[0113] Optionally, if the target type of the target to be detected is determined based on the two-dimensional and three-dimensional features of the target to be detected, the target to be detected may be judged as a static target or a dynamic target by combining the two-dimensional and / or three-dimensional features of the target to be detected.
[0114] For example, based on the two-dimensional and three-dimensional features of the target, it can be determined that the target is a vehicle, which may be a moving vehicle or a parked vehicle. By analyzing the changes in the two-dimensional features of the target in the first image of adjacent frames, such as the changes in the position of the target's pixels in the first image, and / or by analyzing the changes in the depth features contained in the three-dimensional features of the target in the second image, the static or dynamic state of the target can be determined, thus classifying the target as a static or dynamic target.
[0115] S403, determine the location information of the target to be detected based on the two-dimensional features of the target.
[0116] The relative distance between a moving vehicle and a static target in the road environment ahead changes linearly. If the target to be detected is a static target, the position information of the target to be detected, i.e. the relative distance between the target to be detected and the vehicle with the two-dimensional image acquisition device, is determined based on the two-dimensional features of the target to be detected.
[0117] For example, by acquiring the vehicle's speed and the position information of the pixel corresponding to the target to be detected in the first image of the adjacent frame, and then calculating the actual distance traveled by the vehicle when the first image of the adjacent frame was captured based on the interval time between the adjacent frames, and further, by establishing a mapping relationship between the actual distance traveled by the vehicle and the change in the position information of the pixel corresponding to the target to be detected in the first image of the adjacent frame, the position information of the target to be detected can be determined.
[0118] S404, Determine the location information of the target to be detected based on its three-dimensional features.
[0119] The relative distance between a moving vehicle and dynamic targets in the road environment ahead varies. If the target to be detected is dynamic, its position information—that is, the relative distance between the target and the vehicle equipped with the 3D image acquisition device—is determined based on its 3D features. Specifically, the target's position information is directly determined based on the depth features included in its 3D features.
[0120] For example, by obtaining the depth features of the target in the second image and then combining the mapping relationship between the depth features and the location information in the real world, the location information of the target can be determined.
[0121] In this embodiment, the target to be detected is classified. For static targets, the 2D imaging data corresponding to the first image is used first, and for dynamic targets, the 3D imaging data corresponding to the second image is used first. This can effectively reduce the number of layers in the neural network, thereby improving the efficiency of the neural network in calculating the position information of the target to be detected.
[0122] To further optimize the location information of the target to be detected, based on the above embodiments, in one embodiment, SLAM mapping technology can be used to determine more precise location information of the target to be detected, thereby refining the above-mentioned S203. For example... Figure 5 As shown, the specific steps may include:
[0123] S501 performs real-time localization and SLAM mapping based on two-dimensional and three-dimensional features to obtain an environmental map of the data collection area.
[0124] Based on the two-dimensional features of the target to be detected extracted from the first image, and the three-dimensional features composed of the image features and depth features of the target to be detected extracted from the second image, SLAM (Simultaneous Localization and Mapping) is performed on the road environment in front of the vehicle to obtain an environmental map of the road environment in front, that is, an environmental map of the acquisition area of the two-dimensional image acquisition device and the three-dimensional image acquisition device.
[0125] Among them, the visual SLAM algorithm can build a 3D map in real time and simultaneously track the position and orientation of the camera. The environmental map built based on the visual SLAM algorithm is marked with the location information of the target to be detected.
[0126] S502 matches the environmental map of the collected area with the actual map to obtain the location information of the target to be detected in the actual map.
[0127] When a high-precision real-world map exists, the environmental map constructed based on the visual SLAM algorithm can be matched with the real-world map. For example, if there are iconic targets in the road environment ahead, such as landmarks and signs, the environmental map and the real-world map can be matched based on these iconic targets.
[0128] Furthermore, based on the location information of the target being detected marked on the environmental map, the location information of the target being detected in the actual map can be determined.
[0129] This embodiment performs SLAM mapping based on two-dimensional and three-dimensional features, thereby mapping the target to be detected in the road environment in front of the vehicle to the actual map of the real world. It can not only determine the relative distance between the target to be detected and the vehicle, but also determine the actual location information of the target to be detected in the real world, thus improving the accuracy of target detection.
[0130] In order to extract accurate two-dimensional and three-dimensional features, based on the above embodiments, in one embodiment, two-dimensional scene semantic segmentation and three-dimensional scene semantic segmentation can be performed on the first image and the second image respectively, so as to achieve the refinement of S202 above. Figure 6 As shown, the specific steps may include:
[0131] S601, perform two-dimensional scene semantic segmentation on the first image to obtain the two-dimensional bounding box of the target to be detected in the first image, and perform three-dimensional scene semantic segmentation on the second image to obtain the three-dimensional bounding box of the target to be detected in the second image.
[0132] Scene semantic segmentation refers to classifying each pixel in an image to determine the class to which each pixel belongs, thereby segmenting the image into multiple bounding boxes, each bounding box corresponding to one or more targets to be detected.
[0133] Semantic segmentation of the first image is performed on the two-dimensional scene to obtain the two-dimensional bounding boxes of all targets to be detected in the first image. Correspondingly, semantic segmentation of the second image is performed on the three-dimensional scene to obtain the three-dimensional bounding boxes of all targets to be detected in the second image.
[0134] It is understandable that the first image and the second image are images of the road environment in front of the vehicle acquired at the same time, so the targets to be detected in the first image and the second image should correspond one-to-one.
[0135] S602, extract features from the image region within the two-dimensional bounding box to obtain the two-dimensional features of the target to be detected.
[0136] CNN algorithms are used to extract two-dimensional features of the target object. For example, each pixel in an image region within a two-dimensional bounding box stores image information. These pixels together form a numerical matrix, and specific features are extracted from the image region through one or more convolutional layers. Specifically, the convolutional kernel of the convolutional layer is multiplied by the corresponding bits of the numerical matrix and then added together to obtain the output of the convolutional layer.
[0137] S603: Obtain the depth features of the 3D bounding box and perform 3D modeling on the image region within the 3D bounding box to obtain the 3D modeling result of the target to be detected.
[0138] The depth features of each 3D bounding box in the captured second image are directly acquired by the SPAD camera and used as the depth features of the target to be detected. The image features of the image region within the 3D bounding box are extracted by the CNN algorithm or a simpler algorithm. Furthermore, the depth features and image features are combined to perform 3D modeling of the image region within the 3D bounding box, and the 3D modeling result of the target to be detected within the 3D bounding box is obtained, such as the shape information of the target to be detected.
[0139] This embodiment uses artificial intelligence and 3D modeling technology to extract two-dimensional and three-dimensional features. In terms of algorithm, it can automatically iterate and upgrade through self-learning, which can improve the accuracy and efficiency of feature extraction and has a greater advantage in target detection.
[0140] To enable real-time tracking of targets in the road environment ahead of the vehicle, in one embodiment, based on the above embodiments, static targets can be tracked based on the two-dimensional features corresponding to the first image, thereby refining the above target detection method. For example... Figure 7 As shown, the specific steps may include:
[0141] S701, acquire multiple first images obtained by a two-dimensional image acquisition device at different times.
[0142] Two-dimensional image acquisition equipment captures images of the road environment in front of the vehicle at a certain frame rate and trigger signal, obtaining several first images at different times.
[0143] Since the vehicle is assumed to be moving, the target in the first image acquired at different times will change. For example, if there is a stationary traffic cone in the road ahead, the cone will appear larger in the first image as the vehicle moves, indicating that the distance between the vehicle and the cone is decreasing. If there is a target vehicle traveling in the same direction in the road ahead, the target vehicle may appear larger or smaller in the first image as the vehicle moves, depending on the speed of both the target vehicle and the vehicle being captured.
[0144] S702, extract the two-dimensional features of the target to be detected in each first image.
[0145] For the first images acquired at different times, extract the two-dimensional features of all targets to be detected in each first image. Optionally, if type detection has already been performed on the targets to be detected in the first image, extract only the two-dimensional features of static targets in each first image.
[0146] S703, based on the similarity between the two-dimensional features, tracks each target to be detected in each first image.
[0147] The similarity between the two-dimensional features of the target to be detected in the first images of adjacent frames is calculated to identify the same target to be detected in different first images.
[0148] For example, first image a and first image b are adjacent frames. For the target to be detected in first image a, a unique ID is assigned to that target. Further, based on the two-dimensional features of each target in first image a and each target in first image b, the similarity between the two-dimensional features is calculated. If the similarity between two two-dimensional features is higher than a preset threshold, then the target corresponding to these two features is determined to be the same target. In the case of determining the same target, the ID of that target in first image a is used as the ID of that target in first image b, thereby enabling tracking of each target in each of the first images.
[0149] This embodiment compares and matches each target to be detected based on the similarity of the two-dimensional features of each target to be detected in the first image of adjacent frames. When the same target is determined, the same ID is assigned. Specifically, when the target to be detected is determined to be a static target, the first image is used to track each target to be detected, which can effectively reduce the number of layers in the neural network model and improve the computational efficiency of the neural network.
[0150] To enable real-time tracking of targets in the road environment ahead of the vehicle, in one embodiment, based on the above embodiments, dynamic targets can be tracked based on the three-dimensional features corresponding to the second image, thereby refining the above target detection method. For example... Figure 8 As shown, the specific steps may include:
[0151] S801 acquires multiple second images obtained by a 3D image acquisition device at different times.
[0152] The 3D image acquisition device captures images of the road environment in front of the vehicle at a certain frame rate and trigger signal, obtaining several second images acquired at different times.
[0153] S802, perform 3D modeling of the target to be detected in each second image.
[0154] For the second images acquired at different times, 3D modeling is performed on all targets to be detected in each second image. Optionally, if type detection has already been performed on the targets to be detected in the second images, 3D modeling is performed only on dynamic targets in each second image.
[0155] S803 tracks each target to be detected in each second image based on the 3D modeling results.
[0156] Based on the 3D modeling results, the same target to be detected is identified from different second images. Then, when the same target to be detected is identified, the same ID is assigned to the same target to be detected, thereby enabling the tracking of each target to be detected in each second image.
[0157] Based on the 3D modeling results, this embodiment compares and matches each target to be detected. When the same target is identified, the same ID is assigned. Specifically, when the target to be detected is determined to be a dynamic target, the second image is used first to track each target to be detected, which can effectively improve the reliability of target tracking.
[0158] In one embodiment, an alternative example of an object detection method is provided, such as... Figure 9 As shown, the specific process is as follows:
[0159] S901, acquire the first image and the second image.
[0160] The first image and the second image were acquired at the same time from the same area by a two-dimensional image acquisition device and a three-dimensional image acquisition device, respectively; the two-dimensional image acquisition device is a complementary metal-oxide-semiconductor (CMOS) camera, and the three-dimensional image acquisition device is a single-photon avalanche diode (SPAD) camera.
[0161] S902, perform two-dimensional scene semantic segmentation on the first image to obtain the two-dimensional bounding box of the target to be detected in the first image, and perform three-dimensional scene semantic segmentation on the second image to obtain the three-dimensional bounding box of the target to be detected in the second image.
[0162] S903 extracts features from the image region within the two-dimensional bounding box to obtain the two-dimensional features of the target to be detected, and obtains the depth features of the three-dimensional bounding box. It then performs three-dimensional modeling on the image region within the three-dimensional bounding box to obtain the three-dimensional modeling result of the target to be detected.
[0163] The three-dimensional features of the target to be detected include the three-dimensional modeling results and depth features of the target.
[0164] S904: Based on the two-dimensional features, perform type detection on the target to be detected to obtain the first type detection result; and based on the three-dimensional features, perform type detection on the target to be detected to obtain the second type detection result.
[0165] The first type of detection result includes the first type and the first confidence value of the target belonging to the first type, and the second type of detection result includes the second type and the second confidence value of the target belonging to the second type.
[0166] S905, if the first type in the first type detection result and the second type in the second type detection result are the same, the first type or the second type shall be taken as the target type to which the target to be detected belongs; if the first type and the second type are different, the target type to which the target to be detected belongs shall be selected from the first type and the second type according to the first confidence value and the second confidence value.
[0167] S906, when the target to be detected is determined to be a static target, the position information of the target to be detected is determined based on the two-dimensional features of the target to be detected; and when the target to be detected is determined to be a dynamic target, the position information of the target to be detected is determined based on the three-dimensional features of the target to be detected.
[0168] S907 performs SLAM mapping based on two-dimensional and three-dimensional features to obtain an environmental map of the data collection area.
[0169] The environmental map includes the location information of the targets to be detected.
[0170] S908 matches the environmental map of the collected area with the actual map to obtain the location information of the target to be detected in the actual map.
[0171] S909, based on the similarity between the two-dimensional features in multiple first images acquired by the two-dimensional image acquisition device at different times, each target to be detected in each first image is tracked; and based on the three-dimensional modeling results of each target to be detected in multiple second images acquired by the three-dimensional image acquisition device at different times, each target to be detected is tracked.
[0172] The specific process of the above steps can be found in the description of the above method embodiments. The implementation principle and technical effect are similar, and will not be repeated here.
[0173] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0174] Based on the same inventive concept, this application also provides a target detection apparatus for implementing the target detection method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more target detection apparatus embodiments provided below can be found in the limitations of the target detection method described above, and will not be repeated here.
[0175] In one embodiment, such as Figure 10 As shown, a target detection device 1 is provided, comprising: an acquisition module 10, an extraction module 20, and a determination module 30, wherein:
[0176] The acquisition module 10 is used to acquire the first image and the second image.
[0177] The first image and the second image were acquired at the same time from the same area using a two-dimensional image acquisition device and a three-dimensional image acquisition device, respectively.
[0178] The extraction module 20 is used to extract two-dimensional features of the target to be detected in the first image and to extract three-dimensional features of the target to be detected in the second image.
[0179] The determination module 30 is used to determine the target information of the target to be detected based on the two-dimensional features and / or the three-dimensional features.
[0180] In one embodiment, the target information includes type information, in Figure 10 On the basis of, such as Figure 11 As shown, the determining module 30 may include:
[0181] The first type detection unit 31 is used to perform type detection on the target to be detected based on two-dimensional features, and obtain the first type detection result.
[0182] The second type detection unit 32 is used to perform type detection on the target to be detected based on the three-dimensional features, and obtain the second type detection result.
[0183] The type determination unit 33 is used to determine the type information of the target to be detected based on the first type detection result and the second type detection result.
[0184] In one embodiment, the first type of detection result includes a first type and a first confidence value that the target to be detected belongs to the first type, and the second type of detection result includes a second type and a second confidence value that the target to be detected belongs to the second type.
[0185] Specifically, the type determination unit 33 can be used to determine the target type of the target to be detected when the first type and the second type are the same, and to select the target type of the target to be detected from the first type and the second type according to the first confidence value and the second confidence value when the first type and the second type are different.
[0186] In one embodiment, the target information includes location information, Figure 10 On the basis of, such as Figure 12 As shown, the determining module 30 may further include:
[0187] The first position detection unit 34 is used to determine the position information of the target to be detected based on the two-dimensional features of the target when it is determined that the target to be detected is a static target.
[0188] The second position detection unit 35 is used to determine the position information of the target to be detected based on the three-dimensional features of the target when it is determined that the target to be detected is a dynamic target.
[0189] In one embodiment, in Figure 10 On the basis of, such as Figure 13 As shown, the determining module 30 may further include:
[0190] Mapping unit 36 is used to perform real-time localization and map building using SLAM based on two-dimensional and three-dimensional features to obtain an environmental map of the collected area.
[0191] The environmental map includes the location information of the targets to be detected.
[0192] The third location determination unit 37 is used to match the environmental map of the collection area with the actual map to obtain the location information of the target to be detected in the actual map.
[0193] In one embodiment, the three-dimensional features of the target to be detected include the three-dimensional modeling result and depth features of the target. Figure 10 On the basis of, such as Figure 14 As shown, the extraction module 20 described above may include:
[0194] The semantic segmentation unit 21 is used to perform two-dimensional scene semantic segmentation on the first image to obtain the two-dimensional bounding box of the target to be detected in the first image, and to perform three-dimensional scene semantic segmentation on the second image to obtain the three-dimensional bounding box of the target to be detected in the second image.
[0195] The two-dimensional feature extraction unit 22 is used to extract features from the image region within the two-dimensional bounding box to obtain the two-dimensional features of the target to be detected.
[0196] The three-dimensional feature extraction unit 23 is used to obtain the depth features of the three-dimensional bounding box and perform three-dimensional modeling of the image region within the three-dimensional bounding box to obtain the three-dimensional modeling result of the target to be detected.
[0197] In one embodiment, the target detection device 1 may further include:
[0198] The first tracking module is used to track each target to be detected in each first image based on the similarity between each two-dimensional feature.
[0199] In one embodiment, the target detection device 1 may further include:
[0200] The second tracking module is used to track each target to be detected in each second image based on the 3D modeling results.
[0201] In one embodiment, the first image and the second image are acquired simultaneously from the same area by a two-dimensional image acquisition device and a three-dimensional image acquisition device, respectively; the two-dimensional image acquisition device is a complementary metal-oxide-semiconductor (CMOS) camera, and the three-dimensional image acquisition device is a single-photon avalanche diode (SPAD) camera.
[0202] In one embodiment, such as Figure 15 As shown, a target detection device 2 is provided, including a SPAD camera, a CMOS camera, a surround-view camera, and a side-view camera connected to a perception computing and fusion SoC module via an interface chip, and a millimeter-wave radar, an ultrasonic radar, and a control unit connected to a planning and control MCU module via another interface chip. The perception computing and fusion SoC module is connected to the planning and control MCU module, and both are connected to an ETH Switch module. A high-precision positioning map can also be connected to the ETH Switch module, which can be connected to a vehicle gateway.
[0203] Each module in the aforementioned target detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0204] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 16As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data such as a first image and a second image. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a target detection method.
[0205] Those skilled in the art will understand that Figure 16 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0206] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the target detection method described above.
[0207] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the target detection method described above.
[0208] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the target detection method described above.
[0209] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0210] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0211] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A target detection method characterized by, include: Acquire a first image and a second image; the first image and the second image are acquired at the same time from the same area using a two-dimensional image acquisition device and a three-dimensional image acquisition device, respectively. Two-dimensional features of the target to be detected in the first image are extracted, and three-dimensional features of the target to be detected in the second image are extracted. The three-dimensional features include the three-dimensional modeling result and depth features of the target to be detected. Based on the two-dimensional features and / or the three-dimensional features, the target information of the target to be detected is determined, wherein the target information includes location information; Based on the two-dimensional features or the three-dimensional features, the target information of the target to be detected is determined, including: If the target to be detected is determined to be a static target, the position information of the target to be detected is determined based on the changes in the two-dimensional features of the target to be detected in the first image of adjacent frames. If the target to be detected is determined to be a dynamic target, the position information of the target to be detected is determined based on the changes in the three-dimensional features of the target to be detected in the second image of adjacent frames.
2. The method of claim 1, wherein, The target information includes type information. Based on the two-dimensional features and the three-dimensional features, the target information of the target to be detected is determined, including: Based on the two-dimensional features, the target to be detected is subjected to type detection to obtain a first type detection result; Based on the three-dimensional features, the target to be detected is subjected to type detection to obtain a second type detection result; Based on the first type of detection result and the second type of detection result, the type information of the target to be detected is determined.
3. The method of claim 2, wherein, The first type of detection result includes a first type and a first confidence value that the target to be detected belongs to the first type; the second type of detection result includes a second type and a second confidence value that the target to be detected belongs to the second type. The step of determining the type information of the target to be detected based on the first type of detection result and the second type of detection result includes: If the first type and the second type are the same, then the first type or the second type shall be regarded as the target type to which the target to be detected belongs; If the first type and the second type are different, then the target type to which the target to be detected belongs is selected from the first type and the second type based on the first confidence value and the second confidence value.
4. The method of claim 1, wherein, Based on the two-dimensional features and the three-dimensional features, the target information of the target to be detected is determined, including: Based on the two-dimensional and three-dimensional features, real-time localization and map building (SLAM) are performed to obtain an environmental map of the acquisition area; wherein the location information of the target to be detected is marked in the environmental map. The environmental map of the collected area is matched with the actual map to obtain the location information of the target to be detected in the actual map.
5. The method of claim 1, wherein, The step of extracting two-dimensional features of the target to be detected in the first image and extracting three-dimensional features of the target to be detected in the second image includes: Two-dimensional scene semantic segmentation is performed on the first image to obtain the two-dimensional bounding box of the target to be detected in the first image, and three-dimensional scene semantic segmentation is performed on the second image to obtain the three-dimensional bounding box of the target to be detected in the second image. Feature extraction is performed on the image region within the two-dimensional bounding box to obtain the two-dimensional features of the target to be detected; The depth features of the three-dimensional bounding box are obtained, and the image region within the three-dimensional bounding box is modeled in three dimensions to obtain the three-dimensional modeling result of the target to be detected.
6. The method of claim 1, wherein, The method further includes: Acquire multiple first images obtained by a two-dimensional image acquisition device at different times; Extract the two-dimensional features of the target to be detected from each first image; Based on the similarity between the two-dimensional features, each target to be detected in each first image is tracked.
7. The method of claim 1, wherein, The method further includes: Acquire multiple second images at different times using a 3D image acquisition device; Perform 3D modeling of the target to be detected in each second image; Based on the 3D modeling results, each target to be detected in each second image is tracked.
8. The method of claim 1, wherein, The two-dimensional image acquisition device is a complementary metal-oxide-semiconductor (CMOS) camera, and the three-dimensional image acquisition device is a single-photon avalanche diode (SPAD) camera.
9. A target detection apparatus characterized by comprising: include: The acquisition module is used to acquire a first image and a second image; the first image and the second image are acquired at the same time from the same area by a two-dimensional image acquisition device and a three-dimensional image acquisition device, respectively. An extraction module is used to extract two-dimensional features of the target to be detected in the first image and to extract three-dimensional features of the target to be detected in the second image. The three-dimensional features include the three-dimensional modeling result and depth features of the target to be detected. The determining module is configured to determine the target information of the target to be detected based on the two-dimensional features and / or the three-dimensional features, wherein the target information includes location information; The determining module is specifically used for: If the target to be detected is determined to be a static target, the position information of the target to be detected is determined based on the changes in the two-dimensional features of the target to be detected in the first image of adjacent frames. If the target to be detected is determined to be a dynamic target, the position information of the target to be detected is determined based on the changes in the three-dimensional features of the target to be detected in the second image of adjacent frames.
10. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The steps of implementing the method of any one of claims 1-8 when the processor executes a computer program.
11. A computer readable storage medium having stored thereon a computer program, characterized in that When a computer program is executed by a processor, it implements the steps of the method of any one of claims 1-8.