High-dynamic target detection method and device, vehicle and storage medium

By combining image acquisition equipment and event cameras to obtain image and event data, target detection and supplementary detection are performed, solving the problem of poor detection performance of ordinary cameras under high dynamic conditions and achieving high-accuracy target detection under extreme lighting conditions.

CN115223142BActive Publication Date: 2026-01-23WUHAN LOTUS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210777409.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2026-01-23
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

In existing autonomous vehicles, ordinary cameras are prone to motion blur when sensing high-speed moving objects, and have a small dynamic range of light sensitivity, making it difficult to adapt to sudden changes in lighting or scenes with excessively bright or dark lighting, resulting in reduced target detection performance.

Method used

By combining image acquisition equipment and event cameras, multiple frames of images and event data are acquired. Target object information is obtained through image detection, and event data is used to supplement the detection of missed objects, thereby improving detection accuracy.

Benefits of technology

Under high dynamic conditions, the accuracy of target detection is improved, especially under extreme lighting conditions, it can effectively detect dynamic targets and enhance the robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223142B_ABST
    Figure CN115223142B_ABST
Patent Text Reader

Abstract

The application discloses a high-dynamic target detection method and device, a vehicle and a storage medium, relates to the technical field of target detection, and can improve the accuracy of target detection under high dynamics. The specific scheme comprises the following steps: acquiring a plurality of images collected by an image collection device in a preset time period and image sampling time; acquiring a plurality of event data collected by an event camera in the preset time period; performing target detection on each frame of image to obtain target object information, wherein the target object information comprises a detection frame of an object in the image; determining, according to the image sampling time and each event data, detection event data corresponding to the detection frame of each frame of image in the plurality of event data; determining target event stream data according to each detection event data, wherein the target event stream data comprises data other than the detection event data in the plurality of event data; performing target detection on the target event stream data to obtain missed detection object information; and obtaining a target detection result according to the target object information and the missed detection object information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of target detection, and particularly relates to a high-dynamic target detection method and device, a vehicle and a storage medium. BACKGROUND

[0002] With the rapid development of the automobile industry, automatic driving car technology has received extensive attention from the academic and industrial circles in recent years. Vehicle target detection is a challenging task in automatic driving vehicle technology. It is an important application in the field of automatic driving car technology and intelligent transportation system. It plays a key role in automatic driving technology. The purpose of vehicle target detection is to accurately locate the position of the remaining vehicles in the surrounding environment to avoid accidents with other vehicles.

[0003] At present, the mainstream sensor scheme of automatic driving vehicles all adopts ordinary cameras as visual perception sensors, but the sensing frame rate of ordinary cameras determines that ordinary cameras are prone to motion blur when perceiving high-speed moving objects, which reduces the target detection effect. At the same time, the photosensitive dynamic range of ordinary cameras is small, which is difficult to adapt to sudden changes in light or over-bright / over-dark scenes, and overexposure or underexposure may occur, for example, when entering or exiting a tunnel, the system cannot normally perceive the target in the area near the tunnel entrance, which also reduces the target detection effect. SUMMARY

[0004] The present application provides a high-dynamic target detection method, device, vehicle and storage medium, which can improve the accuracy of target detection under high dynamics.

[0005] To achieve the above purpose, the present application adopts the following technical solutions:

[0006] In an embodiment of the present application, a high-dynamic target detection method is provided, which comprises the following steps:

[0007] Obtaining a plurality of images collected by an image collection device in a preset time period, and image sampling times of each image;

[0008] Obtaining a plurality of event data collected by an event camera in a preset time period;

[0009] Performing target detection on each image to obtain target object information of each image, wherein the target object information includes a detection box of an object in the image;

[0010] According to the image sampling times of each image and each event data, determining the detection event data corresponding to the detection box of each image in the plurality of event data;

[0011] According to each detection event data, determining target event stream data, wherein the target event stream data includes data other than the detection event data in the plurality of event data;

[0012] Target detection is performed on the target event stream data to obtain information on missed objects;

[0013] The target detection result is obtained based on the target object information and the missed object information.

[0014] In one embodiment, each event data includes the event sampling time and the location information of multiple pixels;

[0015] Based on the image sampling time and event data of each frame, the detection event data corresponding to the detection box of each frame is determined from multiple event data, including:

[0016] Based on the image sampling time and event sampling time of each frame, multiple candidate event data corresponding to each frame are determined from multiple event data. The time interval between the image sampling time of the image and the event sampling time of each corresponding candidate event data is less than a preset threshold.

[0017] Based on the position of each detection box in each frame image and the position of multiple pixels in multiple candidate event data corresponding to each frame image, the detection event data corresponding to each detection box is determined from the multiple candidate event data.

[0018] In one embodiment, the detection event data corresponding to each detection box is determined from the multiple candidate event data based on the positions of each detection box in each frame image and the positions of multiple pixels in the multiple candidate event data corresponding to each frame image, including:

[0019] Acquire spatial calibration parameters between the image acquisition device and the event camera;

[0020] Based on the position and spatial calibration parameters of each detection box in each frame image, the event coordinate range corresponding to each detection box is obtained;

[0021] The candidate event data whose pixel coordinates fall within the event coordinate range among multiple candidate event data are determined as the detection event data.

[0022] In one embodiment, before acquiring multiple frames of images collected within a preset time period, and before the image sampling time of each frame, the method further includes:

[0023] The image acquisition device and the event camera are spatially calibrated using a preset dual-camera calibration algorithm to obtain spatial calibration parameters, including the spatial rotation matrix and spatial translation matrix of the event camera relative to the image acquisition device.

[0024] In one embodiment, a preset dual-camera calibration algorithm is used to spatially calibrate the image acquisition device and the event camera to obtain spatial calibration parameters, including:

[0025] Acquire the grayscale image captured by the event camera and the target image captured by the image acquisition device. The grayscale image and the target image are sampled at the same time.

[0026] A dual-camera calibration algorithm is used to spatially calibrate the grayscale image and the target image to obtain spatial calibration parameters.

[0027] In one embodiment, object detection is performed on each frame of the image to obtain the detection bounding box of the object in each frame of the image, including:

[0028] Target detection is performed on each frame of the image to obtain the center point coordinates of the detection box corresponding to each frame of the image, as well as the height and width of each detection box;

[0029] The detection bounding box of the object in each frame image is determined based on the position, height, and width of the center point.

[0030] In one embodiment, before acquiring images of the target device during its motion process captured in real time by the first device, the method further includes:

[0031] Perform time registration on the image acquisition device and the event camera.

[0032] A second aspect of this application provides a high dynamic target detection device, the device comprising:

[0033] The first acquisition module is used to acquire multiple frames of images acquired by the image acquisition device within a preset time period, as well as the image sampling time of each frame of images;

[0034] The second acquisition module is used to acquire multiple event data collected by the event camera within a preset time period;

[0035] The first detection module is used to perform target detection on each frame of the image to obtain target object information for each frame of the image, including the detection box of the object in the image.

[0036] The first determining module is used to determine the detection event data corresponding to the detection box of each frame image from multiple event data based on the image sampling time of each frame image and each event data;

[0037] The second determining module is used to determine the target event stream data based on each detection event data. The target event stream data includes data other than each detection event data from multiple event data.

[0038] The second detection module is used to perform target detection on the target event stream data and obtain information on missed objects.

[0039] The third determination module is used to obtain the target detection result based on the target object information and the missed object information.

[0040] In a third aspect of this application, a vehicle is provided, the vehicle including a memory and a processor, the memory storing a computer program, which, when executed by the processor, implements the high dynamic target detection method of the first aspect of this application.

[0041] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the high dynamic target detection method of the first aspect of this application.

[0042] The beneficial effects of the technical solutions provided in this application include at least the following:

[0043] The high dynamic range target detection method provided in this application acquires multiple frames of images captured by an image acquisition device within a preset time period, along with the image sampling time of each frame, and multiple event data collected by an event camera within the preset time period. Then, target detection is performed on each frame to obtain target object information, which includes detection bounding boxes of objects in the image. Next, based on the image sampling time and event data of each frame, detection event data corresponding to the detection bounding boxes of each frame is determined from the multiple event data. Target event stream data is then determined based on the detection event data, which includes data other than the detection event data from the multiple event data. Target detection is then performed on the target event stream data to obtain information on missed objects. Finally, the target detection result is obtained based on the target object information and the missed object information. This application improves the accuracy of target detection under high dynamic range by first performing target detection on the acquired images to obtain target object information, and then performing target detection on the event stream data excluding the event data corresponding to the target object information. Attached Figure Description

[0044] Figure 1 This is a schematic diagram of the internal structure of a vehicle-mounted terminal provided in an embodiment of this application;

[0045] Figure 2 The flowchart of a high dynamic target detection method provided in the embodiments of this application Figure One ;

[0046] Figure 3 The flowchart of a high dynamic target detection method provided in the embodiments of this application Figure Two ;

[0047] Figure 4 The flowchart of a high dynamic target detection method provided in the embodiments of this application Figure Three ;

[0048] Figure 5This is a structural diagram of a high dynamic target detection device provided in an embodiment of this application. Detailed Implementation

[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0050] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.

[0051] In addition, the use of “based on” or “according to” implies openness and inclusivity, because processes, steps, calculations or other actions “based on” or “according to” one or more conditions or values ​​can in practice be based on additional conditions or values ​​beyond those conditions.

[0052] With the rapid development of the automotive industry, autonomous vehicle technology has received widespread attention from academia and industry in recent years. Vehicle object detection is a challenging task in autonomous vehicle technology. It is an important application in the fields of autonomous vehicle technology and intelligent transportation systems. It plays a crucial role in autonomous driving technology. The purpose of vehicle object detection is to accurately locate the position of remaining vehicles in the surrounding environment to avoid accidents with other vehicles.

[0053] Currently, mainstream sensor solutions for autonomous vehicles all use ordinary cameras as visual perception sensors. However, the frame rate of ordinary cameras means they are prone to motion blur when sensing high-speed moving objects, leading to reduced target detection performance. Furthermore, ordinary cameras have a small dynamic range, making them ill-suited for sudden changes in lighting or excessively bright / dark scenes, resulting in overexposure or underexposure. For example, when entering or exiting a tunnel, the system may fail to properly detect targets near the tunnel entrance, further reducing target detection effectiveness.

[0054] To address the aforementioned problems, this application provides a high dynamic range target detection method. This method acquires multiple frames of images captured by an image acquisition device over a preset time period, along with the image sampling time of each frame, and multiple event data collected by an event camera over the same preset time period. Then, target detection is performed on each frame to obtain target object information, including detection bounding boxes for objects in the image. Next, based on the image sampling time and event data of each frame, detection event data corresponding to the detection bounding boxes of each frame is determined from the multiple event data. Target event stream data is then determined based on the detection event data, including data other than the detection event data from the multiple event data. Target detection is then performed on the target event stream data to obtain information on missed objects. Finally, the target detection result is obtained based on the target object information and the missed object information. This application improves the accuracy of target detection under high dynamic range by first performing target detection on the acquired images to obtain target object information, and then performing target detection on the event stream data excluding the event data corresponding to the target object information.

[0055] The high dynamic target detection method provided in this application can be executed by a vehicle, a server, or a server cluster, etc., and this application does not limit this. Specifically, when the executing entity is a vehicle, it is usually an in-vehicle terminal.

[0056] Figure 1 This is a schematic diagram of the internal structure of a vehicle-mounted terminal provided in an embodiment of this application. Figure 1 As shown, the vehicle-mounted terminal includes a processor and a memory connected via a system bus. The processor provides computing and control capabilities. The memory may include non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. These computer programs can be executed by the processor to implement the steps of the high-dynamic target detection method provided in the above embodiments. The internal memory provides a cached runtime environment for the operating system and computer programs in the non-volatile storage media.

[0057] Those skilled in the art will understand that Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the vehicle terminal to which the present application is applied. A specific vehicle terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0058] Based on the aforementioned execution entity, embodiments of this application provide a high-dynamic target detection method. For example... Figure 2 As shown, the method includes the following steps:

[0059] Step 201: Obtain multiple frames of images acquired by the image acquisition device within a preset time period, as well as the image sampling time of each frame.

[0060] The image acquisition device can be a camera or an RGB camera.

[0061] Step 202: Obtain multiple event data collected by the event camera within a preset time period.

[0062] Among them, the event camera is a new type of biomimetic vision sensor based on events. Unlike the working mechanism and output method of traditional cameras, the pixels of the event-based vision sensor can detect changes in light intensity individually, and output event information including position, time and polarity when the change exceeds a certain threshold. It has the advantages of low latency, high dynamic range and low power consumption, and is suitable for occasions with high-speed movement, large changes in lighting conditions or low energy consumption.

[0063] The event camera offers selectable data output modes, including grayscale image frame mode and event stream data mode. The coordinates of the event stream data are consistent with those of the grayscale image. When pixel changes occur due to object movement or lighting variations in the scene, a series of events are generated, and these events are output as an event stream.

[0064] Specifically, the event data in the event stream is discrete pulse data, formatted as an [n*4] matrix. Here, n represents the number of discrete pulses, 4 represents the dimension of the event data (x, y, a, t), and x, y, a, t represent the pixel coordinates x and y, the event polarity a, and the event occurrence time t, respectively. The event polarity indicates the nature of the event: positive polarity +1 indicates increased light intensity, and negative polarity -1 indicates decreased light intensity. The data stream dataset can be represented as Λ.

[0065] It should be noted that the event camera and image acquisition device are located in the same target device, and are used to collect images and event data of the target device during its movement, respectively. Specifically, the target device can be a vehicle, or it can also be a drone, aircraft, or automated inspection equipment, etc., which is not limited in this application.

[0066] Taking autonomous driving scenarios as an example, the image acquisition device and event camera are installed on the vehicle to acquire the vehicle's forward-facing view. The image acquisition device and event camera are installed at corresponding positions within the same forward-facing view of the vehicle, and the fields of view of the two cameras overlap. Simultaneously, the image acquisition device and event camera can also be installed at corresponding positions in other viewpoints of the target device to acquire other viewpoints of the target device; this application does not impose any limitations on this comparison.

[0067] Step 203: Perform target detection on each frame of the image to obtain the target object information of each frame of the image. The target object information includes the detection box of the object in the image.

[0068] Optionally, image-appropriate target detection algorithms, such as deep learning algorithms, can be used to detect targets in each frame of the image, obtaining information about the detected target objects in each frame. This target object information includes: detection bounding box information, the target object's category, confidence level, and the target object's identifier.

[0069] The target objects can include traffic participants, obstacles, and traffic facilities required by the autonomous driving system, such as people, motor vehicles, non-motor vehicles, traffic cones, and traffic lights. The algorithm outputs information such as the location, size, category, and confidence level of each target in each frame of the image.

[0070] It should be noted that by performing target detection on each frame of the image, the coordinates of the center point of the detection box corresponding to each frame of the image can be obtained, as well as the height and width of each detection box. Then, based on the position, height and width of the center point, the detection box of the object in each frame of the image can be determined.

[0071] For example, given a frame image P with a timestamp of sampling time Tp, and a target m is identified, the information data vector of target m is [u m v m w m h m , ...], where u m v m w represents the position of the center point of the target m detection box in the image pixels. m h m Let represent the width and height of the bounding box for target m, respectively. The "..." in the data vector represents other terms of interest, such as target category, confidence level, object identifier ID, etc. An image frame contains several targets; the dataset of target information contained in P is denoted as: Φp={[u...} m v m w m h m ,...], m=1, 2, 3...}.

[0072] Step 204: Based on the image sampling time and event data of each frame, determine the detection event data corresponding to the detection box of each frame from multiple event data.

[0073] It should be noted that there is a corresponding spatial transformation relationship between the image acquisition device and the event camera. Therefore, based on the image sampling time of each frame and each event data, the detection event data corresponding to the detection box of each frame can be determined from multiple event data.

[0074] Step 205: Determine the target event stream data based on each detection event data. The target event stream data includes data from multiple event data excluding each detection event data.

[0075] Step 206: Perform target detection on the target event stream data to obtain information on missed objects.

[0076] Optionally, target detection algorithms suitable for event data can be used to perform target detection on the target event stream data to obtain information on missed objects that were not detected based on image detection.

[0077] Since images contain more texture information, and current image and video perception algorithms are more mature than those for event cameras, target detection is mainly performed using image acquisition devices, with event cameras used as a supplement in extreme lighting conditions and motion blur situations.

[0078] Step 207: Obtain the target detection result based on the target object information and the missed object information.

[0079] The target detection results include: the category information of the detected target object, the detection box of the target object, the confidence level of the target object, and the identification information of the target object.

[0080] This application provides a high dynamic range target detection method. It acquires multiple frames of images captured by an image acquisition device over a preset time period, along with the image sampling time of each frame and multiple event data collected by an event camera over the same time period. Then, it performs target detection on each frame to obtain target object information, including detection bounding boxes for objects in the image. Next, based on the image sampling time and event data, it determines the detection event data corresponding to the detection bounding boxes in each frame from the multiple event data. Based on the detection event data, it determines target event stream data, which includes data from the multiple event data excluding the detection event data. It then performs target detection on the target event stream data to obtain information on missed objects. Finally, it obtains the target detection result based on the target object information and the missed object information. This application improves the accuracy of target detection under high dynamic range by first performing target detection on the acquired images to obtain target object information, and then performing target detection on the event stream data excluding the event data corresponding to the target object information.

[0081] Furthermore, due to the high dynamic range of the event camera, it can detect moving targets even when there are only minor changes in light. This allows it to detect moving targets even under extreme lighting conditions, such as sudden changes in light or scenes that are too bright or too dark, thereby improving the accuracy of target detection.

[0082] Optionally, the event data in step 202 above includes the event sampling time and the position information of multiple pixels. Correspondingly, such as... Figure 3 As shown, step 204 above, which determines the detection event data corresponding to the detection box of each frame image from multiple event data based on the image sampling time of each frame image and each event data, can be described as follows:

[0083] Step 301: Based on the image sampling time and event sampling time of each frame, determine multiple candidate event data corresponding to each frame from multiple event data.

[0084] In this case, the time interval between the image sampling time of the image and the event sampling time of each corresponding candidate event data is less than a preset threshold.

[0085] Step 302: Based on the position of each detection box in each frame image and the position of multiple pixels in the multiple candidate event data corresponding to each frame image, determine the detection event data corresponding to each detection box in the multiple candidate event data.

[0086] In other words, candidate event data corresponding to each image is first determined based on the image acquisition time of each frame. The time interval between the sampling time of these candidate event data and the image acquisition time is less than a preset threshold, which can be the time interval between two adjacent frames. Then, based on the coordinateless positions of the detection boxes in the image, the detection event data corresponding to each detection box is determined from the candidate event data. Since event stream data is continuous pulse data, and there is a sampling time interval between two adjacent images, when determining the corresponding event data based on the image sampling time, event data whose sampling time is less than the time interval between two adjacent frames can be identified as candidate event data. This ensures that the determined candidate event data is relatively complete.

[0087] Specifically, such as Figure 4 As shown, step 302 above, which determines the detection event data corresponding to each detection box from multiple candidate event data based on the position of each detection box in each frame image and the position of multiple pixels in multiple candidate event data corresponding to each frame image, can be as follows:

[0088] Step 401: Obtain the spatial calibration parameters between the image acquisition device and the event camera.

[0089] Step 402: Based on the position and spatial calibration parameters of each detection box in each frame image, obtain the event coordinate range corresponding to each detection box.

[0090] Step 403: Select the candidate event data whose pixel coordinates are within the event coordinate range from the multiple candidate event data as the detection event data.

[0091] For example, for the object detection box vertex dataset: Φ′p={[u m1 v m1 u m2 v m2 u m3 v m3 u m4 v m4 For a target m in the bounding box [u], m = 1, 2, 3…, the coordinates of the four vertices of the detection box are [u m1 v m1 u m2 v m2 u m3 v m3 u m4 v m4 Using the calibration parameters of the image acquisition device and the event camera, as well as the rotation matrix R and translation matrix T, calculate [u]. m1 v m1 u m2 v m2 u m3 v m3 u m4 v m4 The corresponding event camera coordinates [x] m1 y m1 x m2 y m2 x m3 y m3 x m4 y m4 A transformation operation is performed on each detection box to obtain the projected position ψ′p={[x m1 y m1 x m2 y m2 x m3 y m3 x m4 y m4 ],m=1,2,3…}。 Then, delete the event data that is within the range of the four vertices of the detection box and is within the time interval [Tp-Δt,Tp]. Where Δt is less than or equal to the time interval between two frames. For example, event data (x0, y0, a, t0), (x0, y0) is within the detection box of target 1 [x 11 y 11 x 12 y 12 x 13 y 13 x 14 y 14If the event stream data is within the range of [Tp-Δt, Tp], and t0∈[Tp-Δt, Tp], then delete (x0, y0, a, t0) from the event stream data Λ. The filtered event stream data is denoted as Λ′, which is the target event stream data.

[0092] Based on the target event stream data Λ′, target detection is performed using a target detection algorithm suitable for event camera data, which can detect targets that were missed by the image algorithm.

[0093] It should be noted that before performing step 201 above, spatial calibration can be performed on the image acquisition device and the event camera to obtain spatial calibration parameters, which include the spatial rotation matrix and spatial translation matrix of the event camera relative to the image acquisition device.

[0094] Optionally, since the pixel coordinates of the grayscale image are the same as the pixel coordinates of the event data, in order to facilitate spatial calibration, the grayscale image captured by the event camera and the target image captured by the image acquisition device can be acquired. The grayscale image and the target image are sampled at the same time. Then, the dual-camera calibration algorithm is used to spatially calibrate the grayscale image and the target image to obtain the spatial calibration parameters.

[0095] Meanwhile, after spatially calibrating the image acquisition device and the event camera, time registration is also required so that the image acquisition device and the event camera can share the time of the autonomous driving computing platform and maintain a synchronized time axis.

[0096] This application provides a high dynamic range target detection method. It acquires multiple frames of images captured by an image acquisition device over a preset time period, along with the image sampling time of each frame and multiple event data collected by an event camera over the same time period. Then, it performs target detection on each frame to obtain target object information, including detection bounding boxes for objects in the image. Next, based on the image sampling time and event data, it determines the detection event data corresponding to the detection bounding boxes in each frame from the multiple event data. Based on the detection event data, it determines target event stream data, which includes data from the multiple event data excluding the detection event data. It then performs target detection on the target event stream data to obtain information on missed objects. Finally, it obtains the target detection result based on the target object information and the missed object information. This application improves the accuracy of target detection under high dynamic range by first performing target detection on the acquired images to obtain target object information, and then performing target detection on the event stream data excluding the event data corresponding to the target object information.

[0097] Furthermore, due to the high dynamic range of the event camera, it can detect moving targets even when there are only minor changes in light. This allows it to detect moving targets even under extreme lighting conditions, such as sudden changes in light or scenes that are too bright or too dark, thereby improving the accuracy of target detection.

[0098] like Figure 5 As shown in the figure, this application embodiment also provides a high dynamic target detection device, which includes:

[0099] The first acquisition module 11 is used to acquire multiple frames of images acquired by the image acquisition device within a preset time period, as well as the image sampling time of each frame of images;

[0100] The second acquisition module 12 is used to acquire multiple event data collected by the event camera within a preset time period;

[0101] The first detection module 13 is used to perform target detection on each frame of the image to obtain target object information of each frame of the image, including the detection box of the object in the image.

[0102] The first determining module 14 is used to determine the detection event data corresponding to the detection box of each frame image from multiple event data based on the image sampling time of each frame image and each event data;

[0103] The second determining module 15 is used to determine target event stream data based on each detection event data. The target event stream data includes data other than each detection event data from multiple event data.

[0104] The second detection module 16 is used to perform target detection on the target event stream data and obtain information on missed objects.

[0105] The third determination module 17 is used to obtain the target detection result based on the target object information and the missed object information.

[0106] In one embodiment, the first determining module 14 is specifically used for:

[0107] Based on the image sampling time and event sampling time of each frame, multiple candidate event data corresponding to each frame are determined from multiple event data. The time interval between the image sampling time of the image and the event sampling time of each corresponding candidate event data is less than a preset threshold.

[0108] Based on the position of each detection box in each frame image and the position of multiple pixels in multiple candidate event data corresponding to each frame image, the detection event data corresponding to each detection box is determined from the multiple candidate event data.

[0109] In one embodiment, the first determining module 14 is specifically used for:

[0110] Acquire spatial calibration parameters between the image acquisition device and the event camera;

[0111] Based on the position and spatial calibration parameters of each detection box in each frame image, the event coordinate range corresponding to each detection box is obtained;

[0112] The candidate event data whose pixel coordinates fall within the event coordinate range among multiple candidate event data are determined as the detection event data.

[0113] In one embodiment, the device further includes a calibration module 18, which is used for:

[0114] The image acquisition device and the event camera are spatially calibrated using a preset dual-camera calibration algorithm to obtain spatial calibration parameters, including the spatial rotation matrix and spatial translation matrix of the event camera relative to the image acquisition device.

[0115] In one embodiment, the calibration module 18 is specifically used for:

[0116] Acquire the grayscale image captured by the event camera and the target image captured by the image acquisition device. The grayscale image and the target image are sampled at the same time.

[0117] A dual-camera calibration algorithm is used to spatially calibrate the grayscale image and the target image to obtain spatial calibration parameters.

[0118] In one embodiment, the first detection module 13 is specifically used for:

[0119] Target detection is performed on each frame of the image to obtain the center point coordinates of the detection box corresponding to each frame of the image, as well as the height and width of each detection box;

[0120] The detection bounding box of the object in each frame image is determined based on the position, height, and width of the center point.

[0121] In one embodiment, the apparatus further includes a registration module 19, which is used to perform time registration between the image acquisition device and the event camera.

[0122] The high dynamic target detection device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described in detail here.

[0123] Specific limitations regarding the high dynamic target detection device can be found in the limitations of the high dynamic target detection method described above, and will not be repeated here. Each module in the aforementioned high dynamic target detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the vehicle terminal in hardware form or independently of it, or stored in the memory of the vehicle terminal in software form, so that the processor can call and execute the corresponding operations of each module.

[0124] In another embodiment of this application, a vehicle is also provided, including a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the steps of the high dynamic target detection method as described in the embodiments of this application.

[0125] In another embodiment of this application, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the high dynamic target detection method as described in the embodiments of this application.

[0126] In another embodiment of this application, a computer program product is also provided, which includes computer instructions that, when executed on a high dynamic target detection device, cause the high dynamic target detection device to perform each step of the high dynamic target detection method in the method flow shown in the above method embodiment.

[0127] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs).

[0128] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0129] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A high dynamic target detection method, characterized in that, The method includes: Acquire multiple frames of images captured by the image acquisition device within a preset time period, as well as the image sampling time of each frame; Acquire multiple event data collected by the event camera during the preset time period; Target detection is performed on each frame of the image to obtain target object information for each frame of the image, wherein the target object information includes the detection bounding box of the object in the image; Based on the image sampling time of each frame and each event data, the detection event data corresponding to the detection box of each frame is determined from the plurality of event data; Target event stream data is determined based on each of the detected event data, and the target event stream data includes data other than each of the detected event data from the plurality of event data; Target detection is performed on the target event stream data to obtain information on missed objects; The target detection result is obtained based on the target object information and the missed object information; Each event data includes the event sampling time and the position information of multiple pixels; The step of determining the detection event data corresponding to the detection box of each frame image from the plurality of event data based on the image sampling time of each frame image and each of the event data includes: Based on the image sampling time and event sampling time of each frame, multiple candidate event data corresponding to each frame are determined from the multiple event data. The time interval between the image sampling time of the image and the event sampling time of each corresponding candidate event data is less than a preset threshold, which is the time interval between two adjacent frames. Based on the position of each detection box in each frame image and the position of multiple pixels in multiple candidate event data corresponding to each frame image, the detection event data corresponding to each detection box is determined from the multiple candidate event data.

2. The method according to claim 1, characterized in that, The step of determining the detection event data corresponding to each detection box from the multiple candidate event data based on the position of each detection box in each frame image and the position of multiple pixels in the multiple candidate event data corresponding to each frame image includes: Obtain the spatial calibration parameters between the image acquisition device and the event camera; Based on the position of each detection box in each frame image and the spatial calibration parameters, the event coordinate range corresponding to each detection box is obtained; The candidate event data whose pixel coordinates are located within the event coordinate range among the multiple candidate event data are determined as the detection event data.

3. The method according to claim 2, characterized in that, Before acquiring multiple frames of images captured by the image acquisition device within a preset time period, the method further includes: The image acquisition device and the event camera are spatially calibrated using a preset dual-camera calibration algorithm to obtain spatial calibration parameters, which include the spatial rotation matrix and spatial translation matrix of the event camera relative to the image acquisition device.

4. The method according to claim 3, characterized in that, The spatial calibration of the image acquisition device and the event camera using a preset dual-camera calibration algorithm to obtain the spatial calibration parameters includes: The grayscale image captured by the event camera and the target image captured by the image acquisition device are obtained, wherein the grayscale image and the target image are sampled at the same time; The dual-camera calibration algorithm is used to spatially calibrate the grayscale image and the target image to obtain the spatial calibration parameters.

5. The method according to any one of claims 1-4, characterized in that, Target detection is performed on each frame of the image to obtain the bounding boxes of objects in each frame, including: Target detection is performed on each frame of the image to obtain the center point coordinates of the detection box corresponding to each frame of the image, as well as the height and width of each detection box; The detection bounding box of the object in each frame image is determined based on the center point coordinates, the height, and the width.

6. The method according to claim 1, characterized in that, Before acquiring multiple frames of images captured by the image acquisition device within a preset time period, the method further includes: Time registration is performed on the image acquisition device and the event camera.

7. A high dynamic target detection device, characterized in that, The device includes: The first acquisition module is used to acquire multiple frames of images acquired by the image acquisition device within a preset time period, as well as the image sampling time of each frame of images; The second acquisition module is used to acquire multiple event data collected by the event camera during the preset time period; The first detection module is used to perform target detection on each frame of the image to obtain target object information of each frame of the image, wherein the target object information includes the detection box of the object in the image; The first determining module is used to determine the detection event data corresponding to the detection box of each frame image from the plurality of event data based on the image sampling time of each frame image and each event data; The second determining module is used to determine target event stream data based on each of the detected event data, wherein the target event stream data includes data other than each of the detected event data among the plurality of event data; The second detection module is used to perform target detection on the target event stream data to obtain information on missed objects. The third determining module is used to obtain the target detection result based on the target object information and the missed detection object information; Each event data includes the event sampling time and the position information of multiple pixels; the first determining module is specifically used for: Based on the image sampling time and event sampling time of each frame, multiple candidate event data corresponding to each frame are determined from the multiple event data. The time interval between the image sampling time of the image and the event sampling time of each candidate event data is less than a preset threshold, which is the time interval between two adjacent frames. Based on the position of each detection box in each frame and the position of multiple pixels in the multiple candidate event data corresponding to each frame, detection event data corresponding to each detection box is determined from the multiple candidate event data.

8. A vehicle, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, which, when executed by the processor, implements the high dynamic target detection method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the high dynamic target detection method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Driving environment sensing method and device, electronic equipment and storage medium

    CN113326820A

  • High-altitude parabolic object detection method and device based on mixed vision

    CN114170295A