Object Detection Method, Device, Medium and Equipment
Through the method of time synchronization and pixel-level data fusion, combined with millimeter-wave radar and camera information, the problem of insufficient accuracy and reliability of single sensor detection in smart traffic is solved, and target recognition and motion state detection are achieved with higher accuracy.
Patent Information
- Application Number
- CN202111137606.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-09-27
AI Technical Summary
The existing single sensor detection method is difficult to meet the accuracy and reliability requirements of target detection in smart transportation, especially the fusion detection technology based on monocular vision sensors and millimeter wave radars has the problem of low accuracy and reliability of target detection.
By obtaining time-synchronized millimeter-wave radar point cloud information and camera image information, data fusion is performed based on the pixel coordinate system, the target and its area of interest are determined using the image detection model, and the point cloud information in the area of interest is clustered to obtain the target's motion state detection information.
The accuracy and reliability of object detection are improved, the target loss problem caused by relying solely on the respective sensor algorithms is avoided, and the target recognition and motion state detection are achieved with higher accuracy.
Smart Images

Figure CN113887376B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent transportation technology, and particularly to a target detection method, device, medium and equipment. Background Art
[0002] With the development of intelligent transportation, the method of road target detection based on the data collected by a single sensor has been difficult to meet the growing application requirements of intelligent transportation. Therefore, the research focus has shifted to multi-sensor fusion detection technology to improve the perception ability of the traffic environment.
[0003] Among them, the multi-sensor fusion detection technology mainly focuses on the fusion of a monocular vision sensor (such as a camera) and a millimeter-wave radar. However, currently, this fusion detection technology simultaneously relies on the target detection algorithms of the monocular vision sensor and the millimeter-wave radar respectively, resulting in low accuracy and reliability of target detection. Summary of the Invention
[0004] In order to improve the accuracy and reliability of target detection, this application provides a target detection method, device, medium and equipment. The technical solutions are as follows:
[0005] In a first aspect, this application provides a target detection method, which includes:
[0006] Obtain the point cloud information collected by the millimeter-wave radar and the image information collected by the camera; the point cloud information and the image information meet the preset time synchronization condition;
[0007] Based on the pixel coordinate system, perform data fusion on the point cloud information and the image information to obtain a key image;
[0008] Input the key image into an image detection model for image detection processing to determine the target and the region of interest corresponding to the target in the key image;
[0009] Cluster the point cloud information within the region of interest to obtain the motion state detection information of the target.
[0010] Optionally, the obtaining the point cloud information collected by the millimeter-wave radar and the image information collected by the camera includes:
[0011] Obtain the original point cloud information collected by the millimeter-wave radar based on a first time period; the original point cloud information carries a first timestamp;
[0012] Obtain the original image information collected by the camera based on a second time period; the original image information carries a second timestamp;
[0013] Determine a processing period according to the first time period and the second time period;
[0014] Based on the processing period, the first timestamp, and the second timestamp, determine the point cloud information from the original point cloud information and determine the image information from the original image information.
[0015] Optionally, the determining the point cloud information from the original point cloud information and determining the image information from the original image information based on the processing period, the first timestamp, and the second timestamp includes:
[0016] Respectively determine target original point cloud information and target original image information that are in the same processing period from the original point cloud information and the original image information;
[0017] Determine a timestamp difference according to the first timestamp of the target original point cloud information and the second timestamp of the target original image information;
[0018] When the timestamp difference meets a preset condition, use the target original point cloud information as the point cloud information and use the target original image information as the image information.
[0019] Optionally, the method further includes:
[0020] Determine a first target obtained based on a first key image corresponding to the current processing period and a first region of interest corresponding to the first target;
[0021] Determine a second target obtained based on a second key image corresponding to the previous processing period and a second region of interest corresponding to the second target;
[0022] Match the first target and the second target according to the first region of interest and the second region of interest to obtain a third target that appears successively in the second key image and the first key image;
[0023] Determine motion state detection information of the third target in the current processing period and the previous processing period to track the third target.
[0024] Optionally, the method further includes:
[0025] Input the motion state detection information of the target into an extended Kalman filter to optimize and estimate the motion state detection information to obtain optimized motion state detection information of the target.
[0026] Optionally, the data fusion of the point cloud information and the image information based on the pixel coordinate system to obtain a key image includes:
[0027] Determine the spatial position information of the millimeter-wave radar and the camera;
[0028] Determine the coordinate system conversion relationship based on the spatial position information;
[0029] According to the coordinate system conversion relationship, determine the point cloud data of the point cloud information in the pixel coordinate system;
[0030] Determine the image data of the image information in the pixel coordinate system;
[0031] Perform data fusion on the point cloud data and the image data to obtain a key image.
[0032] Optionally, the clustering of the point cloud information in the region of interest to obtain the motion state detection information of the target includes:
[0033] Obtain a preset neighborhood radius threshold and density threshold;
[0034] According to the neighborhood radius threshold and the density threshold, perform density clustering on the point cloud information in the region of interest to obtain one or more clusters, where the cluster is a set of sampling points;
[0035] According to the point cloud information corresponding to the sampling points in the cluster, calculate the motion state detection information of the target, and the motion state detection information includes at least one of position information, size information, and row direction information.
[0036] In a second aspect, the present application provides a target detection device, and the device includes:
[0037] An information acquisition module, configured to acquire point cloud information collected by a millimeter-wave radar and acquire image information collected by a camera; the point cloud information and the image information meet a preset time synchronization condition;
[0038] An information fusion module, configured to perform data fusion on the point cloud information and the image information based on a pixel coordinate system to obtain a key image;
[0039] An image detection module, configured to input the key image into an image detection model, perform image detection processing, and determine a target and a region of interest corresponding to the target in the key image;
[0040] A motion state detection module, configured to cluster the point cloud information in the region of interest to obtain the motion state detection information of the target.
[0041] In a third aspect, the present application provides a computer-readable storage medium storing at least one instruction or at least one program segment, which is loaded and executed by a processor to implement an object detection method as described in the first aspect.
[0042] In a fourth aspect, the present application provides a computer device including a processor and a memory, where the memory stores at least one instruction or at least one program segment, and the at least one instruction or at least one program segment is loaded and executed by the processor to implement an object detection method as described in the first aspect.
[0043] The object detection method, device, medium, and equipment provided by the present application have the following technical effects:
[0044] The technical solution provided by the present application first obtains point cloud information and image information that are synchronized in time. The point cloud information can be collected by a millimeter-wave radar, and the image information can be collected by a camera. The synchronization in time ensures the accuracy and reliability of subsequent detection results. Secondly, the point cloud information and the image information are mapped into a pixel coordinate system to perform pixel-level data fusion of the point cloud information and the image information, obtaining a fused key image. Performing pixel-level data fusion first can retain richer collected data, provide more information for object detection, thus making fuller use of the original information and improving the detection accuracy and reliability. Then, the key image is input into an image detection model based on deep learning to detect the object that conforms to the model detection category and the region of interest of the object in the key image. This is the recognition and detection of the object and its category. Using the image detection model based on deep learning to identify and detect the key image has better real-time performance and higher detection accuracy, further effectively improving the accuracy of recognizing and detecting the object and its category. Finally, clustering calculation is performed on the point cloud information included in the region of interest to obtain the motion state detection information of the object, that is, the detection of the object's motion state is completed based on the previous model detection result, making the detected object position, size, heading, etc. information more accurate.
[0045] The technical solution provided by the present application performs object recognition and detection on the key image after pixel-level data fusion and clustering calculation on the region of interest of the object in the key image to complete the detection of the object's motion state, without relying on the respective object detection algorithms of the camera and the millimeter-wave radar, and can avoid problems such as object loss caused by mismatches between the two parts of object detection results.
[0046] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings
[0047] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0048] Figure 1 It is a schematic diagram of the implementation environment of a target detection method provided by an embodiment of the present application;
[0049] Figure 2 It is a schematic flowchart of a target detection method provided by an embodiment of the present application;
[0050] Figure 3 It is a schematic diagram of the positions of sampling points in a polar coordinate system and a millimeter-wave radar coordinate system provided by an embodiment of the present application;
[0051] Figure 4 It is a schematic flowchart of a method for obtaining time-synchronized point cloud information and image information provided by an embodiment of the present application;
[0052] Figure 5 It is another schematic flowchart of a method for obtaining time-synchronized point cloud information and image information provided by an embodiment of the present application;
[0053] Figure 6 It is a schematic flowchart of a method for matching original point cloud information and original image information based on timestamps provided by an embodiment of the present application;
[0054] Figure 7 It is a schematic flowchart of a method for data fusion based on a pixel coordinate system provided by an embodiment of the present application;
[0055] Figure 8 It is a schematic diagram of the spatial positions of a camera and a millimeter-wave radar provided by an embodiment of the present application;
[0056] Figure 9 It is a schematic diagram of the working postures of a camera and a millimeter-wave radar provided by an embodiment of the present application;
[0057] Figure 10 It is a schematic diagram of the conversion of a radar coordinate system to a world coordinate system provided by an embodiment of the present application;
[0058] Figure 11 It is a schematic diagram of the conversion of a world coordinate system to a pixel coordinate system provided by an embodiment of the present application;
[0059] Figure 12 It is a schematic flowchart of a method for detection based on density clustering provided by an embodiment of the present application;
[0060] Figure 13 It is a schematic diagram of the effect of object detection provided by an embodiment of the present application;
[0061] Figure 14 It is a schematic diagram of the process for continuously detecting an object provided by an embodiment of the present application;
[0062] Figure 15 It is a schematic diagram of the effect of continuously detecting an object provided by an embodiment of the present application;
[0063] Figure 16 It is a schematic diagram of the process for continuously detecting and optimizing an object provided by an embodiment of the present application;
[0064] Figure 17 It is a schematic diagram of an object detection device provided by an embodiment of the present application;
[0065] Figure 18 It is a schematic diagram of the hardware structure of a device for implementing an object detection method provided by an embodiment of the present application. Detailed implementation manners
[0066] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics.
[0067] The solution provided by the embodiment of the present application involves technologies such as computer vision (CV) and deep learning (DL) in artificial intelligence.
[0068] Among them, computer vision technology (CV) is a science that studies how to enable machines to "see". Further speaking, it refers to using cameras and computers to replace human eyes for machine vision such as target recognition, tracking, and measurement, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, the theories and technologies related to computer vision research attempt to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content recognition, behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0069] Among them, deep learning (DL) is a major research direction in the field of machine learning (ML). It is introduced into machine learning to make it closer to the original goal - artificial intelligence. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during the learning process is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analysis and learning like humans, and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed the previous related technologies. Deep learning has achieved many results in search technology, data mining, machine learning, machine translation, natural language processing, multimedia learning, speech, recommendation and personalization technology, and other related fields. Deep learning enables machines to imitate human activities such as audiovisual and thinking, solves many complex pattern recognition problems, and makes great progress in artificial intelligence-related technologies.
[0070] Intelligent transportation makes full use of new generation information technologies such as the Internet of Things, spatial perception, cloud computing, and mobile Internet in the entire transportation field, and comprehensively applies theories and tools such as traffic science, systematic methods, artificial intelligence, and knowledge mining. With the goals of comprehensive perception, deep integration, active service, and scientific decision-making, by building a real-time dynamic information service system, deeply mining transportation-related data, forming a problem analysis model, realizing the improvement of the industry's resource allocation optimization ability, public decision-making ability, industry management ability, and public service ability, promoting the safer, more efficient, more convenient, more economical, more environmentally friendly, and more comfortable operation and development of transportation, and driving the transformation and upgrading of transportation-related industries.
[0071] The solution provided by the embodiments of this application relates to technologies such as the vehicle networking and autonomous driving in intelligent transportation.
[0072] Among them, the concept of vehicle networking originates from the Internet of Things, that is, the Internet of Vehicles. It takes the vehicles in motion as the information perception objects and, with the help of the new generation of information and communication technologies, realizes the network connection between the vehicle and X (i.e., vehicle-to-vehicle, vehicle-to-person, vehicle-to-road, vehicle-to-service platform), improves the overall intelligent driving level of the vehicle, provides users with safe, comfortable, intelligent, and efficient driving experiences and traffic services, and at the same time improves the traffic operation efficiency and the intelligent level of social traffic services. Among them, the communication between vehicles refers to the information exchange and sharing between vehicles, including vehicle state information such as vehicle position and driving speed, which can be used to judge the traffic flow conditions on the road; the communication between vehicle and road refers to the information exchange between the vehicle and the road with the help of the fixed communication facilities on the ground road, which is used to monitor the road surface conditions and guide the vehicle to select the best driving route.
[0073] Among them, the acquisition of environmental information and intelligent decision-making control, which are key links in autonomous driving, rely on a series of high-tech technologies such as sensor technology, image recognition technology, electronics and computer technology, and control technology. With the progress of machine vision (such as 3D camera technology), pattern recognition software (such as optical character recognition programs), and lidar systems (which have combined global positioning technology and spatial data), in-vehicle computers can control the driving of vehicles by combining machine vision, sensor data, and spatial data.
[0074] The solution provided by the embodiments of this application can also be deployed in the cloud, and cloud technology and the like are also involved.
[0075] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. It can also be understood as the general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system. Therefore, cloud technology needs to be supported by cloud computing. Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services according to needs. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a basic capability provider of cloud computing, a cloud computing resource pool platform will be established, abbreviated as the cloud platform, generally referred to as Infrastructure as a Service (IaaS). Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (which can be virtual machines, including operating systems), storage devices, and network devices.
[0076] To improve the accuracy and reliability of object detection, the embodiments of the present application provide an object detection method, device, medium, and equipment. The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end.
[0077] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0078] To facilitate the understanding of the technical solutions described in the embodiments of the present application and the technical effects produced thereby, the embodiments of the present application explain the relevant professional terms involved:
[0079] Point cloud: Point Cloud, is a massive point set that expresses the spatial distribution of the target and the surface characteristics of the target under the same spatial reference system. After obtaining the spatial coordinates of each sampling point on the object surface, the set of points obtained is called a "point cloud".
[0080] Clustering: The process of dividing a set of physical or abstract objects into multiple classes composed of similar objects is called clustering. The clusters generated by clustering are a set of data objects, and these objects are similar to each other within the same cluster and different from those in other clusters.
[0081] Density clustering: The density-based clustering algorithm assumes that the clustering structure can be determined by the tightness of the sample distribution, and performs clustering based on the density of the dataset in the spatial distribution. That is, as long as the sample density in a region is greater than a certain threshold, it is assigned to the cluster similar to it.
[0082] DBSCAN: Density-Based Spatial Clustering of Applications with Noise, a density-based clustering method with noise, is based on a set of neighborhood parameters to describe the tightness of the sample distribution. Compared with the partitioning-based clustering method and the hierarchical clustering method, the DBSCAN algorithm defines a cluster as the largest set of density-connected samples, can divide regions with high enough density into clusters, does not require a given number of clusters, and can discover clusters of any shape in a spatial dataset with noise.
[0083] Kalman filtering: Kalman filtering, is an algorithm that uses a linear system state equation to optimally estimate the system state through system input and output observation data.
[0084] Extended Kalman Filter (EKF) is an extended form of the standard Kalman filter in the non - linear case and is a highly efficient recursive filter. Its basic idea is to linearize the non - linear system by using Taylor series expansion and then use the Kalman filter framework to filter the signal.
[0085] Hungarian algorithm: It is a combinatorial optimization algorithm that solves the task assignment problem in polynomial time and has promoted the subsequent primal - dual method.
[0086] Please refer to Figure 1 , which is a schematic diagram of the implementation environment of an object detection method provided by an embodiment of the present application. As Figure 1 shown, the implementation environment may at least include a millimeter - wave radar, a camera, and an object detection device. The camera and the millimeter - wave radar may be installed on a vehicle or on the roadside of a road. The object detection device may be correspondingly installed on the roadside of the road or on the vehicle, or may be placed in a remote server. The camera, the millimeter - wave radar, and the object detection device may be connected through a wireless network or a wired network.
[0087] The following introduces an object detection method provided by the present application. Figure 2 is a flowchart of an object detection method provided by an embodiment of the present application. The present application provides the method operation steps as described in the embodiment or the flowchart, but based on routine or non - creative labor, there may be more or fewer operation steps. The step order listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or server product executes, it may execute in the order of the method shown in the embodiment or the drawings, or execute in parallel (for example, in an environment of parallel processors or multi - thread processing). Please refer to Figure 2 , an object detection method provided by an embodiment of the present application may include the following steps:
[0088] S210: Obtain the point cloud information collected by the millimeter - wave radar and obtain the image information collected by the camera; the point cloud information and the image information meet the preset time synchronization condition.
[0089] It is understandable that in both the vehicle networking field and the autonomous driving field, the environmental perception ability is a key link. Only by perceiving the surrounding environment based on the data collected by sensors can a basis for decision-making and control be provided. In the autonomous driving field, object detection is an essential part of the autonomous driving environmental perception system, and the fusion perception based on millimeter-wave radar and camera has become the research focus. In the driving environment, taking vehicle perception as the research object, by detecting targets such as vehicles, pedestrians, obstacles, and road signs in the surrounding driving environment, an information basis for vehicle control and decision-making is provided. In the vehicle networking field, the perception system at the roadside can be used to monitor the transportation status in real time and accurately, realize the intelligent management of traffic. At the same time, the perception results at the roadside can be transmitted to the vehicle, liberating the computing power of part of the in-vehicle perception system and providing more accurate and comprehensive environmental information for the driving vehicle.
[0090] In the embodiments of the present application, the millimeter-wave radar and the camera can be installed on the vehicle, or can be installed on roadside gantries, street lights and other facilities. The method provided by the embodiments of the present application can be executed by an in-vehicle device, equipment or system, or can be executed by a module or device with computing power configured on the roadside of the road, or can be executed by a remote server or a cloud server. The present application does not make any limitations in this regard. In the embodiments of the present application, the millimeter-wave radar is used to collect data on the traffic road to obtain the original point cloud information. The millimeter-wave radar is a radar that operates in the millimeter-wave band. Usually, millimeter waves refer to the frequency domain of 30 GHz to 300 GHz (wavelength is 1 mm to 10 mm). The in-vehicle radar will also have an allocated exclusive frequency band. The wavelength of millimeter waves is between centimeter waves and light waves. Therefore, millimeter waves have both the advantages of microwave guidance and optoelectronic guidance, and its seeker has the characteristics of small volume, light weight and high spatial resolution. In addition, the millimeter-wave seeker has strong ability to penetrate fog, smoke and dust. The millimeter-wave radar mainly obtains information such as the distance, speed and angle of an object by sending electromagnetic waves and receiving echoes. The position data of the sampling points on the object surface can be the coordinates in the polar coordinate system with the millimeter-wave radar as the origin, or the coordinates in the three-dimensional space coordinate system with the millimeter-wave radar as the origin. In the embodiments of the present application, the visual sensor (that is, the camera) is used to collect data on the traffic road to obtain the original image information. The camera can be a monocular visual sensor, a binocular visual sensor, a panoramic vision or an infrared camera, etc. The collected image can cover all the road environment information within the line of sight, such as lane lines, traffic signs, traffic lights, pedestrians and vehicles, etc.
[0091] Figure 3 An exemplary relative position between a millimeter-wave radar and a sampling point is shown, such as Figure 3As shown, the millimeter-wave radar is at Or, the sampling point is point P, the radial distance between the millimeter-wave radar and the sampling point is R, and the horizontal azimuth angle between the sampling point and the millimeter-wave radar is α. In the polar coordinate system, the position coordinates of the sampling point can be expressed as (R, α). In addition to using the polar coordinate system, a self-coordinate system XrOrYr with the position of the millimeter-wave radar itself as the origin can also be used, where the Xr axis points forward to the millimeter wave, that is, the direction of the detected road, and the Yr axis points to the left side of the millimeter-wave radar. Therefore, the coordinates (xr, yr) of the sampling point in the millimeter-wave radar coordinate system XrOrYr can be as shown in formula (1) and formula (2):
[0092] xr = R × cosα; (1)
[0093] yr = R × sinα; (2)
[0094] In the embodiment of the present application, the point cloud information and the image information are a set of data to be fused after time synchronization, and the time synchronization is the basic guarantee for subsequent data fusion and target detection. There are differences in the data acquisition frequencies of different sensors and different acquisition time starting points. One way is to synchronize the time of the millimeter-wave radar and the video detector, so that the two describe the information at the same moment, which is convenient for subsequent processing; another way is to screen and match the original point cloud information and the original image information to obtain the synchronized point cloud information and the image information.
[0095] Figure 4 Shows a schematic flow diagram of an exemplary method for obtaining synchronized point cloud information and image information. Specifically, as Figure 4 shown, the step S210 may include the following steps:
[0096] S211: Obtain the original point cloud information collected by the millimeter-wave radar based on the first time period; the original point cloud information carries a first timestamp.
[0097] S213: Obtain the original image information collected by the camera based on the second time period; the original image information carries a second timestamp.
[0098] S215: Determine the processing period according to the first time period and the second time period.
[0099] For example, the processing period can be the least common multiple period of the first time period and the second time period.
[0100] S217: Based on the processing period, the first timestamp, and the second timestamp, determine the point cloud information from the original point cloud information and determine the image information from the original image information.
[0101] In a feasible implementation manner, as Figure 5As shown, the step S217 may specifically include the following steps:
[0102] S2171: Corresponding target original point cloud information and target original image information within the same processing cycle are respectively determined from the original point cloud information and the original image information.
[0103] S2173: According to the first timestamp of the target original point cloud information and the second timestamp of the original image information, a timestamp difference is determined.
[0104] S2175: When the timestamp difference meets a preset condition, the target original point cloud information is used as the point cloud information and the target original image information is used as the image information.
[0105] Exemplarily, as Figure 6 shown, the point cloud information within the same processing cycle includes data for four cycles (T0, T1, T2, and T3), and the image information includes at least one frame of image ( Figure 6 only one frame is shown here for illustration). The timestamp differences between this frame of image and the point cloud information within the four cycles are respectively calculated. If the timestamp difference between this frame of image and the point cloud information corresponding to the T0 cycle is the smallest and does not exceed the preset threshold, then these two are used as a set of information to be fused.
[0106] It can be understood that the so-called time synchronization does not limit the timestamps of the point cloud information and the image information to be exactly the same, that is, absolute time synchronization. In practice, a certain difference in timestamps between the point cloud information and the image information that are used as a set of data to be fused is allowed.
[0107] In an exemplary embodiment of the present application, when performing time synchronization processing, a multi-threaded method can be adopted. The first thread uses a vector (container) to receive the original point cloud information of the millimeter-wave radar. Specifically, when receiving the first frame of message in each cycle, the local timestamp is added, and the point cloud information of the sampling points is periodically pushed into the container. At the same time, the data size in the container needs to be detected. When the data volume is greater than 1000, the container needs to be cleared for protection measures; the second thread is used to save one frame of image information in real time; the third thread compares the timestamps according to the two types of saved data, and extracts the two types of data with a timestamp difference not exceeding 10 ms within the same processing cycle for subsequent processing. After the subsequent processing is completed, the previous data in the radar data container can be cleared. Feasibly, preprocessing can also be performed on the point cloud information and the image information after time synchronization, such as filtering out the outlier points in the point cloud information and using spatial filtering technology to remove the redundant noise in the image.
[0108] In the above embodiments, based on the acquisition timestamps of the original data, the problem of clock asynchronization between different sensors can be solved, providing effective and accurate data for subsequent data fusion and target detection, ensuring the accuracy of the detection results. At the same time, the point cloud information and image information are determined from the original point cloud information and original image information for subsequent fusion, so that the fused data can retain more original and underlying data, avoiding the loss of valid data.
[0109] S230: Based on the pixel coordinate system, perform data fusion on the point cloud information and image information to obtain a key image.
[0110] It can be understood that for fusion perception based on a millimeter-wave radar and a camera, in addition to synchronization in time to achieve a matching data rate, spatial synchronization is also required, that is, unifying the two types of data into the same coordinate system. Only after spatio-temporal synchronization can the data fusion give full play to the dual advantages of the millimeter-wave radar and the camera, be able to obtain environmental information more fully, provide more information for target detection, and improve the accuracy and reliability of the detection.
[0111] In an embodiment of the present application, before performing data fusion on the point cloud information and image information based on the pixel coordinate system, both types of information need to be converted to the pixel coordinate system. As Figure 7 shown, the step of performing data fusion on the point cloud information and the image information based on the pixel coordinate system to obtain a key image may specifically include the following steps:
[0112] S231: Determine the spatial position information of the millimeter-wave radar and the camera.
[0113] Specifically, the spatial position information can be determined according to the relative height, spacing, etc. of the installation of the millimeter-wave radar and the camera.
[0114] S233: Determine the coordinate system conversion relationship based on the spatial position information.
[0115] The coordinate system conversion relationship represents the conversion coefficient when the position information is mapped from one coordinate system to another. In addition, the coordinate system conversion relationship is also related to the setting of the origin and direction of each coordinate system.
[0116] S235: According to the coordinate system conversion relationship, determine the point cloud data of the point cloud information in the pixel coordinate system.
[0117] The point cloud information may include position data in the polar coordinate system with the millimeter-wave radar as the origin, or may include position data in the millimeter-wave radar coordinate system with the millimeter-wave radar as the origin, while the point cloud data is coordinate data in the pixel coordinate system. The pixel coordinate system is a coordinate system u-v established with the upper left corner of the image as the origin and in pixels. The abscissa u and ordinate v of the pixel are respectively the column number and row number where the pixel point is located in its image array.
[0118] In order to accurately project a point in the space detected by the millimeter-wave radar onto a point in the image plane captured by the camera, it involves the conversion between a total of five coordinate systems: the millimeter-wave radar coordinate system, the world coordinate system, the camera coordinate system, the image coordinate system, and the pixel coordinate system. Through computer vision theory and the camera model, the conversion relationships between the four coordinate systems related to the camera can be obtained, and with the help of the Zhang Zhengyou camera calibration method, the internal parameters and external parameter values of the camera in the conversion relationships can be obtained, thus realizing the conversion from the millimeter-wave radar coordinate system to the pixel coordinate system.
[0119] S237: Determine the image data of the image information in the pixel coordinate system.
[0120] The image information may include the position information of the pixel points in the image coordinate system, while the image data is the position information of the pixel points in the pixel coordinate system. Both the pixel coordinate system and the image coordinate system are on the imaging plane, but their respective origins and measurement units are different. The origin of the image coordinate system is the intersection point of the camera optical axis and the imaging plane, usually the midpoint of the imaging plane or called the principal point. The unit of the image coordinate system is mm, which belongs to a physical unit, while the unit of the pixel coordinate system is pixel.
[0121] S239: Perform data fusion on the point cloud data and the image data to obtain the key image.
[0122] Based on the above operations, the millimeter-wave radar point cloud information in space can be matched to the visual image, and on this basis, the motion state information of the sampling points in the point cloud information can be output.
[0123] Feasible. The point cloud data and image data are synthesized using a fusion algorithm to obtain a consistent interpretation and description of pixel points. Among them, the fusion algorithm can have robustness and parallel processing capabilities, and can also meet other requirements such as the operation speed and accuracy of the algorithm, the interface performance with the previous spatio-temporal synchronization processing and subsequent target recognition, etc. Generally, non-linear mathematical methods with fault tolerance, self-adaptability, associative memory, and parallel processing capabilities can be used as fusion methods. The common methods can basically be divided into two categories: random and artificial intelligence. Among them, the random category can include weighted average method, Kalman filtering method, multi-Bayesian estimation method, etc., and the artificial intelligence category can include fuzzy logic reasoning, neural network method. The embodiments of the present application do not limit this.
[0124] In the above embodiments, the spatio-temporal synchronization of point cloud information and image information is performed, that is, the two types of data are unified into the same coordinate system. The pixel-level data fusion after spatio-temporal synchronization can give full play to the dual advantages of millimeter-wave radar and camera. In addition to being able to obtain environmental information more fully, it can also eliminate the uncertain factors of a single end, provide more accurate observation results and comprehensive information, and thus can improve the accuracy and reliability of detection.
[0125] Figures 8 to 9 Fig. shows a schematic diagram of the spatial positions of a millimeter-wave radar and a camera in an embodiment, where Figure 8 Fig. shows a millimeter-wave radar and a camera installed on the crossbar of a roadside gantry, which are horizontally installed in parallel. The height positions of the two devices relative to the ground and the relative distance between the two devices can be as Figure 8 shown; where Figure 9 Fig. shows the direction angles when the millimeter-wave radar and the camera perform data acquisition. Based on this, Figure 10 Fig. shows a schematic diagram of converting from the millimeter-wave radar coordinate system to the world coordinate system. Taking the coordinates (xr, yr) of the sampling point in the millimeter-wave radar coordinate system in the previous embodiment as an example, the values of xr and yr can be as shown in formulas (1) and (2), which will not be elaborated here. The world coordinate system is the absolute coordinate system of the objective three-dimensional world, also known as the objective coordinate system. In Figures 10 - 11 a world coordinate system XwYwZw - Ow is established with the position of the camera as the origin. h is the installation height of the millimeter-wave radar. The coordinates (xw, yw, zw) of the sampling point of this millimeter-wave radar in the world coordinate system XwYwZw - Ow can be as shown in formulas (3) to (5):
[0126] xw = -R×sinα; (3)
[0127]
[0128] zw = -h; (5)
[0129] Figure 11 Shows a schematic diagram of the conversion from the world coordinate system to the pixel coordinate system. In the conversion from the world coordinate system to the pixel coordinate system, as Figure 10 shown, where α is the pitch angle of the camera coordinate system XcYcZc - Ow relative to the world coordinate system, that is, the rotation angle around the Xw axis, and β is the yaw angle of the camera coordinate system relative to the world coordinate system, that is, the rotation angle around the Zw axis. Denote s1 = sinα, s2 = sinβ, c1 = cosα, c2 = cosβ, f u is the vertical focal length of the camera, f v is the horizontal focal length of the camera, c u and c v represent the position of the image center point in the pixel coordinate system, u and v are the coordinate values in the pixel coordinate system, P i is the position of the sampling point in the pixel coordinate system, P g is the position of the sampling point in the world coordinate system. Based on the above formulas (3) to (4), the coordinate components xg, yg, zg of P g in the world coordinate system are numerically respectively equivalent to xw, yw, zw. Indicates the conversion relationship from the world coordinate system to the pixel coordinate system, and the expressions of each item can be shown as follows:
[0130]
[0131]
[0132]
[0133]
[0134] After conversion, the coordinates of each sampling point in the pixel coordinate system can be obtained.
[0135] S250: Input the key image into the image detection model, perform image detection processing, and determine the target and the region of interest corresponding to the target in the key image.
[0136] In the embodiments of the present application, an image detection model is used to detect, identify, and classify targets, and the region of interest (ROI) of the targets is delineated. Among them, the targets to be detected may include, but are not limited to, vehicles, pedestrians, road signs, road lines, obstacles, etc. The categories of the targets can be further subdivided. For example, vehicles can be divided into trucks, sedans, motorcycles, etc. Among them, the region of interest (ROI) refers to an image area selected from an image in the field of image processing. This area is the focus of image analysis. Using ROI to delineate the target can reduce processing time and increase accuracy. Feasibly, a deep learning algorithm is adopted in the post-fusion image processing part. Compared with traditional image processing, it has better real-time performance and higher detection accuracy. Deep learning has a powerful feature learning ability, the extracted features are more abundant, the expression ability is stronger, and the results of target detection, identification, and classification are more accurate.
[0137] In the embodiments of the present application, the model can be trained with the fused historical key image samples and the annotation information of the targets in the samples to obtain the required image detection model. Exemplarily, a YOLO5 (You Only Look Once) model based on deep learning is adopted. The network of the YOLO5 model mainly consists of three main components:
[0138] (1) Backbone: A convolutional neural network that aggregates and forms image features at different image granularities.
[0139] (2) Neck: A series of network layers that mix and combine image features and transfer the image features to the prediction layer.
[0140] (3) Output: Predicts the image features, generates bounding boxes, and predicts the categories.
[0141] Furthermore, specific samples and the target categories to be detected can be selectively and specifically selected for model training according to the application scenarios.
[0142] S270: Cluster the point cloud information within the region of interest to obtain the motion state detection information of the target.
[0143] In the embodiments of the present application, by virtue of the role of the image detection model in this process, the category information of the target is obtained and the region of interest of the target is delineated, and the motion state detection information of the target is calculated by clustering based on the point cloud information within the region of interest. The motion state detection information may include, but is not limited to, position information, size information, heading information, etc. Image detection and clustering calculations jointly complete the process of predicting the category of the target and detecting the motion state.
[0144] In the embodiments of the present application, multiple targets can be detected simultaneously to obtain the category information and motion state detection information of the multiple targets. Hereinafter, taking a single target as an example for illustration, the embodiments of the present application will not be elaborated.
[0145] In an embodiment of the present application, as Figure 12 shown, the step of clustering the point cloud information in the region of interest to obtain the motion state detection information of the target may specifically include the following steps:
[0146] S271: Obtain a preset neighborhood radius threshold and density threshold.
[0147] S273: According to the neighborhood radius threshold and density threshold, perform density clustering on the point cloud information in the region of interest to obtain one or more clusters, where a cluster is a set of sampling points.
[0148] S275: According to the point cloud information corresponding to the sampling points in the cluster, calculate the motion state detection information of the target, and the motion state detection information includes at least one of position information, size information, and heading information.
[0149] In a feasible implementation manner, a density clustering algorithm of DBSCAN (Density-Based Spatial Clustering of Applications with Noise) can be used. This algorithm can divide regions with high enough density into "clusters" and can discover clusters of any shape in data with "noise".
[0150] In the above embodiments, the motion state detection information of the target is calculated by using a density clustering method. From the perspective of sample density, the connectivity between points in the region of interest of the target is examined and continuously expanded until the final clustering clusters are obtained. Clusters of any shape can be discovered and it is not sensitive to noise data, thereby effectively and accurately calculating the motion state of the target.
[0151] Figure 13 Shows a schematic diagram of the results of target category detection and motion state detection for a key image. As Figure 13 shown, the method provided by the embodiments of the present application can effectively identify that the front targets are a truck and a car. The aspect ratios of the predicted regions of interest of the targets are 0.83 and 0.98 respectively, the driving speeds are 17.76 m / s (meters per second) and 14.52 m / s respectively, and the radial distances are 64.70 meters and 29.07 meters respectively. In addition, Figure 13 the road boundary line can also be detected in
[0152] In an embodiment of the present application, the target can be continuously tracked and detected. Figure 14The flowchart shows a method for continuously detecting and tracking a target. As shown in Figure 14 the method provided by the embodiments of the present application may further include the following steps:
[0153] S310: Obtain a first key image corresponding to the current processing cycle and a second key image corresponding to the previous processing cycle.
[0154] S330: Obtain a first target and a first region of interest corresponding to the first target based on the first key image corresponding to the current processing cycle.
[0155] S350: Obtain a second target and a second region of interest corresponding to the second target based on the second key image corresponding to the previous processing cycle.
[0156] S370: Match the first target and the second target according to the first region of interest and the second region of interest to obtain a third target, and the third target appears successively in the second key image and the first key image.
[0157] S390: Determine the motion state detection information of the third target in the current processing cycle and the previous processing cycle to track the third target.
[0158] In a feasible implementation, the targets in two processing cycles can be marked based on the IOU (Intersection over Union) between the predicted bounding box and the true bounding box in the region of interest. Then, the Hungarian algorithm can be used to perform maximum matching on multiple targets in two processing cycles to determine the same target that appears successively in two processing cycles, that is, the consistency detection of the target is completed. Further, a threshold can be set for the IOU in the target detection result, and the consistency detection for targets such as road lines can be omitted. Exemplarily, taking the detection result in Figure 13 as the detection result of the previous processing cycle, Figure 15 and the detection result in
[0159] as the detection result of the current processing cycle, the method provided by the embodiments of the present application can track and detect a car.
[0160] In another embodiment of the present application, the detection results of adjacent processing cycles can be used for recursive optimization of the results. Exemplarily, the motion state detection information of the target can be input into an extended Kalman filter to optimize and estimate the motion state detection information, and the optimized detection information of the target's motion state can be obtained. That is, the motion state detection information is used as a measurement value, and the measurement value is adjusted and optimized based on the Kalman gain. Figure 16 FIG. shows a schematic flow chart of an object detection method that combines object consistency detection and extended Kalman filtering. Specifically, taking the previous processing cycle and the current processing cycle as examples, by performing consistency detection on the detection results of the previous processing cycle (including the category detection information and motion state detection information of the object) and the detection results of the current processing cycle, the objects that can be continuously detected are determined. In the extended Kalman filter, the optimized value of the detection results of the previous processing cycle is used to optimize and estimate the measurement value of the detection results of the current processing cycle, and the optimized value of the detection results of the object in the current processing cycle is obtained and saved as a reference for optimizing the detection results of the next processing cycle. This recursive loop continues until the object disappears. Specifically, if the previous processing cycle is the first processing cycle after startup, the measurement value of the detection results of the previous processing cycle can be directly used as the optimized value and saved. In the above embodiment, through the extended Kalman filter, the detection results are optimized, which can make the detection more accurate. Moreover, the extended Kalman filter requires less memory and has a fast calculation speed, making it more suitable for real-time situations and the needs of embedded devices.
[0161] An embodiment of the present application further provides an object detection device 1700, as Figure 17 shown, the device 1700 may include:
[0162] An information acquisition module 1710, configured to acquire point cloud information collected by a millimeter-wave radar and image information collected by a camera; the point cloud information and the image information meet a preset time synchronization condition;
[0163] An information fusion module 1720, configured to perform data fusion on the point cloud information and the image information based on a pixel coordinate system to obtain a key image;
[0164] An image detection module 1730, configured to input the key image into an image detection model for image detection processing to determine an object and an interested region corresponding to the object in the key image;
[0165] A motion state detection module 1740, configured to cluster the point cloud information in the interested region to obtain the motion state detection information of the object.
[0166] In an embodiment of the present application, the information acquisition module 1710 may include:
[0167] An original point cloud information acquisition unit, configured to acquire original point cloud information collected by the millimeter-wave radar based on a first time period; the original point cloud information carries a first timestamp;
[0168] An original image information acquisition unit, configured to acquire original image information collected by the camera based on a second time period; the original image information carries a second timestamp;
[0169] A processing period determination unit, configured to determine a processing period according to the first time period and the second time period;
[0170] An information determination unit, configured to determine the point cloud information from the original point cloud information and determine the image information from the original image information based on the processing period, the first timestamp, and the second timestamp.
[0171] In an embodiment of the present application, the information determination unit may include:
[0172] An information division sub-unit, configured to respectively determine target original point cloud information and target original image information in the same processing period from the original point cloud information and the original image information;
[0173] A timestamp difference calculation sub-unit, configured to determine a timestamp difference according to the first timestamp of the target original point cloud information and the second timestamp of the target original image information;
[0174] An information matching sub-unit, configured to use the target original point cloud information as the point cloud information and the target original image information as the image information when the timestamp difference meets a preset condition.
[0175] In an embodiment of the present application, the apparatus 1700 may further include:
[0176] A current processing period image detection unit, configured to determine a first target obtained based on a first key image corresponding to the current processing period and a first region of interest corresponding to the first target;
[0177] A previous processing period image detection unit, configured to determine a second target obtained based on a second key image corresponding to the previous processing period and a second region of interest corresponding to the second target;
[0178] A target matching unit, configured to match the first target and the second target according to the first region of interest and the second region of interest to obtain a third target, and the third target appears successively in the second key image and the first key image;
[0179] A continuous tracking unit, configured to determine the motion state detection information of the third target in the current processing cycle and the previous processing cycle, so as to track the third target.
[0180] In an embodiment of the present application, the device 1700 may further include:
[0181] An optimization unit, configured to input the motion state detection information of the target into an extended Kalman filter, optimize and estimate the motion state detection information, and obtain the optimized motion state detection information of the target.
[0182] In an embodiment of the present application, the information fusion module 1720 may include:
[0183] A spatial position determination unit, configured to determine the spatial position information of the millimeter-wave radar and the camera;
[0184] A coordinate system relationship determination unit, configured to determine the coordinate system conversion relationship based on the spatial position information;
[0185] A first coordinate conversion unit, configured to determine the point cloud data of the point cloud information in the pixel coordinate system according to the coordinate system conversion relationship;
[0186] A second coordinate conversion unit, configured to determine the image data of the image information in the pixel coordinate system;
[0187] A data fusion unit, configured to perform data fusion on the point cloud data and the image data to obtain a key image.
[0188] In an embodiment of the present application, the motion state detection module 1740 may include:
[0189] A preset information acquisition unit, configured to acquire a preset neighborhood radius threshold and density threshold;
[0190] A density clustering unit, configured to perform density clustering on the point cloud information in the region of interest according to the neighborhood radius threshold and the density threshold to obtain one or more clusters, where the cluster is a set of sampling points;
[0191] A calculation unit, configured to calculate the motion state detection information of the target according to the point cloud information corresponding to the sampling points in the cluster, where the motion state detection information at least includes one of position information, size information, and row direction information.
[0192] It should be noted that when the device provided in the above embodiments realizes its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0193] An embodiment of the present application provides a computer device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement a target detection method as provided in the above method embodiment.
[0194] Figure 18 A schematic diagram of the hardware structure of a device for implementing a target detection method provided in an embodiment of the present application is shown. The device may participate in constituting or include the device or system provided in the embodiment of the present application. As Figure 18 shown, the device 10 may include one or more processors (the processors may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA, shown as 1002a, 1002b,..., 1002n in the figure), a memory 1004 for storing data, and a transmission device 1006 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 18 the structure shown is only for illustration and does not limit the structure of the above electronic device. For example, the device 10 may further include more or fewer components than those Figure 18 shown, or have a different configuration from that Figure 18 shown.
[0195] It should be noted that the above one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the device 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is used for processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0196] The memory 1004 can be used to store software programs and modules of application software, such as the program instruction / data storage device corresponding to a target detection method as described in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 1004, that is, to implement the above-mentioned method. The memory 1004 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 1004 may further include a memory remotely disposed relative to the processor, and these remote memories can be connected to the device 10 through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and their combinations.
[0197] The transmission device 1006 is used to receive or send data via a network. Specific examples of the above network may include the wireless network provided by the communication provider of the device 10. In one instance, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission device 1006 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0198] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the device 10 (or mobile device).
[0199] The embodiments of the present application also provide a computer-readable storage medium, which can be disposed in a server to store at least one instruction or at least one segment of program related to a target detection method in the method embodiments. The at least one instruction or the at least one segment of program is loaded and executed by the processor to implement the target detection method provided by the above method embodiments.
[0200] Optionally, in this embodiment, the above storage medium may be located in at least one of multiple network servers in a computer network. Optionally, in this embodiment, the above storage medium may include but are not limited to: USB flash drive, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk, or optical disc, etc., various media that can store program codes.
[0201] An embodiment of the present invention also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes a target detection method provided in the above various alternative embodiments.
[0202] It should be noted that: the above sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of the present application have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0203] The various embodiments in this application are all described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the apparatus, device, and storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0204] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0205] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A target detection method, characterized in that, The method includes: Obtaining point cloud information collected by a millimeter-wave radar and obtaining image information collected by a camera; the point cloud information and the image information meet a preset time synchronization condition; Based on a pixel coordinate system, performing data fusion on the point cloud information and the image information to obtain a key image; Inputting the key image into an image detection model, performing image detection processing, and determining a target and an interested region corresponding to the target in the key image; Clustering the point cloud information within the interested region to obtain motion state detection information of the target; The method further includes: Determining a first target obtained based on a first key image corresponding to a current processing cycle and a first interested region corresponding to the first target; Determining a second target obtained based on a second key image corresponding to a previous processing cycle and a second interested region corresponding to the second target; Matching the first target and the second target according to the first interested region and the second interested region to obtain a third target, and the third target appears successively in the second key image and the first key image; Determining motion state detection information of the third target in the current processing cycle and the previous processing cycle to track the third target.
2. The object detection method according to claim 1, wherein The obtaining the point cloud information collected by the millimeter-wave radar and obtaining the image information collected by the camera includes: Obtaining original point cloud information collected by the millimeter-wave radar based on a first time period; the original point cloud information carries a first timestamp; Obtaining original image information collected by the camera based on a second time period; the original image information carries a second timestamp; Determining a processing cycle according to the first time period and the second time period; Based on the processing cycle, the first timestamp, and the second timestamp, determining the point cloud information from the original point cloud information and determining the image information from the original image information.
3. The object detection method according to claim 2, wherein The determining the point cloud information from the original point cloud information and determining the image information from the original image information based on the processing cycle, the first timestamp, and the second timestamp includes: Correspondingly determining target original point cloud information and target original image information that are within the same processing cycle from the original point cloud information and the original image information respectively; Determining a timestamp difference according to the first timestamp of the target original point cloud information and the second timestamp of the target original image information; When the timestamp difference meets a preset condition, using the target original point cloud information as the point cloud information and using the target original image information as the image information.
4. The object detection method according to claim 1, wherein The method further includes: Inputting the motion state detection information of the target into an extended Kalman filter to perform optimized estimation on the motion state detection information and obtain optimized motion state detection information of the target.
5. The object detection method according to claim 1, characterized in that, The performing data fusion on the point cloud information and the image information based on a pixel coordinate system to obtain a key image includes: Determining spatial position information of the millimeter-wave radar and the camera; Determine the coordinate system conversion relationship based on the spatial position information; Determine the point cloud data of the point cloud information in the pixel coordinate system according to the coordinate system conversion relationship; Determine the image data of the image information in the pixel coordinate system; Perform data fusion on the point cloud data and the image data to obtain a key image.
6. The object detection method according to claim 1, wherein, The clustering of the point cloud information in the region of interest to obtain the motion state detection information of the target includes: Obtain a preset neighborhood radius threshold and density threshold; Perform density clustering on the point cloud information in the region of interest according to the neighborhood radius threshold and the density threshold to obtain one or more clusters, where the cluster is a set of sampling points; Calculate the motion state detection information of the target according to the point cloud information corresponding to the sampling points in the cluster, and the motion state detection information includes at least one of position information, size information, and row direction information.
7. A target detection device, characterized in that, The device includes: An information acquisition module, configured to acquire the point cloud information collected by the millimeter-wave radar and the image information collected by the camera; the point cloud information and the image information meet the preset time synchronization condition; An information fusion module, configured to perform data fusion on the point cloud information and the image information based on the pixel coordinate system to obtain a key image; An image detection module, configured to input the key image into an image detection model for image detection processing to determine the target and the region of interest corresponding to the target in the key image; A motion state detection module, configured to cluster the point cloud information in the region of interest to obtain the motion state detection information of the target; The device further includes: A current processing cycle image detection unit, configured to determine a first target obtained based on a first key image corresponding to the current processing cycle and a first region of interest corresponding to the first target; A previous processing cycle image detection unit, configured to determine a second target obtained based on a second key image corresponding to the previous processing cycle and a second region of interest corresponding to the second target; A target matching unit, configured to match the first target and the second target according to the first region of interest and the second region of interest to obtain a third target, and the third target appears successively in the second key image and the first key image; A continuous tracking unit, configured to determine the motion state detection information of the third target in the current processing cycle and the previous processing cycle to track the third target.
8. A computer-readable storage medium, characterized in that, At least one instruction or at least one program segment is stored in the computer-readable storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement a target detection method according to any one of claims 1 to 6.
9. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement a target detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Detection method, device, equipment and system and storage medium
CN111856468A
Road running state detection method and system based on radar and video fusion
CN112946628A
Obstacle detection method based on millimeter wave radar and camera information fusion
CN113156421A