Real-time target detection method fused with big data analysis
By combining lidar and high-resolution cameras to collect data on urban roads, and using the target image detection model and PointPillars algorithm for target detection, the problem of the contradiction between the target detection algorithm and embedded equipment computing power in the existing technology is solved, and more efficient target detection and tracking effects are achieved.
Patent Information
- Application Number
- CN202510023747.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, the visual object detection algorithm is inconsistent with the computing power of embedded devices, resulting in poor target detection and tracking effects in complex backgrounds and multi-target environments on urban roads.
A real-time object detection method integrating big data analysis is adopted to collect three-dimensional point cloud data and two-dimensional image data through lidar and high-resolution cameras, combined with a preset target image detection model and a PointPillars algorithm that introduces attention mechanisms, target recognition and detection are carried out, and the effective fusion of multimodal data is achieved through external parameter matrix projection and IOU interleaving method.
It improves the accuracy and robustness of object detection, alleviates the problem of limited computing resources of embedded devices, enhances real-time processing capabilities, and adapts to the needs of low-computing environments.
Smart Images

Figure CN119964120A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and in particular relates to a real-time target detection method integrating big data analysis. Background Art
[0002] On urban roads, target detection technology can be used to realize real-time detection and identification of vehicles, pedestrians, bicycles and other targets in urban road scenes, thereby improving the efficiency and accuracy of traffic management, reducing the occurrence of traffic accidents such as traffic accidents and trampling accidents, alleviating road congestion, and improving the operating efficiency and safety of urban traffic, providing important data support and decision-making basis for traffic management and traffic safety.
[0003] At present, the commonly used sensors for target detection include RGB cameras and LiDAR. Among them, the camera has a fast detection speed and can capture rich texture information of the target to be detected, but it is difficult to directly measure the shape and position of the object. At the same time, as a passive sensor, it is easily affected by changes in the intensity of ambient light. Compared with RGB cameras, LiDAR detects the surrounding environment through lasers, can accurately measure the distance and shape of objects, and has strong robustness to light changes, but even high-resolution LiDARs collect relatively sparse point cloud data. However, target detection on urban roads still faces the following challenges: there are not only various types of vehicles on urban roads, but also other targets such as pedestrians, bicycles, and motorcycles. The shape, color, texture and other features of these targets vary greatly, which increases the complexity of target detection and tracking; the background on urban roads is often dynamically changing, including factors such as lighting changes, shadows, reflections, rain and fog, which will affect the visibility and differentiation of targets and reduce the effect of target detection and tracking.
[0004] Therefore, it is necessary to propose a real-time target detection method that integrates big data analysis to solve the problem that the visual target detection algorithm used in the prior art is inconsistent with the low computing power of embedded devices.
[0005] The above information disclosed in this background technology is only used to increase the understanding of the background technology of the present invention and therefore, it may include information that does not constitute the prior art known to ordinary technicians in this field. Summary of the invention
[0006] The purpose of the present invention is to provide a real-time target detection method integrating big data analysis to solve the problems raised in the above-mentioned background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A real-time target detection method integrating big data analysis, comprising:
[0009] The deployed LiDAR and high-resolution cameras collect raw data of the road ahead, including 3D point cloud data and 2D image data;
[0010] Using a preset target image detection model to perform target recognition on the two-dimensional image data to obtain an image recognition result;
[0011] Use the PointPillars algorithm that introduces an attention mechanism to perform target detection on the three-dimensional point cloud data to obtain a point cloud detection result;
[0012] Projecting the point cloud detection result onto the two-dimensional image data using the external parameter matrix obtained by calibrating the laser radar and the high-resolution camera to obtain a projected image;
[0013] According to the projection image combined with the image recognition result, the target is matched using the IOU intersection-union method to obtain the final target detection result.
[0014] Preferably, the three-dimensional point cloud data and the two-dimensional image data are collected by a laser radar and a high-resolution camera respectively;
[0015] Convert the two-dimensional image data into a data format and adjust it to a preset size to obtain the pre-processed two-dimensional image data;
[0016] The three-dimensional point cloud data is optimized by using voxel grid filtering technology, thereby achieving downsampling of the three-dimensional point cloud data and obtaining the preprocessed three-dimensional point cloud data.
[0017] Preferably, the target image detection model is constructed based on the YOLOv5n network model by integrating the GhostNet lightweight structure and introducing a parameter-free attention mechanism;
[0018] GhostNet is used to replace the CBS module of the YOLOv5n network model to reduce the number of model parameters;
[0019] And GhostNet is introduced into the C3 module of the YOLOv5n network model to share weights between multiple convolutional layers and reduce the computational complexity of the model;
[0020] A parameter-free attention mechanism is introduced to measure the linear separability between the target neuron and its surrounding neurons in the model, and to identify the most distinctive neurons in each channel;
[0021] The constructed target image detection model is used to perform target recognition on the two-dimensional image data to obtain the image recognition result.
[0022] Preferably, the attention mechanism is introduced into the two-dimensional feature extraction network of the PointPillars algorithm to obtain a new PointPillars algorithm;
[0023] Use the new PointPillars algorithm to perform target detection on the three-dimensional point cloud data to obtain the point cloud detection result;
[0024] When performing target detection, the Softplus activation function is used to capture negative information and details of small targets.
[0025] Preferably, the coordinate systems O1 and O2 of the laser radar and the high-resolution camera are respectively established, and time synchronization of the three-dimensional point cloud data and the two-dimensional image data is achieved based on a timestamp;
[0026] The coordinate system O1 and the coordinate system O2 are transformed into the same coordinate system by rotation and translation to obtain the external parameter matrix;
[0027] The point cloud detection result is projected into the two-dimensional image data using the external parameter matrix to obtain the projected image.
[0028] Preferably, the comparison is performed by comparing the relationship between the projection image and the image recognition result based on the IOU intersection-union method to generate a comparison result;
[0029] The final target detection result is generated according to the comparison result and the actual situation, and is output in real time according to a preset output format.
[0030] Preferably, the multi-target tracking algorithm is used to track the real-time target, and the target is dynamically updated according to the real-time updated three-dimensional point cloud data and the two-dimensional image data to generate a travel trajectory;
[0031] Analyze multiple travel trajectories through big data analysis to discover potential movement patterns or abnormal behaviors;
[0032] The big data platform is used to store and analyze all the above data, and predictions and decisions are made based on the acquired environmental information.
[0033] Preferably, the data processing process is accelerated using a GPU or FPGA hardware platform, or multi-threaded parallel computing is used to simultaneously analyze the three-dimensional point cloud data and the two-dimensional image data.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] The present invention uses a preset target image detection model to perform target recognition on two-dimensional image data, and can quickly obtain target information in the image. The PointPillars algorithm with an attention mechanism is used to process three-dimensional point cloud data, thereby improving the detection accuracy and robustness of point cloud data. Based on the calibration results of the laser radar and the camera, the three-dimensional point cloud detection results are projected into the two-dimensional image plane, ensuring the effective fusion of multimodal data. Through the target matching method based on the IOU intersection and union, the image and point cloud are precisely aligned, thereby obtaining more accurate target detection results. This method effectively alleviates the problem of limited computing resources of embedded devices. Through algorithm optimization and data fusion, the real-time processing capability is greatly improved while ensuring detection accuracy, and it adapts to the needs of low computing power environments.
[0036] The above summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present invention will be readily apparent by reference to the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a flow chart of the real-time target detection method integrating big data analysis of the present invention. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0039] Embodiment 1:
[0040] See also Figure 1 As shown, a real-time target detection method integrating big data analysis includes:
[0041] The deployed LiDAR and high-resolution cameras collect raw data of the road ahead, including 3D point cloud data and 2D image data;
[0042] Use a preset target image detection model to perform target recognition on the two-dimensional image data to obtain an image recognition result;
[0043] Use the PointPillars algorithm with attention mechanism to detect targets on 3D point cloud data and obtain point cloud detection results.
[0044] The point cloud detection results are projected onto the two-dimensional image data using the external parameter matrix obtained by calibrating the laser radar and the high-resolution camera to obtain the projected image;
[0045] According to the projection image combined with the image recognition result, the target matching is completed using the IOU intersection-union method to obtain the final target detection result.
[0046] Collect 3D point cloud data and 2D image data through LiDAR and high-resolution camera respectively;
[0047] Convert the two-dimensional image data to a pre-set size to obtain pre-processed two-dimensional image data.
[0048] Voxel grid filtering technology is used to optimize the three-dimensional point cloud data, thereby achieving downsampling of the three-dimensional point cloud data and obtaining preprocessed three-dimensional point cloud data.
[0049] Before using voxel grid filtering technology to optimize the 3D point cloud data, the 3D point cloud data is segmented based on the road surface plane fitting algorithm. The specific steps are as follows:
[0050] (1) Divide the front ground into N sub-planes along the x-axis;
[0051] (2) First select the N with the lowest height value low point clouds, take the average value of the point cloud height N low Add it to the seed threshold, and the point with the smallest height is used as the seed point;
[0052] (3) Using the seed point set E = H obtained in the second step 3×3 , the linear model formula for pavement plane model fitting is as follows:
[0053] ax+by+dz+s=0
[0054] The model parameters (a, b, d) are determined by the covariance matrix F∈H of the initial seed point set 3×3 Solving, the calculation formula of matrix F is as follows, where Represents the mean height of all point clouds:
[0055]
[0056] (4) The initial plane model is obtained through the third step. The vertical distances of other laser point clouds in the sub-plane to the initial plane are calculated. If the distances are below the threshold, they are road planes. Otherwise, they are obstacles. The laser point clouds judged as road surface points are used as new initial seed point sets. After optimization and iteration, point cloud segmentation is performed on N sub-planes in turn to obtain the point cloud segmentation of the entire area in front of the vehicle.
[0057] Based on the YOLOv5n network model, a target image detection model is constructed by integrating the GhostNet lightweight structure and introducing a parameter-free attention mechanism;
[0058] GhostNet is a lightweight convolutional neural network architecture that reduces computational complexity and the number of parameters mainly through innovative Ghost modules and depth-separable convolutions. It is particularly suitable for devices and application scenarios with limited computing resources, such as mobile devices, embedded devices, IoT devices, etc.
[0059] GhostNet is used to replace the CBS module of the YOLOv5n network model to reduce the number of model parameters;
[0060] The CBS module consists of convolution, batch normalization, and Swish activation functions, which are used to process and transform feature maps, providing stable training and stronger expression capabilities.
[0061] And GhostNet is introduced in the C3 module of the YOLOv5n network model to share weights between multiple convolutional layers and reduce the computational complexity of the model;
[0062] The C3 module is mainly used for the separation and fusion of feature maps. It combines residual connection and CSP techniques to improve the network's feature extraction capability and training efficiency.
[0063] A parameter-free attention mechanism is introduced to measure the linear separability between the target neuron and its surrounding neurons in the model, and to identify the most distinctive neurons in each channel;
[0064] The constructed target image detection model is used to perform target recognition on the two-dimensional image data to obtain the image recognition result.
[0065] For an input image X∈R C×H×W Where C, H, and W represent the number of channels, input image height, and input image width, respectively. The operation of any convolutional layer for generating n feature maps can be written as the following equation.
[0066] Y=X*f+b
[0067] In the formula, * represents the convolution operation, b represents the bias, Y is the output, and f represents a convolution layer with n convolution kernels of size k×k. The number of computations required in traditional convolution can be calculated as n×c×h×w×k×k, which is usually hundreds of thousands because the number of convolution layers n and the number of channels c are very large.
[0068] The computational complexity of the convolution operation proposed by GhostNet is significantly reduced compared to the traditional convolution. Assume that the input is set to X∈R C×H×W, first perform conventional convolution on the input feature vector and output p feature maps Y′.
[0069] Y′∈X*f′
[0070] In order to obtain redundant (np) feature maps, each feature map in the p outputs is subjected to (q-1) linear operations, so that each feature map generates (q-1) similar feature maps through mapping, where q = n / p. Finally, the obtained p feature maps are stacked with the (q-1) feature maps obtained by linear operation to obtain n feature maps. The linear operation y i ' ,j It can be expressed as follows, so the ratio r of GhostNet to the conventional convolution operation can be expressed as follows.
[0071]
[0072] In the formula, y i ′ represents the i-th feature map in y′, φ i,j Where j represents the jth linear operation on the i-th feature map. When the convolution kernel size of the linear operation is the same as the convolution kernel size of the conventional convolution operation, we can get:
[0073]
[0074] Finally, we can get the estimated value q of the ratio of the computational amount of GhostNet to the conventional convolution operation. This value is much smaller than c, so the model has the characteristic of lightweight.
[0075] The attention mechanism is introduced into the two-dimensional feature extraction network of the PointPillars algorithm to obtain a new PointPillars algorithm;
[0076] PointPillars is a 3D object detection algorithm based on deep learning. Its core idea is to convert point cloud data into "pillars" and process them with CNN.
[0077] Use the new PointPillars algorithm to perform target detection on 3D point cloud data and obtain point cloud detection results;
[0078] When performing target detection, the Softplus activation function is used to capture negative information and details of small targets.
[0079] Embodiment 2:
[0080] See also Figure 1As shown, this embodiment is basically the same as the above embodiment, except that the coordinate systems O1 and O2 of the laser radar and the high-resolution camera are established respectively, and the time synchronization of the three-dimensional point cloud data and the two-dimensional image data is achieved based on the timestamp;
[0081] The coordinate system O1 and the coordinate system O2 are transformed into the same coordinate system by rotation and translation to obtain the external parameter matrix;
[0082] The point cloud detection results are projected into the two-dimensional image data using the external parameter matrix to obtain the projected image.
[0083] Use the IOU intersection-union method to compare the relationship between the projection image and the image recognition result to generate a comparison result;
[0084] IOU (Intersection over Union) is a commonly used indicator to measure the degree of overlap between two objects or regions, especially in the field of computer vision, to evaluate the accuracy of target detection models. Specifically, IOU calculates the ratio of the area of the intersection of two regions (such as target detection boxes) to the area of their union.
[0085] The IOU calculation formula of the projection image and the image recognition result is as follows:
[0086]
[0087] In the formula, S L is the projected image area, S Z is the area of image recognition result;
[0088] 0.5≤IOU≤1, indicating that the coverage areas of the two target images are from the same target, and the category and location information detected by the high-resolution camera are used as the output information of the target;
[0089] 0<IOU<0.5, although the real target exists, due to the complexity and variability of the scene in the actual environment, there may be different weather conditions such as sunny days, rainy days, snowy days, and foggy days. At this time, it is necessary to consider the possibility of missed detection by the high-resolution camera due to the influence of weather. Finally, the detection category of the high-resolution camera and the orientation information detected by the lidar are used as the target output information;
[0090] When IOU=0, the coverage areas of the two target images do not overlap, S Z ≠0 and S L = 0, the surface laser radar missed detection, and finally output high-resolution camera information; S L ≠0 and S Z =0, indicating that the high-resolution camera missed the detection and outputs the lidar information; S L = 0 and S Z=0, indicating that the target does not exist and the judgment ends.
[0091] The final target detection result is generated based on the comparison result and the actual situation, and is output in real time according to the preset output format.
[0092] Use multi-target tracking algorithms to track real-time targets, and dynamically update targets based on real-time updated 3D point cloud data and 2D image data to generate travel trajectories;
[0093] Analyze multiple travel trajectories through big data analysis to discover potential movement patterns or abnormal behaviors;
[0094] The big data platform is used to store and analyze all the above data, and predictions and decisions are made based on the acquired environmental information.
[0095] Use GPU or FPGA hardware platform to accelerate the above data processing process, or use multi-threaded parallel computing to analyze 3D point cloud data and 2D image data at the same time.
[0096] As can be seen from the above, the present invention uses a preset target image detection model to perform target recognition on two-dimensional image data, which can quickly obtain target information in the image, and processes three-dimensional point cloud data by introducing the PointPillars algorithm with an attention mechanism, thereby improving the detection accuracy and robustness of point cloud data; based on the calibration results of the laser radar and the camera, the three-dimensional point cloud detection results are projected into the two-dimensional image plane, ensuring the effective fusion of multimodal data; through the target matching method based on the IOU intersection and union, the image and point cloud are accurately aligned, thereby obtaining more accurate target detection results. This method effectively alleviates the problem of limited computing resources of embedded devices, and through algorithm optimization and data fusion, it greatly improves the real-time processing capability while ensuring detection accuracy, and adapts to the needs of low computing power environments.
[0097] Embodiment 3:
[0098] The embodiment of the present invention also provides a computer-readable storage medium, on which a program of a real-time target detection method integrating big data analysis as described above is stored. When the program is executed by a processor, each process of the above-mentioned real-time target detection method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. Among them, the computer-readable storage medium, such as a read-only memory (Read-Only Memory, referred to as ROM), a random access memory (Random Access Memory, referred to as RAM), a disk or an optical disk, etc.
[0099] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0100] In the drawings of the embodiments disclosed in the present invention, only the structures involved in the embodiments disclosed in the present invention are involved, and other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of the present invention can be combined with each other.
[0101] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0102] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A real-time target detection method integrating big data analysis, characterized in that: include: The deployed LiDAR and high-resolution cameras collect raw data of the road ahead, including 3D point cloud data and 2D image data; Using a preset target image detection model to perform target recognition on the two-dimensional image data to obtain an image recognition result; Use the PointPillars algorithm that introduces an attention mechanism to perform target detection on the three-dimensional point cloud data to obtain a point cloud detection result; Projecting the point cloud detection result onto the two-dimensional image data using the external parameter matrix obtained by calibrating the laser radar and the high-resolution camera to obtain a projected image; According to the projection image combined with the image recognition result, the target is matched using the IOU intersection-union method to obtain the final target detection result.
2. The real-time target detection method integrating big data analysis according to claim 1 is characterized in that: The raw data collected by the deployed laser radar and high-resolution camera in front of the road, including three-dimensional point cloud data and two-dimensional image data, includes: Collect the three-dimensional point cloud data and the two-dimensional image data respectively by using a laser radar and a high-resolution camera; Convert the two-dimensional image data into a data format and adjust it to a preset size to obtain the pre-processed two-dimensional image data; The three-dimensional point cloud data is optimized by using voxel grid filtering technology, thereby achieving downsampling of the three-dimensional point cloud data and obtaining the preprocessed three-dimensional point cloud data.
3. The real-time target detection method integrating big data analysis according to claim 2 is characterized in that: The using a preset target image detection model to perform target recognition on the two-dimensional image data to obtain an image recognition result includes: Based on the YOLOv5n network model, the target image detection model is constructed by integrating the GhostNet lightweight structure and introducing a parameter-free attention mechanism; GhostNet is used to replace the CBS module of the YOLOv5n network model to reduce the number of model parameters; And GhostNet is introduced into the C3 module of the YOLOv5n network model to share weights between multiple convolutional layers and reduce the computational complexity of the model; A parameter-free attention mechanism is introduced to measure the linear separability between the target neuron and its surrounding neurons in the model, and to identify the most distinctive neurons in each channel; The constructed target image detection model is used to perform target recognition on the two-dimensional image data to obtain the image recognition result.
4. The real-time target detection method integrating big data analysis according to claim 3 is characterized in that: The PointPillars algorithm using the attention mechanism is used to perform target detection on the three-dimensional point cloud data to obtain a point cloud detection result, including: Introducing the attention mechanism into the two-dimensional feature extraction network of the PointPillars algorithm to obtain the new PointPillars algorithm; Use the new PointPillars algorithm to perform target detection on the three-dimensional point cloud data to obtain the point cloud detection result; When performing target detection, the Softplus activation function is used to capture negative information and details of small targets.
5. The real-time target detection method integrating big data analysis according to claim 4 is characterized in that: The step of projecting the point cloud detection result onto the two-dimensional image data using the external parameter matrix obtained by calibrating the laser radar and the high-resolution camera to obtain a projected image includes: Establishing coordinate systems O1 and O2 of the laser radar and the high-resolution camera respectively, and realizing time synchronization of the three-dimensional point cloud data and the two-dimensional image data based on a timestamp; The coordinate system O1 and the coordinate system O2 are transformed into the same coordinate system by rotation and translation to obtain the external parameter matrix; The point cloud detection result is projected into the two-dimensional image data using the external parameter matrix to obtain the projected image.
6. The real-time target detection method integrating big data analysis according to claim 5 is characterized in that: The target matching is completed by using an IOU intersection-union method based on the projection image combined with the image recognition result to obtain a final target detection result, including: Compare the relationship between the projection image and the image recognition result using an IOU-based intersection-union method to generate a comparison result; The final target detection result is generated according to the comparison result and the actual situation, and is output in real time according to a preset output format.
7. The real-time target detection method integrating big data analysis according to claim 6 is characterized in that: The method further comprises: Tracking the target in real time using a multi-target tracking algorithm, and dynamically updating the target according to the three-dimensional point cloud data and the two-dimensional image data updated in real time to generate a travel trajectory; Analyze multiple travel trajectories through big data analysis to discover potential movement patterns or abnormal behaviors; The big data platform is used to store and analyze all the above data, and predictions and decisions are made based on the acquired environmental information.
8. The real-time target detection method integrating big data analysis according to claim 7 is characterized in that: The method may also use a GPU or FPGA hardware platform to accelerate the data processing process, or use multi-threaded parallel computing to analyze the three-dimensional point cloud data and the two-dimensional image data simultaneously.