Three-dimensional target detection method and system of object, electronic equipment and storage medium
By combining object velocity information and timing characteristics in the three-dimensional object detection technology, the object is inferred in the area of the historical frame and generating global features, the problem of poor combination of timing information and object motion information in the prior art is solved, the detection accuracy is improved and the calculation cost is reduced, and it is suitable for the actual deployment of intelligent driving tasks.
Patent Information
- Application Number
- CN202510104030.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-09
AI Technical Summary
When processing the time point cloud sequence of lidar, it is difficult to effectively combine timing information and object motion information, resulting in insufficient detection accuracy and stability, and high calculation cost, making it difficult to actually deploy in intelligent driving tasks.
By determining the object area and velocity information of each object in the current frame, inferring it in the object area of each consecutive historical frame, and generating global features based on this information, thereby realizing the calculation of the object's classification confidence and regression prediction box.
This method effectively avoids the problem of dynamic object drag caused by the direct superposition of multi-frame point clouds, improves the accuracy of three-dimensional detection, and greatly reduces the calculation cost, reduces the cost of three-dimensional target detection, and improves the actual deployment potential of intelligent driving tasks.
Smart Images

Figure CN119964114A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent driving technology, and more specifically, to a three-dimensional target detection method, system, electronic device and storage medium for an object. Background Art
[0002] LiDAR is an important sensor for intelligent driving technology, especially the three-dimensional target detection technology for LiDAR data, which is one of the core technologies for intelligent driving environment perception.
[0003] The current 3D target detection technology mainly uses single-frame laser point cloud technology as input, and processes it through voxelization, voxel feature extraction, multi-scale feature encoding, detection head and other modules to obtain the bounding boxes of several targets. However, in actual application scenarios, laser radar can also provide time point cloud sequences. Therefore, in order to improve the accuracy and stability of 3D target detection, time sequence information can also be added to the original 3D target detection technology.
[0004] In the existing technology, the temporal point cloud can be directly superimposed in space through the pose information to form a denser point cloud as the input of the detection model; or, the feature of each frame of point cloud can be extracted independently, and then the temporal features can be fused at a deeper feature level; or, the local point cloud of the object can be detected for each frame of point cloud, and then the temporal features of the object level can be fused through the attention mechanism or Transformer structure. However, the direct superposition of point clouds is prone to the problem of dynamic object smearing, and the object motion information is not fully considered through the method of temporal feature fusion, and the method of frame-by-frame point cloud feature extraction or frame-by-frame object detection will greatly increase the computational cost, which is not conducive to the actual deployment of intelligent driving tasks.
[0005] Therefore, how to provide a three-dimensional target detection method based on the fusion of object speed information and timing features to combine timing information and object motion information, improve the accuracy and stability of three-dimensional target detection, avoid the occurrence of dynamic object ghosting and reduce the computational cost, and facilitate the actual deployment of intelligent driving tasks is a problem that urgently needs to be solved in this application. Summary of the invention
[0006] In view of this, the present invention provides a three-dimensional target detection method, system, electronic device and storage medium for an object, so as to combine timing information and object motion information to improve the accuracy and stability of three-dimensional target detection, avoid the occurrence of dynamic object ghosting and reduce the calculation cost, so as to facilitate the actual deployment of intelligent driving tasks.
[0007] A first aspect of the present application provides a three-dimensional target detection method for an object, the method comprising:
[0008] Determine the object area and velocity information of each object in the current frame;
[0009] Determine the object region of the object in each consecutive historical frame based on the object region and speed information of the object in the current frame;
[0010] Determine a local point cloud region of the object in the current frame and a local point cloud region of the object in each historical frame based on the object region of the object in the current frame and the object region of each historical frame;
[0011] A three-dimensional object detection model is used to generate corresponding global features according to the object area of the object and the local point cloud area of each continuous frame, and the global features are processed to obtain the classification confidence and regression prediction box of the object; wherein each continuous frame includes a current frame and each historical frame.
[0012] Optionally, determining the object area and speed information of each object in the current frame includes:
[0013] Get the current frame point cloud of each object;
[0014] The current frame point cloud of the object is used by a three-dimensional target detection task to determine the object area and speed information of the object in the current frame; wherein the three-dimensional target detection model is a task that uses the current frame point cloud of the object as input and the target detection result as output, and the target detection result includes the object area and speed information of the object in the current frame.
[0015] Optionally, determining the object region of the object in each consecutive historical frame based on the object region and speed information of the object in the current frame includes:
[0016] Obtaining the first T-1 consecutive historical frames connected to the current frame from the cache; wherein T is greater than 1;
[0017] Controlling the object to move in a straight line at a uniform speed based on the object area of the object in the current frame and the speed information;
[0018] When the object is moving in a straight line at a uniform speed, the object region of the object in the historical frame is estimated according to the time difference between the current frame and the historical frame.
[0019] Optionally, determining the local point cloud area of the object in the current frame and the local point cloud area of each consecutive historical frame based on the object area of the object in the current frame and the object area in each historical frame includes:
[0020] Acquire the width and length of the object, and determine the area radius of the object in the current frame and the area radius of each of the historical frames according to the length and width of the object;
[0021] Determine a local point cloud area of the object in the current frame according to the object area and area radius of the object in the current frame;
[0022] According to the object area and area radius of each of the objects in the historical frame, a local point cloud area of each of the objects in the historical frame is determined.
[0023] Optionally, the three-dimensional object detection model includes a feature extraction module and a multi-layer perceptron module;
[0024] Generate corresponding global features according to the object area of the object and the local point cloud area of each continuous frame through the 3D object detection model, and process the global features to obtain the classification confidence and regression prediction box of each object, including:
[0025] Extracting local point cloud features of each coordinate point in the object area of the object from the local point cloud areas of each continuous frame of the object by the feature extraction module, fusing the local point cloud features of each coordinate point with its spatial information to obtain the spatial features of each coordinate point, and fusing the spatial features of each coordinate point with its time information to obtain the global features;
[0026] The global features are processed by the multi-layer perceptron module to obtain the classification confidence and regression prediction box of each object.
[0027] Optionally, the feature extraction module includes a local feature mixer, a spatial feature mixer and a temporal feature mixer;
[0028] The feature extraction module extracts local point cloud features of each coordinate point in the object area of the object from the local point cloud area of each continuous frame of the object, performs spatial information fusion on the local point cloud features of each coordinate point to obtain the spatial features of each coordinate point, and fuses the spatial features of each coordinate point with its time information to obtain the global features, including:
[0029] Extracting local point cloud features of each coordinate point in the object area of the object from the local point cloud areas of each continuous frame of the object by the local feature mixer;
[0030] The local point cloud features of each of the coordinate points are fused with the spatial information thereof by the spatial feature mixer to obtain the spatial features of each of the coordinate points;
[0031] The spatial features of each of the coordinate points are fused with their time information by the time feature mixer to obtain a global feature;
[0032] The local point cloud features, the spatial features and the temporal features are integrated to obtain global features.
[0033] Optionally, the multi-layer perceptron module includes a classification multi-layer perceptron and a regression multi-layer perceptron;
[0034] The global features are processed by a multi-layer perceptron module to obtain the classification confidence and regression prediction box of each object, including:
[0035] Processing the global features by the classification multilayer perceptron to obtain the classification confidence of each of the objects;
[0036] The global features are processed by the reply multi-layer perceptron to obtain a regression prediction box for each of the objects.
[0037] A second aspect of the present application provides a three-dimensional target detection system for an object, the system comprising:
[0038] A first determining unit, used to determine the object area and speed information of the object in the current frame;
[0039] A second determining unit, configured to determine an object region of the object in each historical frame continuous with the current frame based on the object region and speed information of the object in the current frame;
[0040] A third determining unit, configured to determine a local point cloud area of the object in the current frame and a local point cloud area of each historical frame based on the object area of the object in the current frame and the object area of each historical frame;
[0041] A processing unit is used to generate corresponding global features according to the object area of the object and the local point cloud area of each continuous frame through a three-dimensional object detection model, and process the global features to obtain the classification confidence and regression prediction box of the object; wherein each continuous frame includes a current frame and each historical frame.
[0042] The third aspect of the present invention provides an electronic device, comprising: a processor and a memory, wherein the processor and the memory are connected via a communication bus; wherein the processor is used to call and execute a program stored in the memory; and the memory is used to store a program, wherein the program is used to implement the three-dimensional target detection method of an object provided in the first aspect of the present invention.
[0043] A fourth aspect of the present invention provides a computer-readable storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are used to execute the three-dimensional target detection method of the object provided in the first aspect of the present invention.
[0044] The embodiment of the present invention provides a three-dimensional target detection method, system, electronic device and storage medium for an object, which determines the object area and speed information of each object in the current frame so as to infer the object area of the object in each continuous historical frame based on the object area and speed information of the object in the current frame, thereby avoiding the problem of dynamic object smear caused by the direct superposition of multi-frame point clouds, thereby improving the three-dimensional detection accuracy of the object; based on the object area of the object in the current frame and the object area in each historical frame, the local point cloud area of the object in the current frame and the local point cloud area in each historical frame are determined, and corresponding global features are generated according to the object area of the object and the local point cloud areas of each continuous frame through a three-dimensional target detection model, and the obtained global features are processed to obtain the classification confidence and regression prediction box of each object; wherein each continuous frame includes the current frame and each historical frame, and each continuous frame includes the current frame and each historical frame. It can be seen that the technical solution provided by the present invention, by combining the speed information of the object and the timing features contained in multiple consecutive frames, can not only avoid the problem of dynamic object ghosting caused by the direct superposition of multi-frame point clouds and improve the three-dimensional detection accuracy of the object, but also can greatly reduce the computational cost required for the detection method of multi-frame time-series point clouds, reduce the cost of three-dimensional target detection, and improve the potential for application to implementation scenarios, that is, improve the actual deployment in intelligent driving tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0046] Figure 1 A schematic diagram of a flow chart of a three-dimensional target detection method for an object provided by an embodiment of the present invention;
[0047] Figure 2 An example diagram of an object region and a local point cloud region of an object is provided for an embodiment of the present invention;
[0048] Figure 3 A structural schematic diagram of a feature extraction module provided by an embodiment of the present invention;
[0049] Figure 4 An example diagram of a three-dimensional target detection method for an object provided by an embodiment of the present invention;
[0050] Figure 5 A schematic diagram of the structure of a three-dimensional target detection system for an object provided by an embodiment of the present invention;
[0051] Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0052] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0053] In this application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0054] From the above background technology, it can be seen that in the prior art, the temporal point clouds can be directly superimposed in space through posture information to form a denser point cloud as the input of the detection model; or, the features of each frame of point cloud can be extracted independently, and then the temporal features are fused at a deeper feature level; or, the local point cloud of the object is first detected for each frame of point cloud, and then the object-level temporal features are fused through the attention mechanism or Transformer structure.
[0055] However, the method of directly superimposing time-series point clouds in space is essentially the same as single-frame target detection. Although the density of the point cloud itself is increased, the motion information of the object is lost, and the smears produced by dynamic objects during the superposition process may affect the detection accuracy of these objects. For the method that does not directly superimpose point clouds, although it can solve the smear problem of dynamic objects to a certain extent, each frame of point cloud is independently subjected to feature extraction or local object point cloud detection, and the use of complex computing units such as Transformer will greatly increase the computational cost of the deep model, which is not conducive to its actual deployment in intelligent driving tasks.
[0056] Therefore, an embodiment of the present invention provides a three-dimensional target detection method and system for an object, which determines the object area and speed information of each object in the current frame, so as to infer the object area of the object in each continuous historical frame based on the object area and speed information of the object in the current frame, thereby avoiding the problem of dynamic object dragging caused by the direct superposition of multi-frame point clouds, thereby improving the three-dimensional detection accuracy of the object; based on the object area of the object in the current frame and the object area in each historical frame, the local point cloud area of the object in the current frame is determined in the local point cloud area of each historical frame, and a corresponding global feature is generated according to the object area of the object and the local point cloud area of each continuous frame through a three-dimensional target detection model, and the obtained global feature is processed to obtain the classification confidence and regression prediction box of each object, which can greatly reduce the computational cost required for the detection method of multi-frame time series point clouds, reduce the cost of three-dimensional target detection, and improve the landing potential of application to implementation scenarios, that is, improve the actual deployment in intelligent driving tasks.
[0057] See also Figure 1 , shows a flow chart of a three-dimensional target detection method for an object provided by an embodiment of the present invention, and the three-dimensional target detection method for an object specifically comprises the following steps:
[0058] S101: Determine the object area and speed information of each object in the current frame.
[0059] In the embodiment of the present invention, the inventors have found that by making full use of the motion information of the object, the problem of dynamic object drag caused by the direct superposition of multiple frame point clouds can be effectively avoided. Therefore, a three-dimensional target detection task can be set. For each object, the current frame point cloud of the object can be obtained, and the object area and speed information of the object can be determined using the current frame point cloud of the object. Among them, the current frame point cloud of the object can be the point cloud of the object in the current frame. Among them, the three-dimensional target detection task can be CenterPoint, which is not limited in this embodiment of the present invention.
[0060] Optionally, the process of determining the object area and speed information of each object in the current frame can be specifically as follows: pre-setting a three-dimensional target detection task, and for each object, obtaining the current frame point cloud of the object; and determining the object area and speed information of the object in the current frame by using the current frame point cloud through the three-dimensional target detection task.
[0061] It should be noted that the three-dimensional target detection task takes the current frame point cloud of the object as input and takes the target detection result as output. The target detection result is the detection box of the object and the information with speed. Therefore, after determining the current frame point cloud of the object, the current frame point cloud of the object can be input into the three-dimensional target detection task, so that the three-dimensional target detection task can use the current frame point cloud of the object to predict the object area and speed information of the object in the current frame.
[0062] It should also be noted that the object area predicted by the three-dimensional object detection task is only the approximate location of the object at present, not the exact location. In addition, the three-dimensional object detection task is a three-dimensional object detection task without time sequence, and the corresponding three-dimensional object detection task can be set according to the actual application, which is not limited in this embodiment of the present invention.
[0063] S102: Based on the object region and speed information of the object in the current frame, determine the object region of the object in each consecutive historical frame.
[0064] In the embodiment of the present invention, since the acquisition frequency of the laser radar sensor is fixed, the time point from each historical frame to the current frame can constitute an arithmetic progression. Therefore, in actual use, there is no need to perform a specific selection operation on the historical frame. It is only necessary to read the first T-1 historical frames connected to the current frame acquired by the laser radar sensor in the cache, and use the read T-1 historical frames as a continuous plurality of historical frames. Of course, the present invention can also be a continuous plurality of historical frames consisting of T-1 historical frames specially selected and selected at uneven time intervals, which is not limited in this embodiment of the present invention. Wherein, T is greater than 1.
[0065] Optionally, for each object, after determining the object area and speed information of each object in the current frame, the first T-1 consecutive historical frames connected to the current frame can be obtained from the cache using a preset selection method; wherein T is greater than 1; based on the object area and speed information of the object in the current frame, the object is controlled to move in a uniform straight line; during the object's uniform straight line movement, the physical object area in the historical frame is inferred based on the time difference between the current frame and the historical frame.
[0066] In actual application, it can be assumed that the object starts to run in a straight line at a uniform speed from the object area at the speed indicated by the speed information of the object in a short period of time (within a preset time period), and in the process of running in a straight line at a uniform speed, the approximate position of the object in the historical frame can be deduced based on the time difference between the current frame and the historical frame and the running speed of the object, and the inferred approximate position is determined as the object area of the object in the historical frame.
[0067] It should be noted that the preset time period may be preset to 30 ms, and the preset time period and the preset time difference may be set according to actual applications, which is not limited in the embodiment of the present invention.
[0068] S103: Based on the object region of the object in the current frame and the object region of the object in each historical frame, determine the local point cloud region of the object in the current frame and the local point cloud region of each consecutive historical frame.
[0069] In the specific process of executing step S103, for each object, after determining the object area of the object in each historical frame, the length and width of the object itself can be further obtained, and based on the length, width, object area in the current frame and object area in each historical frame of the object, the local point cloud area of the object in the current frame and the local point cloud area in each historical frame can be determined.
[0070] Optionally, for each object, the width and length of the object may be obtained for each historical frame, and the region radius of the object in the current frame and the region radius in each historical frame may be determined based on the length and width of the object; the local point cloud region of the object in the current frame may be determined based on the object region and region radius of the object in the current frame; the local point cloud region of the object in each historical frame may be determined based on the object region and region radius of the object in each historical frame. The calculation method for determining the region radius based on the length and width of the object is shown in formula (1).
[0071] (1)
[0072] Among them, r is the radius of the area, w is the width of the object, and l is the length of the object. is the region extension parameter, and , is the interval frame number from the historical frame to the current frame. If the area radius of the current frame is calculated, then 0.
[0073] It should be noted that The purpose of setting it to be greater than 1 is that, as the time difference between the historical frame and the current frame increases, the error in estimating the object area by running in a straight line at a uniform speed will increase. Therefore, it is necessary to slightly increase the local point cloud area of the object in the historical frame. Therefore, the present invention uses an area expansion parameter The radius of the local point cloud area is increased as the time difference increases, so as to help the subsequent 3D target detection model to better extract the local point cloud features of the object in the local point cloud area of the historical frame.
[0074] In some embodiments, for each historical frame, after determining the area radius of the object in the historical frame, a circular area can be enclosed with the center of the object area in the historical frame as the center of the circle and the area radius in the historical frame as the radius, as the local point cloud area of the object in the historical frame.
[0075] In the embodiment of the present invention, the area radius of the object in the current frame can be calculated according to the calculation method shown in formula (1), and a circular area is enclosed with the center of the object area in the current frame as the center of the circle and the area radius in the current frame as the radius as the local point cloud area of the object in the current frame.
[0076] Further, in an embodiment of the present invention, after obtaining the local point cloud area of the object in the current frame and each historical frame, the points in each local point cloud area can be sampled to obtain the coordinate points of the object, such as Figure 2 shown.
[0077] S104: Generate corresponding global features according to the object area of the object and the local point cloud area of each continuous frame through the three-dimensional object detection model, and process the global features to obtain the classification confidence and regression prediction box of the object; wherein each continuous frame includes the current frame and each historical frame.
[0078] In the specific process of executing step S104, a three-dimensional target detection model can be constructed in advance according to the feature extraction module and the multi-layer perceptron module; for each object, after determining the object region of the object and the local point cloud region of the object in each historical frame, multiple coordinate points in the local point cloud region of each continuous frame of the object region of the object are input into the three-dimensional target detection model, and the local point cloud features of each coordinate point in the object region of the object are extracted from the local point cloud region of each continuous frame of the object by the feature extraction module in the three-dimensional target detection model, and the local point cloud features of each coordinate point are fused with its spatial information to obtain the spatial features of each coordinate point, and the spatial features of each coordinate point are fused with its time information to obtain the global features; the global features are processed by the multi-layer perceptron module to obtain the classification confidence and regression prediction box of each object. Among them, each object region includes the object region of the current frame and the object region of each historical frame.
[0079] The applicant has found through research that in the field of image target detection, the multilayer perceptron mixer uses two sets of multilayer perceptron structures to extract features from the image information and channel information of the blocks respectively. In view of this, in order to avoid introducing too high a computational cost, the present invention follows this design idea and combines the three-dimensional situation of the object to transform the multilayer perceptron mixer in the field of image target detection into a three-dimensional multilayer perceptron mixer, and uses the three-dimensional multilayer perceptron mixer as a feature extraction module of the three-dimensional target detection model. Among them, the three-dimensional multilayer perceptron mixer has the advantages of simple structure, low time and space overhead, and the ability to extract features in multiple dimensions.
[0080] Therefore, the local point cloud features of each coordinate point in the object area of the object can be extracted from the local point cloud areas of each continuous frame of the object through a three-dimensional multi-layer perceptron mixer, and the local point cloud features of each coordinate point are fused with its spatial information to obtain the spatial features of each coordinate point that integrates the local power features and spatial information, and the spatial features of each coordinate point are fused with its time information to obtain the global features that integrate the local point cloud features, spatial features and time features.
[0081] It should be noted that, since the points in the point cloud data are disordered, the local point cloud features of the local point cloud area of an object can be extracted and integrated through the "local feature mixer". Subsequently, this structure can be similarly applied to feature extraction in the spatial and temporal dimensions. It only needs to change the dimensional direction of the input multi-layer perceptron to the spatial direction and the temporal direction to form a "spatial feature mixer" and a "temporal feature mixer", which are used to integrate the spatial features between different objects and the temporal features of each object in the time sequence. Assuming that there are M objects, the input information can be composed according to the N coordinate points in the object area of the M objects and the local point cloud areas of T consecutive frames, and the input information is input into the feature extraction module, so that the feature extraction module extracts the local point cloud features of each coordinate point in the object area of the object from the local point cloud areas of each consecutive frame of the object, fuses the local point cloud features of each coordinate point with its spatial information, obtains the spatial features of each coordinate point, and fuses the spatial features of each coordinate point with its temporal information to obtain the global features that fuse the local point cloud features, spatial features and temporal features. Each mixer adopts a residual structure design, which is more suitable for the training and optimization of deep multi-layer perceptron networks. After connecting the above three feature mixers in series, we get a feature extraction module of MLPMixer3D, such as Figure 3 As shown in the figure, after the input information passes through the feature extraction module, a global feature is obtained that integrates the local point cloud features, spatial features and temporal features of each object.
[0082] Optionally, the feature extraction module includes a local feature mixer, a spatial feature mixer and a temporal feature mixer; the local point cloud features of each coordinate point in the object area of the object are extracted from the local point cloud areas of each continuous frame of the object through the local feature mixer; the local point cloud features of each coordinate point are fused with its spatial information through the spatial feature mixer to obtain the spatial features of each coordinate point; the spatial features of each coordinate point are fused with its temporal information through the temporal feature mixer to obtain the global features; the local point cloud features, spatial features and temporal features are integrated to obtain the global features.
[0083] It should be noted that the N coordinate points in the object area of M objects and the local point cloud area of T consecutive frames can be used to form a three-dimensional M×N×T input information, and the three-dimensional M×N×T input information is input into the feature extraction module; for each object, the local point cloud features of each coordinate point in the object area of each object are extracted from the local point cloud area of T consecutive frames of the object by the local feature mixer in the feature extraction module, and the local point cloud features of each coordinate point are fused with its spatial information by the spatial feature mixer to obtain the spatial features of each coordinate point that fuses the local point cloud features and spatial features, and finally the spatial features of each coordinate point are fused with its temporal information by the temporal feature mixer to obtain the global features that fuse the local point cloud features, spatial features and temporal features. Among them, M represents the object area of M objects, N represents the N coordinate points in each object area, and T represents a current frame and T-1 historical frames.
[0084] In an embodiment of the present invention, a multilayer perceptron module can be constructed based on a classification multilayer perceptron and a regression multilayer perceptron, so that after obtaining the global features, the global features can be processed by the classification multilayer perceptron to obtain the classification confidence of each object; the global features can be processed by the regression multilayer perceptron to obtain the regression prediction box of each object.
[0085] It should be noted that the multilayer perceptron module constructed based on the classification multilayer perceptron and the regression multilayer perceptron is shown in formula (2).
[0086] (2)
[0087] in, The classification loss function of the classification multilayer perceptron is implemented using the cross entropy function; is the regression loss function of the multi-layer perceptron, which is implemented using the L1 loss function. is the weight balancing parameter, and L is the loss function of the multilayer perceptron module consisting of the classification loss function and the regression loss function.
[0088] An embodiment of the present invention provides a three-dimensional target detection method for an object, by determining the object area and speed information of each object in the current frame, so as to infer the object area of the object in each continuous historical frame based on the object area and speed information of the object in the current frame, thereby avoiding the problem of dynamic object smear caused by the direct superposition of multi-frame point clouds, thereby improving the three-dimensional detection accuracy of the object; based on the object area of the object in the current frame and the object area in each historical frame, the local point cloud area of the object in the current frame and the local point cloud area in each historical frame are determined, and corresponding global features are generated according to the object area of the object and the local point cloud areas of each continuous frame through a three-dimensional target detection model, and the obtained global features are processed to obtain the classification confidence and regression prediction box of each object; wherein each continuous frame includes the current frame and each historical frame, and each continuous frame includes the current frame and each historical frame; each local point cloud area includes the local point cloud area of the object in the current frame and the local point cloud area in each historical frame. It can be seen that the technical solution provided by the present invention, by combining the speed information of the object and the timing characteristics contained in multiple consecutive historical frames, can not only avoid the problem of dynamic object ghosting caused by the direct superposition of multi-frame point clouds and improve the three-dimensional detection accuracy of objects, but also can greatly reduce the computational cost required for the detection method of multi-frame time-series point clouds, reduce the cost of three-dimensional target detection, and improve the potential for application to implementation scenarios, that is, improve the actual deployment in intelligent driving tasks.
[0089] In order to better understand the three-dimensional target detection method of the object provided by the above embodiment of the present invention, the following is explained by way of example. Figure 4 .
[0090] For each of the M objects, the current frame point cloud of the object can be obtained, and the object area and speed information of the object in the current frame can be predicted by the three-dimensional target detection task using the current frame point cloud of the object, that is, the M object areas and speed information in the current frame can be obtained.
[0091] Read T-1 historical frames collected by the lidar sensor in the cache, and use the read T-1 historical frames as continuous T-1 historical frames; control the object to move in a straight line at a uniform speed within a preset time period based on the object area and speed information of the object in the current frame, and in the process of the object moving in a straight line at a uniform speed, infer the object area in the historical frame based on the time difference between the current frame and the historical frame.
[0092] Get the width and length of the object, and determine the area radius of the object in each historical frame based on the length and width of the object; determine the local point cloud area of the object in the current frame based on the object radius and area radius of the object in the current frame, and determine the local point cloud area of the object in each historical frame based on the object area and area radius of the object in each historical frame; sample the points in each local point cloud area to obtain N coordinate points in the local point cloud area.
[0093] According to the N coordinate points in the object area of M objects and the local point cloud areas of T continuous frames, M*N*T input information is constructed, and the input information is input into the 3D target detection model; the local point cloud features of each coordinate point in the object area of the object are extracted from the local point cloud areas of each continuous frame of the object through the local feature mixer in the feature extraction module of the 3D target detection model; the local point cloud features of each coordinate point are fused with its spatial information through the spatial feature mixer to obtain the spatial features of each coordinate point; finally, the spatial features of each coordinate point are fused with its temporal information through the temporal feature mixer to obtain the global features.
[0094] Finally, the global features are processed by the classification multilayer perceptron in the multilayer perceptron module to obtain the classification confidence of each object; the global features are processed by the reply multilayer perceptron to obtain the regression prediction box of each object.
[0095] Based on the three-dimensional target detection method of the object provided by the above embodiment of the present invention, the embodiment of the present invention also provides a three-dimensional target detection system of the object, such as Figure 5 As shown, the three-dimensional target detection system of the object includes:
[0096] A first determining unit 51 is used to determine the object area and speed information of the object in the current frame;
[0097] A second determining unit 52, configured to determine the object region of the object in each of the consecutive historical frames based on the object region and speed information of the object in the current frame;
[0098] A third determining unit 53, configured to determine a local point cloud region of the object in the current frame and a local point cloud region of each historical frame based on the object region of the object in the current frame and the object region of each historical frame;
[0099] The processing unit 54 is used to generate corresponding global features according to the object area of the object and the local point cloud area of each continuous frame through the three-dimensional object detection model, and process the global features to obtain the classification confidence and regression prediction box of the object; wherein each continuous frame includes the current frame and each historical frame.
[0100] The specific principles and execution processes of each unit in the three-dimensional target detection system for objects disclosed in the above-mentioned embodiment of the present invention are the same as those of the three-dimensional target detection method for objects disclosed in the above-mentioned embodiment of the present invention. Please refer to the corresponding parts of the three-dimensional target detection method for objects disclosed in the above-mentioned embodiment of the present invention, and will not be repeated here.
[0101] The embodiment of the present invention provides a three-dimensional target detection system for an object, by determining the object area and speed information of each object in the current frame, so as to infer the object area of the object in each continuous historical frame based on the object area and speed information of the object in the current frame, thereby avoiding the problem of dynamic object drag caused by direct superposition of multi-frame point clouds, thereby improving the three-dimensional detection accuracy of the object; based on the object area of the object in each historical frame, the local point cloud area of the object in each historical frame is determined, and the corresponding global features are generated according to the object area of the object and the local point cloud area of each continuous frame through the three-dimensional target detection model, and the global features are processed to obtain the classification confidence and regression prediction box of the object; wherein each continuous frame includes the current frame and each historical frame. The technical solution provided by the present invention, by combining the speed information of the object and the time series features contained in multiple continuous historical frames, can not only avoid the problem of dynamic object drag caused by direct superposition of multi-frame point clouds, improve the three-dimensional detection accuracy of the object, but also can greatly reduce the calculation cost required for the detection method of multi-frame time series point clouds, reduce the cost of three-dimensional target detection, and improve the landing potential of application to the implementation scene, that is, improve the actual deployment on intelligent driving tasks.
[0102] Optionally, the first determining unit includes:
[0103] A static point cloud acquisition unit, used to acquire the current frame point cloud of each object;
[0104] The first determination subunit is used to determine the object area and speed information of the object in the current frame by using the current frame point cloud of the object through a three-dimensional target detection task; wherein the three-dimensional target detection model is a task that uses the current frame point cloud of the object as input and the target detection result as output, and the target detection result includes the object area and speed information of the object in the current frame.
[0105] Optionally, the second determining unit includes:
[0106] A historical frame acquisition unit, used to acquire the previous T-1 consecutive historical frames connected to the current frame from the cache; wherein T is greater than 1;
[0107] A control unit, for controlling the object to move in a straight line at a uniform speed within a preset time period based on the object area and speed information of the object in the current frame;
[0108] The inference unit is used to infer the object area of the object in the historical frame according to the time difference between the current frame and the historical frame when the object is moving in a straight line at a uniform speed.
[0109] Optionally, the third determining unit includes:
[0110] A fourth determination unit is used to obtain the width and length of the object for each historical frame, and determine the area radius of the object in the historical frame according to the length and width of the object;
[0111] a fifth determining unit, configured to determine a local point cloud region of the object in the current frame according to the object region and region radius of the object in the current frame;
[0112] The sixth determining unit is used to determine the local point cloud area of the object in the historical frame according to the object area and area radius of the object in the historical frame.
[0113] Optionally, the three-dimensional object detection model includes a feature extraction module and a multi-layer perceptron module; the processing unit includes:
[0114] A feature extraction and fusion unit is used to extract local point cloud features of each coordinate point in the object area of the object from the local point cloud area of each continuous frame of the object, fuse the local point cloud features of each coordinate point with its spatial information to obtain the spatial features of each coordinate point, and fuse the spatial features of each coordinate point with its time information to obtain the global features;
[0115] The global feature processing unit is used to process the global features through the multi-layer perceptron module to obtain the classification confidence and regression prediction box of each object.
[0116] Optionally, the feature extraction module includes a local feature mixer, a spatial feature mixer and a temporal feature mixer, and the feature extraction and fusion unit includes:
[0117] A first extraction unit is used to extract local point cloud features of each coordinate point in the object area of the object from the local point cloud area of each continuous frame of the object through a local feature mixer;
[0118] The second extraction unit is used to fuse the local point cloud features of each coordinate point with its spatial information through a spatial feature mixer to obtain the spatial features of each coordinate point;
[0119] The third extraction unit is used to fuse the spatial features of each coordinate point with its time information through a time feature mixer to obtain a global feature;
[0120] The feature integration unit is used to integrate local point cloud features, spatial features and temporal features to obtain global features.
[0121] Optionally, a global feature processing unit includes:
[0122] A first global feature processing unit, used to process the global features through a classification multi-layer perceptron to obtain a classification confidence of each object;
[0123] The second global feature processing unit is used to process the global features by replying to the multi-layer perceptron to obtain a regression prediction box for each object.
[0124] The present application embodiment provides an electronic device, such as Figure 6 As shown, the electronic device includes a processor 601 and a memory 602, the memory 602 is used to store program code and data for three-dimensional target detection of objects, and the processor 601 is used to call the program instructions in the memory to execute the steps shown in the three-dimensional target detection method of objects in the above embodiment.
[0125] An embodiment of the present application provides a storage medium, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the three-dimensional target detection method of the object shown in the above embodiment.
[0126] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, in which the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0127] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0128] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0129] The above are only preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A three-dimensional target detection method for an object, characterized in that: The method comprises: Determine the object area and velocity information of each object in the current frame; Determine the object region of the object in each consecutive historical frame based on the object region and speed information of the object in the current frame; Determine a local point cloud region of the object in the current frame and a local point cloud region of the object in each historical frame based on the object region of the object in the current frame and the object region of each historical frame; A three-dimensional object detection model is used to generate corresponding global features according to the object area of the object and the local point cloud area of each continuous frame, and the global features are processed to obtain the classification confidence and regression prediction box of the object; wherein each continuous frame includes a current frame and each historical frame.
2. The method according to claim 1, characterized in that The determining of the object area and speed information of each object in the current frame includes: Get the current frame point cloud of each object; The current frame point cloud of the object is used by a three-dimensional target detection task to determine the object area and speed information of the object in the current frame; wherein the three-dimensional target detection model is a task that uses the current frame point cloud of the object as input and the target detection result as output, and the target detection result includes the object area and speed information of the object in the current frame.
3. The method according to claim 1, characterized in that The determining the object region of the object in each consecutive historical frame based on the object region and speed information of the object in the current frame includes: Obtaining the first T-1 consecutive historical frames connected to the current frame from the cache; wherein T is greater than 1; Controlling the object to move in a straight line at a uniform speed based on the object area of the object in the current frame and the speed information; When the object is moving in a straight line at a uniform speed, the object region of the object in the historical frame is estimated according to the time difference between the current frame and the historical frame.
4. The method according to claim 1, characterized in that: The determining of the local point cloud area of the object in the current frame and the local point cloud area of each continuous historical frame based on the object area of the object in the current frame and the object area in each historical frame comprises: Acquire the width and length of the object, and determine the area radius of the object in the current frame and the area radius of each of the historical frames according to the length and width of the object; Determine a local point cloud area of the object in the current frame according to the object area and area radius of the object in the current frame; According to the object area and area radius of each of the objects in the historical frame, a local point cloud area of each of the objects in the historical frame is determined.
5. The method according to claim 1, characterized in that: The three-dimensional object detection model includes a feature extraction module and a multi-layer perceptron module; Generate corresponding global features according to the object area of the object and the local point cloud area of each continuous frame through the 3D object detection model, and process the global features to obtain the classification confidence and regression prediction box of each object, including: Extracting local point cloud features of each coordinate point in the object area of the object from the local point cloud areas of each continuous frame of the object by the feature extraction module, fusing the local point cloud features of each coordinate point with its spatial information to obtain the spatial features of each coordinate point, and fusing the spatial features of each coordinate point with its time information to obtain the global features; The global features are processed by the multi-layer perceptron module to obtain the classification confidence and regression prediction box of each object.
6. The method according to claim 5, characterized in that The feature extraction module includes a local feature mixer, a spatial feature mixer and a temporal feature mixer; The feature extraction module extracts local point cloud features of each coordinate point in the object area of the object from the local point cloud area of each continuous frame of the object, performs spatial information fusion on the local point cloud features of each coordinate point to obtain the spatial features of each coordinate point, and fuses the spatial features of each coordinate point with its time information to obtain the global features, including: Extracting local point cloud features of each coordinate point in the object area of the object from the local point cloud areas of each continuous frame of the object by the local feature mixer; The local point cloud features of each of the coordinate points are fused with the spatial information thereof by the spatial feature mixer to obtain the spatial features of each of the coordinate points; The spatial features of each of the coordinate points are fused with their time information by the time feature mixer to obtain a global feature; The local point cloud features, the spatial features and the temporal features are integrated to obtain global features.
7. The method according to claim 5, characterized in that The multi-layer perceptron module includes a classification multi-layer perceptron and a regression multi-layer perceptron; The global features are processed by a multi-layer perceptron module to obtain the classification confidence and regression prediction box of each object, including: Processing the global features by the classification multilayer perceptron to obtain the classification confidence of each of the objects; The global features are processed by the reply multi-layer perceptron to obtain a regression prediction box for each of the objects.
8. A three-dimensional target detection system for an object, characterized in that: The system comprises: A first determining unit, used to determine the object area and speed information of the object in the current frame; A second determining unit, configured to determine an object region of the object in each historical frame continuous with the current frame based on the object region and speed information of the object in the current frame; A third determining unit, configured to determine a local point cloud area of the object in the current frame and a local point cloud area of each historical frame based on the object area of the object in the current frame and the object area of each historical frame; A processing unit is used to generate corresponding global features according to the object area of the object and the local point cloud area of each continuous frame through a three-dimensional object detection model, and process the global features to obtain the classification confidence and regression prediction box of the object; wherein each continuous frame includes a current frame and each historical frame.
9. An electronic device, characterized in that: include: A processor and a memory, wherein the processor and the memory are connected via a communication bus; wherein the processor is used to call and execute a program stored in the memory; The memory is used to store a program, and the program is used to implement the three-dimensional target detection method of an object as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to execute the three-dimensional target detection method of an object as described in any one of claims 1-7.