Target detection method and device, terminal equipment and storage medium

By extracting and fusing point cloud data and image data, and using a time alignment model based on a deformable attention mechanism to process features, the time misalignment problem caused by different sensor acquisition rates is solved, and the reliability of target detection is improved.

CN120219700AInactive Publication Date: 2025-06-27VANJEE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311753568.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Due to the different acquisition rates of different sensors, there is a problem of time misalignment between the point cloud and the image, which in turn leads to feature misalignment, reducing the reliability of target detection.

Method used

By obtaining point cloud data and image data of the area to be tested, point cloud features and three-dimensional image features are extracted, and fusion is performed, and the fusion features are processed using a time-aligned model based on the deformable attention mechanism to generate the target fusion features after time alignment.

Benefits of technology

The problem of feature misalignment caused by time misalignment between point cloud features and image features is solved, and the reliability of target detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219700A_ABST
    Figure CN120219700A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of computer application, and provides a target detection method and device, terminal equipment and a storage medium, and the method comprises the steps: obtaining point cloud data and image data corresponding to a to-be-detected region; performing feature extraction on the point cloud data to generate point cloud features corresponding to the to-be-detected region; performing three-dimensional feature extraction on the image data to generate three-dimensional image features corresponding to the to-be-detected region; fusing the point cloud features and the three-dimensional image features to generate fused features corresponding to the to-be-detected region; processing the fusion feature by using a preset time alignment model to generate a target fusion feature after time alignment; and carrying out regression and classification on the target fusion features to generate a target detection result of the to-be-detected region. Therefore, the point cloud features corresponding to the to-be-detected area and the fusion features obtained after the image features are fused are subjected to alignment processing by using the preset time alignment model, the problem that feature fusion time is not aligned is solved, and the reliability of target detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer application technologies, and particularly relates to an object detection method, apparatus, terminal device, and computer-readable storage medium. Background Art

[0002] With the rapid development of fields such as autonomous driving and robot vision, 3D object detection has become a field of great concern. 3D object detection can detect information such as the position, category, and size of an object.

[0003] Traditional 3D object detection methods mainly use deep learning models to process point cloud data to generate object detection results. However, due to the sparsity and irregularity of point cloud data, it is easy to cause problems such as low detection accuracy and robustness.

[0004] In related technologies, in view of the defects existing in the above traditional object detection methods, an object detection method based on multi-sensor fusion has emerged. Usually, the point cloud and the image are respectively input into different networks, and then the features output by each network are fused, and then object detection is performed according to the fused features. However, due to reasons such as different acquisition rates of different sensors, there is a problem of time misalignment between the point cloud and the image, and feature misalignment will occur during feature fusion, resulting in poor reliability of object detection. Summary of the Invention

[0005] The purpose of this application is to provide an object detection method, which can solve the problem in related technologies that due to reasons such as different acquisition rates of different sensors, there is a phenomenon of time misalignment between the point cloud and the image, and feature misalignment will occur during feature fusion, resulting in poor reliability of object detection.

[0006] In a first aspect, an embodiment of this application provides an object detection method, including: obtaining point cloud data and image data corresponding to a region to be detected; extracting features from the point cloud data to generate point cloud features corresponding to the region to be detected; performing 3D feature extraction on the image data to generate 3D image features corresponding to the region to be detected; fusing the point cloud features and the 3D image features to generate fused features corresponding to the region to be detected; using a preset time alignment model to process the fused features to generate target fused features after time alignment; performing regression and classification on the target fused features to generate object detection results corresponding to the region to be detected.

[0007] In a possible implementation manner of the first aspect, the above preset time alignment model is a time alignment model based on a deformable attention mechanism. Before using the preset time alignment model to process the fused features to generate target fused features after time alignment, the above method further includes:

[0008] Obtain a training sample set, where the training sample set includes a plurality of training samples and the annotation data corresponding to each training sample;

[0009] Input each training sample into an initial temporal alignment model based on a deformable attention mechanism to generate a predicted feature vector corresponding to each training sample;

[0010] Determine the current loss value of the initial temporal alignment model based on the deformable attention mechanism according to the difference between the predicted feature vector corresponding to each training sample and the annotation data;

[0011] Adjust the parameters of the initial temporal alignment model based on the deformable attention mechanism according to the current loss value to generate an updated temporal alignment model based on the deformable attention mechanism;

[0012] Input each training sample into the updated temporal alignment model based on the deformable attention mechanism to continue training until the current loss value of the updated temporal alignment model based on the deformable attention mechanism is less than or equal to a loss threshold, then determine the updated temporal alignment model based on the deformable attention mechanism as a preset temporal alignment model.

[0013] Optionally, in another possible implementation manner of the first aspect, before fusing the point cloud feature and the three-dimensional image feature to generate a fused feature corresponding to the region to be measured, it further includes:

[0014] Perform dimensional transformation on the point cloud feature according to the dimension of the three-dimensional image feature so that the dimension of the point cloud feature matches the dimension of the three-dimensional image feature.

[0015] Optionally, in still another possible implementation manner of the first aspect, the extracting three-dimensional features from the image data to generate a three-dimensional image feature corresponding to the region to be measured includes:

[0016] Perform two-dimensional image extraction on the image data to generate a two-dimensional image feature corresponding to the region to be measured;

[0017] Perform three-dimensional mapping on the two-dimensional image feature to generate a three-dimensional image feature.

[0018] Optionally, in yet another possible implementation manner of the first aspect, the performing three-dimensional mapping on the two-dimensional image feature to generate a three-dimensional image feature includes:

[0019] Determine the frustum point cloud data corresponding to the image data according to a preset depth probability vector and the two-dimensional image feature, where the preset depth probability vector includes probability values corresponding to a plurality of preset depth values, and the frustum point cloud data includes the first spatial positions and feature vectors of a plurality of frustum points, and the first spatial position refers to the spatial position of the frustum point in the camera coordinate system;

[0020] Determine the second spatial position of each frustum point in the external coordinate system according to the first spatial position of each frustum point and the internal and external parameters of the camera;

[0021] Determine the grid to which each frustum point belongs according to the second spatial position of each frustum point and the preset bird's-eye view grid map, where the preset bird's-eye view grid map includes a plurality of grids;

[0022] Determine the feature vector corresponding to each grid according to the feature vectors of the respective frustum points included in each grid;

[0023] Generate three-dimensional image features according to the feature vector corresponding to each grid.

[0024] Optionally, in another possible implementation manner of the first aspect, the above-mentioned feature extraction of the point cloud data to generate the point cloud features corresponding to the area to be measured includes:

[0025] Perform voxelization processing on the point cloud data to determine a plurality of voxels corresponding to the point cloud data;

[0026] Perform feature encoding on the point cloud data in each voxel to generate point cloud features, where the point cloud features include the point cloud features corresponding to each voxel.

[0027] Optionally, in yet another possible implementation manner of the first aspect, before the above-mentioned feature encoding of the point cloud data in each voxel to generate point cloud features, it further includes:

[0028] Remove the voxels containing zero point cloud data.

[0029] Optionally, in another possible implementation manner of the first aspect, the above-mentioned feature encoding of the point cloud data in each voxel to generate point cloud features includes:

[0030] Respectively sample a preset number of target point cloud data from the respective point cloud data included in each voxel;

[0031] Perform feature encoding on the target point cloud data in each voxel to generate point cloud features.

[0032] Optionally, in another possible implementation manner of the first aspect, the above-mentioned respectively sampling a preset number of target point cloud data from the respective point cloud data included in each voxel includes:

[0033] When the number of point cloud data included in any voxel is less than the preset number, perform filling processing on the point cloud data included in any voxel to generate a preset number of target point cloud data.

[0034] In a second aspect, the present application also provides an object detection device, including: a first acquisition module for acquiring point cloud data and image data corresponding to a region to be measured; a first feature extraction module for extracting point cloud data features to generate point cloud features corresponding to the region to be measured; a second feature extraction module for extracting three-dimensional features of the image data to generate three-dimensional image features corresponding to the region to be measured; a first fusion module for fusing the point cloud features and the three-dimensional image features to generate fused features corresponding to the region to be measured; a first alignment module for processing the fused features by using a preset time alignment model to generate target fused features after time alignment; and a first classification module for performing regression and classification on the target fused features to generate an object detection result corresponding to the region to be measured.

[0035] In a possible implementation manner of the second aspect, the above-mentioned preset time alignment model is a time alignment model based on a deformable attention mechanism; correspondingly, the above-mentioned object detection device further includes:

[0036] a second acquisition module for acquiring a training sample set, where the training sample set includes a plurality of training samples and annotation data corresponding to each training sample;

[0037] a first input module for inputting each training sample into an initial time alignment model based on a deformable attention mechanism to generate a predicted feature vector corresponding to each training sample;

[0038] a first determination module for determining a current loss value of the initial time alignment model based on a deformable attention mechanism according to the difference between the predicted feature vector corresponding to each training sample and the annotation data;

[0039] an adjustment module for adjusting parameters of the initial time alignment model based on a deformable attention mechanism according to the current loss value to generate an updated time alignment model based on a deformable attention mechanism;

[0040] a second determination module for inputting each training sample into the updated time alignment model based on a deformable attention mechanism to continue training until the current loss value of the updated time alignment model based on a deformable attention mechanism is less than or equal to a loss threshold, and then determining the updated time alignment model based on a deformable attention mechanism as the preset time alignment model.

[0041] Optionally, in another possible implementation manner of the second aspect, the above-mentioned object detection device further includes:

[0042] a dimension transformation module for performing dimension transformation on the point cloud features according to the dimension of the three-dimensional image features so that the dimension of the point cloud features matches the dimension of the three-dimensional image features.

[0043] Optionally, in another possible implementation manner of the second aspect, the above-mentioned second feature extraction module includes:

[0044] A first extraction unit, configured to perform two-dimensional image extraction on the image data to generate two-dimensional image features corresponding to the area to be measured;

[0045] A first generation unit, configured to perform three-dimensional mapping on the two-dimensional image features to generate three-dimensional image features.

[0046] Optionally, in another possible implementation manner of the second aspect, the above-mentioned first generation unit is specifically configured to:

[0047] Determine the frustum point cloud data corresponding to the image data according to a preset depth probability vector and the two-dimensional image features, where the preset depth probability vector includes probability values corresponding to multiple preset depth values, and the frustum point cloud data includes the first spatial positions and feature vectors of multiple frustum points, and the first spatial position refers to the spatial position of the frustum point in the camera coordinate system;

[0048] Determine the second spatial position of each frustum point in the external coordinate system according to the first spatial position of each frustum point and the internal and external parameters of the camera;

[0049] Determine the grid to which each frustum point belongs according to the second spatial position of each frustum point and a preset bird's-eye view grid map, where the preset bird's-eye view grid map includes multiple grids;

[0050] Determine the feature vector corresponding to each grid according to the feature vectors of the respective frustum points included in each grid;

[0051] Generate three-dimensional image features according to the feature vector corresponding to each grid.

[0052] Optionally, in another possible implementation manner of the second aspect, the above-mentioned first feature extraction module includes:

[0053] A first determination unit, configured to perform voxelization processing on the point cloud data to determine multiple voxels corresponding to the point cloud data;

[0054] A second generation unit, configured to perform feature encoding on the point cloud data within each voxel to generate point cloud features, where the point cloud features include the point cloud features corresponding to each voxel.

[0055] Optionally, in another possible implementation manner of the second aspect, the above-mentioned first feature extraction module further includes:

[0056] A removal unit, configured to remove the voxels containing zero point cloud data.

[0057] Optionally, in another possible implementation manner of the second aspect, the above-mentioned second generation unit is specifically configured to:

[0058] Sample a preset number of target point cloud data from each of the point cloud data included in each voxel;

[0059] Perform feature encoding on the target point cloud data in each voxel to generate point cloud features.

[0060] Optionally, in another possible implementation manner of the second aspect, the above-mentioned second generation unit is further specifically configured to:

[0061] When the number of point cloud data included in any voxel is less than the preset number, perform padding processing on the point cloud data included in any voxel to generate a preset number of target point cloud data.

[0062] In a third aspect, the present application further provides a terminal device. The terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method of any one of the implementation manners of the first aspect is implemented.

[0063] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method of any one of the implementation manners of the first aspect is implemented.

[0064] In a fifth aspect, the present application further provides a computer program product. When the computer program product runs on an electronic device, the electronic device is enabled to execute the method of any one of the implementation manners of the first aspect.

[0065] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: By fusing the point cloud features and image features corresponding to the area to be measured to generate fusion features, and using the time alignment model to perform alignment processing on the fusion features, the problem of feature misalignment caused by time misalignment between the point cloud features and the image features is solved, and the reliability of target detection is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0067] Figure 1 It is a schematic flowchart of a target detection method provided by an embodiment of the present application;

[0068] Figure 2 It is a schematic diagram of feature extraction of point cloud data provided by an embodiment of the present application;

[0069] Figure 3 It is a schematic diagram of an aerial view grid map provided by another embodiment of the present application;

[0070] Figure 4 It is a schematic structural diagram of a target detection device provided by an embodiment of the present application;

[0071] Figure 5 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0072] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0073] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0074] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0075] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0076] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0077] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0078] The object detection method, device, terminal device, storage medium and computer program provided by this application will be described in detail below with reference to the accompanying drawings.

[0079] Figure 1 The flowchart of an object detection method provided by an embodiment of this application is shown.

[0080] Step 101, obtain the point cloud data and image data corresponding to the area to be measured.

[0081] It should be noted that the object detection method of the embodiment of this application can be executed by the object detection device of the embodiment of this application. The object detection device of the embodiment of this application can be configured in any terminal device to execute the object detection method of the embodiment of this application.

[0082] As a possible implementation, the area to be measured can be any area. The point cloud data of the area to be measured can be collected by a sensor. The point cloud data can refer to a set of vectors in a three-dimensional coordinate system and can include color information, reflection intensity information, etc. The image data of the area to be measured can be collected by a camera or a scanner, etc. The image data can be a set of gray values of each pixel represented by numerical values.

[0083] As an example, when applied to the field of autonomous driving, the area to be measured can be the area around the vehicle. The lidar can be used to emit laser beams to the surrounding environment of the vehicle and measure the time of reflection back, so as to obtain the three-dimensional coordinates of the surface of the objects in the surrounding environment of the vehicle, and thus obtain the point cloud data of the area around the vehicle. The camera can be used to take pictures of the area around the vehicle and process the taken images, so as to obtain the image data of the area around the vehicle.

[0084] It should be noted that the above-listed methods for obtaining point cloud data and image data and the content included in the point cloud data and image data are only exemplary. In actual use, it can be determined according to actual usage requirements and application scenarios, and the embodiments of this application do not make any limitations in this regard.

[0085] Step 102: Extract features from the point cloud data to generate point cloud features corresponding to the area to be measured.

[0086] As a possible implementation, features can be extracted from the point cloud data. For example, based on normal feature extraction, the local shape of the point cloud can be described by calculating the normal vector of each point in the point cloud data; based on shape feature extraction, the overall shape of the point cloud data can be described using shape features. Commonly used shape feature extraction methods include voxelization method, contour description method, curvature method, etc. After extracting features from the point cloud data, point cloud features corresponding to the area to be measured can be generated. The point cloud features can be a unique representation of the point cloud and can include geometric attributes (such as point cloud density, point cloud dispersion degree, and point cloud curvature, etc.) and statistical attributes (such as the mean, variance, covariance, and maximum and minimum values of the point cloud data, etc.).

[0087] It should be noted that the above-listed feature extraction methods and the content included in the point cloud features are only exemplary. In actual use, they can be determined according to actual usage requirements and application scenarios, and the embodiments of the present application do not make limitations in this regard.

[0088] Furthermore, the amount of point cloud data may be large. Therefore, in order to perform feature extraction more accurately and improve the efficiency of feature extraction, the point cloud data can be voxelized and the point cloud data within each voxel can be feature-encoded. That is, in a possible implementation of the embodiments of the present application, the above Step 102 may include:

[0089] Voxelize the point cloud data to determine multiple voxels corresponding to the point cloud data;

[0090] Feature-encode the point cloud data within each voxel to generate point cloud features, where the point cloud features include the point cloud features corresponding to each voxel.

[0091] As a possible implementation, the point cloud data can be input into the VoxelNet network model to voxelize the point cloud data. That is, the point cloud data can be spatially divided and divided into stacked voxel spaces of equal volume, and multiple voxels corresponding to the point cloud data are obtained. Then, the point cloud data within each voxel is feature-encoded to generate the point cloud features corresponding to the point cloud data within each voxel, thereby generating the point cloud features of the point cloud data corresponding to the area to be measured.

[0092] As a possible implementation, stacked voxel feature encoding can be performed on the point cloud data, and the point cloud data is input into a stacked voxel feature encoding layer for feature extraction, such as Figure 2As shown, the point cloud data 21 can be data-augmented to obtain the augmented point cloud data 22. The augmented point cloud data 22 is input into a fully-connected layer to be transformed into a high-dimensional space, obtaining point-by-point features 23. From the point-by-point features, the point with the largest eigenvalue is taken to obtain local aggregation features 24. The local aggregation features 24 and the point-by-point features 23 are concatenated to obtain point-by-point concatenated features 25. Then, the point-by-point concatenated features 25 are used as the input for the next voxel feature encoding layer. After cyclic feature encoding through multiple stacked voxel feature encoding layers, the point cloud features of the point cloud data corresponding to the area to be measured are generated.

[0093] Further, after voxelizing the point cloud data, there may be a situation where there is no point cloud data in some voxel spaces. Feature encoding of such voxels cannot extract point cloud features. Therefore, such voxels can be removed to further improve the efficiency of feature extraction. That is, in a possible implementation manner of the embodiment of the present application, before the above step 102, it may further include:

[0094] Removing the voxels containing 0 point cloud data.

[0095] As a possible implementation manner, all voxels corresponding to the point cloud data can be screened, and the voxels not containing point cloud data are removed without subsequent feature encoding processing.

[0096] Further, some voxels may contain a large amount of point cloud data. Directly calculating such voxels may waste computing resources and even cause detection deviation. To save computing resources and further improve the accuracy of point cloud feature extraction, the point cloud data in the voxels can be sampled, and then feature encoding is performed after sampling. That is, in a possible implementation manner of the embodiment of the present application, the above step 102 may further include:

[0097] Respectively sampling a preset number of target point cloud data from each of the point cloud data contained in each voxel;

[0098] Performing feature encoding on the target point cloud data in each voxel to generate point cloud features.

[0099] As a possible implementation manner, the preset number can be determined according to actual requirements, actual point cloud quantity, actual computing and processing capabilities, etc. From the point cloud data of each voxel, the preset number of point cloud data is sampled as the target point cloud data, and feature encoding can be performed on the target point cloud data to generate point cloud features.

[0100] Furthermore, if the number of point cloud data contained in the voxel is less than the preset number of samples, the point cloud data may be filled to further improve the accuracy of feature extraction. That is, in a possible implementation of the embodiment of the present application, the above-mentioned sampling of a preset number of target point cloud data from each point cloud data contained in each voxel may also include:

[0101] When the number of point cloud data contained in any voxel is less than a preset number, the point cloud data contained in any voxel is filled to generate a preset number of target point cloud data.

[0102] As a possible implementation method, when the number of point cloud data in a voxel is less than a preset number, the point cloud data in the voxel can be filled. For example, filling can be performed based on the statistical characteristics of the point cloud data contained in the voxel, and the statistical characteristics of each point, such as mean, variance, covariance, etc., can be extracted by statistically analyzing and processing the point cloud data, and then the attributes or state of the current filling point can be inferred based on these statistical characteristics, and the filling point can be filled into the voxel until the number of point cloud data in the voxel reaches the preset number, thereby generating the preset number of target point cloud data.

[0103] Step 103: extract three-dimensional features from the image data to generate three-dimensional image features corresponding to the area to be measured.

[0104] As a possible implementation method, three-dimensional features can be extracted from image data using a pre-trained image feature extraction model, so as to generate three-dimensional image features corresponding to the area to be tested based on the image data of the area to be tested.

[0105] Furthermore, the image data may be first subjected to two-dimensional image feature extraction, and then the two-dimensional image features may be mapped to three dimensions to improve the accuracy of image feature extraction. That is, in a possible implementation of the embodiment of the present application, the above step 103 may include:

[0106] Performing two-dimensional image extraction on the image data to generate two-dimensional image features corresponding to the area to be tested;

[0107] The two-dimensional image features are mapped three-dimensionally to generate three-dimensional image features.

[0108] As a possible implementation of the present application, any feasible encoding method can be used to perform two-dimensional image extraction on the image data, such as the scale-invariant feature transformation method and the directional gradient histogram method, so as to obtain the two-dimensional image features corresponding to the area to be measured. The two-dimensional image features can be mapped to a three-dimensional space to generate the three-dimensional image features corresponding to the area to be measured.

[0109] Further, a bird's-eye view bottom-up method can be adopted to implement the mapping from two-dimensional image features to three-dimensional image features, so as to further improve the efficiency and accuracy of image feature extraction. That is, in a possible implementation manner of the embodiments of the present application, the above-mentioned three-dimensional mapping of two-dimensional image features to generate three-dimensional image features may include:

[0110] According to a preset depth probability vector and two-dimensional image features, determine the frustum point cloud data corresponding to the image data, where the preset depth probability vector includes probability values corresponding to multiple preset depth values, and the frustum point cloud data includes the first spatial positions and feature vectors of multiple frustum points, and the first spatial position refers to the spatial position of the frustum point in the camera coordinate system;

[0111] According to the first spatial position of each frustum point and the internal and external parameters of the camera, determine the second spatial position of each frustum point in the external coordinate system;

[0112] According to the second spatial position of each frustum point and a preset bird's-eye view grid map, determine the grid to which each frustum point belongs, where the preset bird's-eye view grid map includes multiple grids;

[0113] According to the feature vectors of each frustum point included in each grid, determine the feature vector corresponding to each grid;

[0114] Generate three-dimensional image features according to the feature vector corresponding to each grid.

[0115] As a possible implementation manner of the present application, the preset depth probability vector can be obtained through a pre-trained depth probability model, including various possible depth values corresponding to each point in the two-dimensional image features, and the probability values corresponding to each depth value. Among them, the points corresponding to various possible depth values of each point in the two-dimensional image features can be called frustum points, and the spatial position of each frustum point relative to the camera that acquires the image data can be the first spatial position. Feature extraction and downsampling can be performed on the two-dimensional image features to obtain the semantic features of the two-dimensional image features, and the outer product of the semantic features and the preset depth probability vector can be used to obtain the feature vector corresponding to each frustum point, so as to generate frustum point cloud data. The frustum point cloud data can include the first spatial position of each frustum point and the feature vector corresponding to each frustum point.

[0116] As a possible implementation manner, the external coordinate system can be determined according to the actual application scenario, and according to the internal and external parameters of the camera that acquires the image data, the first spatial position of each frustum point can be converted into the coordinate position in the external coordinate system, and this coordinate position can be used as the second spatial position.

[0117] As an example, assume that in the field of autonomous driving, the camera on the vehicle is used to obtain image data. Then, a vehicle coordinate system can be established with the center of the rear axle of the vehicle as the origin. As shown in Figure 3 In the external coordinate system, the first spatial position of each frustum point is converted into the coordinate position in the vehicle coordinate system, thereby generating the second spatial position of each frustum point.

[0118] As a possible implementation, a preset bird's-eye view grid map can be generated according to the actual application scenario. The preset bird's-eye view grid map can correspond to the external coordinate system and include multiple grids. According to the second spatial position of each frustum point, each frustum point can be assigned to the corresponding grid map, thereby determining the grid to which each frustum point belongs. If there are frustum points outside the corresponding preset bird's-eye view grid map, these frustum points are removed.

[0119] For example, as shown in Figure 3 N grids can be divided on the vehicle's top-down plane. If N is 40,000, that is, the preset bird's-eye view grid map includes 200×200 grids, and the length and width of each grid are 0.5 meters, then the preset bird's-eye view grid map can represent a plane range of 100 meters×100 meters around the vehicle. If the center of the vehicle is the center of the preset bird's-eye view grid map, the preset bird's-eye view grid map can represent a range of 50 meters in front of, behind, to the left, and to the right of the vehicle. According to the second spatial position, the frustum points are assigned to the corresponding 40,000 grids, and the frustum points outside the range of 50 meters in front of, behind, to the left, and to the right of the vehicle are ignored.

[0120] As a possible implementation, the feature vectors of the frustum points included in each grid can be added to obtain the feature vector corresponding to each grid. According to the feature vector corresponding to each grid, a three-dimensional image feature is generated.

[0121] Step 104: Fuse the point cloud feature and the three-dimensional image feature to generate a fusion feature corresponding to the area to be measured.

[0122] As a possible implementation, the point cloud feature and the three-dimensional image feature can be spliced to obtain a fusion feature corresponding to the area to be measured.

[0123] Furthermore, in order to improve the accuracy of feature fusion, before feature fusion, the dimensions of the point cloud feature and the image feature can be made to match. That is, in a possible implementation of the embodiment of the present application, before the above step 104, the following may further be included:

[0124] According to the dimension of the three-dimensional image feature, perform a dimension transformation on the point cloud feature so that the dimension of the point cloud feature matches the dimension of the three-dimensional image feature.

[0125] As a possible implementation, the dimensions of the point cloud features can be transformed according to the dimensions of the three-dimensional image features, that is, the number of voxels can be made equal to the number of grids in the preset bird's-eye view grid map, so that the dimensions of the point cloud features are equal to the dimensions of the three-dimensional image features, and thus the dimensions of the point cloud features match the dimensions of the three-dimensional image features.

[0126] Step 105: Process the fused features using a preset time alignment model to generate the target fused features after time alignment.

[0127] As a possible implementation, a preset time alignment model can be used to perform alignment processing on the fused features. The preset time alignment model can be pre-trained and can include a preset alignment vector. The point cloud features or image features or both in the fused features can be multiplied by the preset alignment vector in the preset time alignment model to correspond the respective features in the point cloud features and image features, thereby generating the target fused features after time alignment.

[0128] Furthermore, to make the processing of the fused features by the time alignment model more accurate, improve the reliability of feature fusion, and further improve the reliability of object detection, a time alignment model can be generated based on the deformable attention mechanism and can be pre-trained. That is, in a possible implementation of the embodiments of the present application, the above-mentioned preset time alignment model is a time alignment model based on the deformable attention mechanism. Before the above step 105, the following can also be included:

[0129] Obtain a training sample set, where the training sample set includes multiple training samples and the annotation data corresponding to each training sample;

[0130] Input each training sample into the initial time alignment model based on the deformable attention mechanism to generate a predicted feature vector corresponding to each training sample;

[0131] Determine the current loss value of the initial time alignment model based on the deformable attention mechanism according to the difference between the predicted feature vector corresponding to each training sample and the annotation data;

[0132] Adjust the parameters of the initial time alignment model based on the deformable attention mechanism according to the current loss value to generate an updated time alignment model based on the deformable attention mechanism;

[0133] Input each training sample into the updated time alignment model based on the deformable attention mechanism to continue training until the current loss value of the updated time alignment model based on the deformable attention mechanism is less than or equal to the loss threshold, and then determine the updated time alignment model based on the deformable attention mechanism as the preset time alignment model.

[0134] As a possible implementation method, a training sample set can be obtained, including each training sample and the annotation data corresponding to each training sample. For example, the training sample can be a fusion feature after the time-unaligned point cloud feature and the image feature fusion, etc., and the annotation data corresponding to the training sample can be a standard fusion feature after time alignment, etc. The time alignment model can be generated based on a deformable attention mechanism, which can be used to solve the problem of time misalignment between point cloud features and image features.

[0135] As a possible implementation method of the present application, each training sample can be input into an initial time alignment model based on a deformable attention mechanism. The initial time alignment model can include an initial alignment vector. The training sample can be processed with the initial alignment vector to generate a predicted feature vector corresponding to each training sample. The predicted feature vector corresponding to each training sample can be compared with the labeled data. According to the difference between the predicted feature vector and the labeled data, the current loss value of the initial time alignment model can be determined. The method for calculating the loss value can be selected according to actual needs. For example, the current loss value can be calculated according to the sum of the absolute values ​​of the differences between the predicted feature vector and the labeled data, the current loss value can be calculated according to the mean square error between the predicted feature vector and the labeled data, and the like.

[0136] As a possible implementation method, a loss threshold can be preset according to actual needs. According to the relationship between the current loss value of the initial time alignment model and the loss threshold, when the current loss value is greater than the loss threshold, the parameters of the initial time alignment model can be adjusted. For example, the initial alignment vector in the initial time alignment model can be adjusted to generate an updated time alignment model based on a deformable attention mechanism. The training samples are input into the updated time alignment model based on a deformable attention mechanism for processing, and the predicted feature vector output by the updated time alignment model is compared with the labeled data to obtain the current loss value corresponding to the updated time alignment model. The parameters of the updated time alignment model are adjusted according to the current loss value until the current loss value is less than or equal to the loss threshold. Then, the updated time alignment model based on a deformable attention mechanism corresponding to the current loss value can be determined as the preset time alignment model.

[0137] It should be noted that the above-listed training samples, labeled data and calculation methods of the current loss value are only exemplary. In actual use, they can be determined according to actual usage requirements and application scenarios, and the embodiments of the present application do not limit this.

[0138] Step 106, regressing and classifying the target fusion features to generate a target detection result corresponding to the area to be detected.

[0139] As a possible implementation, the target fusion feature can be input into a pre-trained regression and classification model to generate the target detection result corresponding to the area to be detected. For example, the category of each object in the area to be detected can be obtained through classification, and the position information of each object in the area to be detected can be obtained through regression, etc.

[0140] In the target detection method provided by the embodiments of the present application, by sampling the voxelized point cloud data and removing empty voxels, and then performing two-dimensional feature extraction on the image data and three-dimensional feature mapping, the feature extraction process is made more efficient and reliable, thereby improving the efficiency and reliability of target detection. And by using the time alignment model based on the deformable attention mechanism to process the fusion feature, the problem of time misalignment between the point cloud data and the image data is solved, thereby further improving the accuracy and reliability of target detection.

[0141] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0142] Corresponding to the target detection method described in the above embodiments, Figure 4 The structural block diagram of the target detection device provided by the embodiments of the present application is shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0143] Referring to Figure 4 , the device 40 includes:

[0144] A first acquisition module 41, configured to acquire point cloud data and image data corresponding to the area to be detected;

[0145] A first feature extraction module 42, configured to extract point cloud data features to generate point cloud features corresponding to the area to be detected;

[0146] A second feature extraction module 43, configured to extract three-dimensional features of the image data to generate three-dimensional image features corresponding to the area to be detected;

[0147] A first fusion module 44, configured to fuse the point cloud features and the three-dimensional image features to generate fusion features corresponding to the area to be detected;

[0148] A first alignment module 45, configured to process the fusion feature by using a preset time alignment model to generate a target fusion feature after time alignment;

[0149] A first classification module 46, configured to perform regression and classification on the target fusion feature to generate a target detection result corresponding to the area to be detected.

[0150] In actual use, the target detection device provided by the embodiments of the present application can be configured in any terminal device to execute the foregoing target detection method.

[0151] The target detection device provided by the embodiments of the present application generates a fusion feature by fusing the point cloud feature and the image feature corresponding to the area to be detected, and uses a time alignment model to perform alignment processing on the fusion feature, solving the problem of feature misalignment caused by time misalignment between the point cloud feature and the image feature, and improving the reliability of target detection.

[0152] In a possible implementation form of the present application, the above-mentioned preset time alignment model is a time alignment model based on a deformable attention mechanism; correspondingly, the above-mentioned target detection device 40 further includes:

[0153] A second acquisition module, configured to acquire a training sample set, where the training sample set includes a plurality of training samples and annotation data corresponding to each training sample;

[0154] A first input module, configured to input each training sample into an initial time alignment model based on a deformable attention mechanism to generate a predicted feature vector corresponding to each training sample;

[0155] A first determination module, configured to determine the current loss value of the initial time alignment model based on a deformable attention mechanism according to the difference between the predicted feature vector corresponding to each trained sample and the annotation data;

[0156] An adjustment module, configured to adjust the parameters of the initial time alignment model based on a deformable attention mechanism according to the current loss value to generate an updated time alignment model based on a deformable attention mechanism;

[0157] A second determination module, configured to input each training sample into the updated time alignment model based on a deformable attention mechanism to continue training until the current loss value of the updated time alignment model based on a deformable attention mechanism is less than or equal to a loss threshold, and then determine the updated time alignment model based on a deformable attention mechanism as the preset time alignment model.

[0158] Further, in another possible implementation form of the present application, the above-mentioned target detection device 40 further includes:

[0159] A dimension transformation module, configured to perform dimension transformation on the point cloud feature according to the dimension of the three-dimensional image feature so that the dimension of the point cloud feature matches the dimension of the three-dimensional image feature.

[0160] Further, in another possible implementation form of the present application, the above-mentioned second feature extraction module 43 includes:

[0161] The first extraction unit is configured to perform two-dimensional image extraction on the image data to generate two-dimensional image features corresponding to the area to be measured;

[0162] The first generation unit is configured to perform three-dimensional mapping on the two-dimensional image features to generate three-dimensional image features.

[0163] Furthermore, in another possible implementation form of the present application, the above-mentioned first generation unit is specifically configured to:

[0164] Determine the frustum point cloud data corresponding to the image data according to the preset depth probability vector and the two-dimensional image features, wherein the preset depth probability vector includes probability values corresponding to multiple preset depth values, and the frustum point cloud data includes the first spatial positions and feature vectors of multiple frustum points, and the first spatial position refers to the spatial position of the frustum point in the camera coordinate system;

[0165] Determine the second spatial position of each frustum point in the external coordinate system according to the first spatial position of each frustum point and the internal and external parameters of the camera;

[0166] Determine the grid to which each frustum point belongs according to the second spatial position of each frustum point and the preset bird's-eye view grid map, wherein the preset bird's-eye view grid map includes multiple grids;

[0167] Determine the feature vector corresponding to each grid according to the feature vectors of the respective frustum points included in each grid;

[0168] Generate three-dimensional image features according to the feature vector corresponding to each grid.

[0169] Furthermore, in another possible implementation form of the present application, the above-mentioned first feature extraction module 42 includes:

[0170] The first determination unit is configured to perform voxelization processing on the point cloud data to determine multiple voxels corresponding to the point cloud data;

[0171] The second generation unit is configured to perform feature encoding on the point cloud data within each voxel to generate point cloud features, wherein the point cloud features include the point cloud features corresponding to each voxel.

[0172] Furthermore, in another possible implementation form of the present application, the above-mentioned first feature extraction module 42 further includes:

[0173] The removal unit is configured to remove the voxels containing zero point cloud data.

[0174] Furthermore, in another possible implementation form of the present application, the above-mentioned second generation unit is specifically configured to:

[0175] Sample a preset number of target point cloud data from each of the point cloud data contained in each voxel;

[0176] Perform feature encoding on the target point cloud data within each voxel to generate point cloud features.

[0177] Further, in another possible implementation form of the present application, the above-mentioned second generation unit is specifically further configured to:

[0178] When the number of point cloud data contained in any voxel is less than the preset number, perform filling processing on the point cloud data contained in any voxel to generate the preset number of target point cloud data.

[0179] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of the present application, their specific functions and the technical effects brought about can be specifically referred to the method embodiment part, and will not be elaborated here.

[0180] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments and will not be elaborated here.

[0181] To implement the above embodiments, the present application also proposes a terminal device.

[0182] Figure 5 It is a schematic structural diagram of a terminal device according to an embodiment of the present application.

[0183] As Figure 5 shown, the above-mentioned terminal device 200 includes:

[0184] A memory 210 and at least one processor 220, a bus 230 connecting different components (including the memory 210 and the processor 220), and the memory 210 stores a computer program, and when the processor 220 executes the program, the target detection method described in the embodiments of the present application is implemented.

[0185] Bus 230 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor bus, or a local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include the Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0186] Terminal device 200 typically includes a variety of computer-readable media. These media can be any available media that can be accessed by terminal device 200, including both volatile and nonvolatile media, removable and non-removable media.

[0187] Memory 210 may also include computer-system-readable media in the form of volatile memory, such as random access memory (RAM) 240 and / or cache memory 250. Terminal device 200 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 260 can be used for reading from and writing to non-removable, nonvolatile magnetic media ( Figure 5 not shown and typically called a "hard disk drive"). Although Figure 5 not shown in FIG. , a disk drive for reading from and writing to a removable, nonvolatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from and writing to a removable, nonvolatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) can be provided. In such cases, each drive can be connected to bus 230 by one or more data media interfaces. Memory 210 may include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of the embodiments of the present application.

[0188] A program / utility 280 having a set (at least one) of program modules 270 can be stored, for example, in memory 210, such program modules 270 including—but not limited to—an operating system, one or more application programs, other program modules, and program data, each of which examples or some combination thereof may include an implementation of a network environment. Program modules 270 generally carry out the functions and / or methods of the embodiments described herein.

[0189] The terminal device 200 can also communicate with one or more external devices 290 (such as a keyboard, a pointing device, a display 291, etc.), and can also communicate with one or more devices that enable a user to interact with the terminal device 200, and / or communicate with any device that enables the terminal device 200 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 292. Moreover, the terminal device 200 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 293. As shown in the figure, the network adapter 293 communicates with other modules of the terminal device 200 through a bus 230. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the terminal device 200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0190] The processor 220 executes various functional applications and data processing by running programs stored in the memory 210.

[0191] It should be noted that for the implementation process and technical principle of the terminal device in this embodiment, refer to the foregoing explanation of the object detection method in the embodiments of the present application, and details are not described herein again.

[0192] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0193] The embodiments of the present application provide a computer program product, which when running on a terminal device, enables the terminal device to implement the steps in the above-mentioned various method embodiments when executed.

[0194] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0195] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0196] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0197] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0198] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0199] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A target detection method, characterized in that, Including: Obtain the point cloud data and image data corresponding to the area to be measured; Extract features from the point cloud data to generate point cloud features corresponding to the area to be measured; Extract three-dimensional features from the image data to generate three-dimensional image features corresponding to the area to be measured; Fuse the point cloud features and the three-dimensional image features to generate fused features corresponding to the area to be measured; Process the fused features using a preset time alignment model to generate target fused features after time alignment; Perform regression and classification on the target fused features to generate target detection results corresponding to the area to be measured.

2. The method according to claim 1, wherein The preset time alignment model is a time alignment model based on a deformable attention mechanism. Before processing the fused features using the preset time alignment model to generate target fused features after time alignment, the method further includes: Obtain a training sample set, where the training sample set includes multiple training samples and the annotation data corresponding to each training sample; Input each training sample into the initial time alignment model based on the deformable attention mechanism to generate a predicted feature vector corresponding to each training sample; Determine the current loss value of the initial time alignment model based on the deformable attention mechanism according to the difference between the predicted feature vector corresponding to each training sample and the annotation data; Adjust the parameters of the initial time alignment model based on the deformable attention mechanism according to the current loss value to generate an updated time alignment model based on the deformable attention mechanism; Input each training sample into the updated time alignment model based on the deformable attention mechanism to continue training until the current loss value of the updated time alignment model based on the deformable attention mechanism is less than or equal to the loss threshold, and then determine the updated time alignment model based on the deformable attention mechanism as the preset time alignment model.

3. The method according to claim 1, characterized in that, Before fusing the point cloud features and the three-dimensional image features to generate fused features corresponding to the area to be measured, it further includes: Perform dimensional transformation on the point cloud features according to the dimension of the three-dimensional image features to make the dimension of the point cloud features match the dimension of the three-dimensional image features.

4. The method according to claim 1, wherein The extracting three-dimensional features from the image data to generate three-dimensional image features corresponding to the area to be measured includes: Perform two-dimensional image extraction on the image data to generate two-dimensional image features corresponding to the area to be measured; Perform three-dimensional mapping on the two-dimensional image features to generate the three-dimensional image features.

5. The method according to claim 4, characterized in that, The performing three-dimensional mapping on the two-dimensional image features to generate the three-dimensional image features includes: Determine the frustum point cloud data corresponding to the image data according to a preset depth probability vector and the two-dimensional image features, where the preset depth probability vector includes probability values corresponding to multiple preset depth values, and the frustum point cloud data includes the first spatial positions and feature vectors of multiple frustum points, and the first spatial position refers to the spatial position of the frustum point in the camera coordinate system; Determine the second spatial position of each of the cone points in the external coordinate system according to the first spatial position of each of the cone points, and the internal and external parameters of the camera; Determine the grid to which each of the cone points belongs according to the second spatial position of each of the cone points and a preset bird's-eye view grid map, where the preset bird's-eye view grid map includes a plurality of the grids; Determine the feature vector corresponding to each of the grids according to the feature vectors of the respective cone points included in each of the grids; Generate the three-dimensional image feature according to the feature vector corresponding to each of the grids.

6. The method according to any one of claims 1-5, characterized in that, The feature extraction of the point cloud data to generate the point cloud feature corresponding to the area to be measured includes: Perform voxelization processing on the point cloud data to determine a plurality of voxels corresponding to the point cloud data; Perform feature encoding on the point cloud data within each of the voxels to generate the point cloud feature, where the point cloud feature includes the point cloud features corresponding to the respective voxels.

7. The method according to claim 6, wherein Before performing the feature encoding on the point cloud data within each of the voxels to generate the point cloud feature, it further includes: Remove the voxels that contain zero point cloud data.

8. The method according to claim 6, characterized in that The performing feature encoding on the point cloud data within each of the voxels to generate the point cloud feature includes: Respectively sample a preset number of target point cloud data from each of the point cloud data included in each of the voxels; Perform feature encoding on the target point cloud data within each of the voxels to generate the point cloud feature.

9. The method according to claim 8, wherein The respectively sampling a preset number of target point cloud data from each of the point cloud data included in each of the voxels includes: When the number of point cloud data included in any one voxel is less than the preset number, perform filling processing on the point cloud data included in the any one voxel to generate the preset number of the target point cloud data.

10. A target detection device, characterized in that, Includes: A first acquisition module, configured to acquire point cloud data and image data corresponding to the area to be measured; A first feature extraction module, configured to extract the point cloud data feature to generate the point cloud feature corresponding to the area to be measured; A second feature extraction module, configured to extract the three-dimensional feature of the image data to generate the three-dimensional image feature corresponding to the area to be measured; A first fusion module, configured to fuse the point cloud feature and the three-dimensional image feature to generate the fusion feature corresponding to the area to be measured; A first alignment module, configured to process the fusion feature by using a preset time alignment model to generate a target fusion feature after time alignment; A first classification module, configured to perform regression and classification on the target fusion feature to generate the target detection result corresponding to the area to be measured.

11. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the processor implements the method according to any one of claims 1 to 9.

12. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 9.