Target detection method and device, terminal equipment and storage medium

By using multi-frame point cloud data and preset neural network models for object detection in object detection, the problem of low object detection accuracy caused by laser point cloud sparseness is solved, and higher object detection accuracy and adaptability are achieved.

CN120147601AInactive Publication Date: 2025-06-13VANJEE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311648151.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Due to the characteristics of the target and the sparsity of laser point clouds, it is difficult for the prior art to perform accurate object detection through a small number of point clouds.

Method used

The object detection is performed by obtaining multi-frame point cloud data in the detection area and calling the preset neural network model. The neural network model includes multiple weights, each weight is determined based on the attributes of the target to be detected (such as category, size, and speed). Through the fusion of multi-frame data and feature fusion processing, the accuracy of object detection is improved.

Benefits of technology

By combining multi-frame point cloud data for comprehensive detection, the accuracy of object detection is improved, and the weight of the neural network model is adaptively adjusted to adapt to the detection of targets of different attributes, further improving the detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147601A_ABST
    Figure CN120147601A_ABST
Patent Text Reader

Abstract

The invention provides a target detection method and device, terminal equipment and a storage medium, and the method comprises the steps: firstly obtaining multi-frame point cloud data in a detection region which comprises a to-be-detected target; and then a preset neural network model is called to perform target detection on the multi-frame point cloud data to determine a target detection result of the to-be-detected target, the preset neural network model comprises a plurality of weights, each weight is determined based on attributes of the to-be-detected target, and the attributes of the to-be-detected target comprise at least one of a category, a size and a speed. Therefore, the target detection result of the to-be-detected target is comprehensively obtained by combining the multi-frame point cloud data, and the accuracy of target detection is improved. In addition, the method also considers the influence of the attributes of the to-be-detected target on the merging of the multi-frame point cloud data on the target detection, adaptively adjusts the weight of the neural network model for the attributes of the to-be-detected target to adapt to the detection of targets with different attributes, and further improves the accuracy of target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of detection technology, and particularly relates to an object detection method, device, terminal device and storage medium. Background Art

[0002] Different from images, when representing an object in 3D point cloud, there is only point cloud data on the surface of the object. In the related art, due to the characteristics of the object and the sparsity of the laser point cloud, the number of point clouds that can hit the object may be small. However, it is very difficult to perform object detection with a small number of point clouds. Summary of the Invention

[0003] The embodiments of this application provide an object detection method, device, terminal device and storage medium, which can solve the problem of low object detection accuracy.

[0004] The first aspect of the embodiments of this application provides an object detection method, including: obtaining multiple frames of point cloud data in a detection area, where the detection area includes an object to be detected; calling a preset neural network model to perform object detection on the multiple frames of point cloud data to determine the object detection result of the object to be detected, where the preset neural network model includes multiple weights, and each weight is determined based on the attributes of the object to be detected, and the attributes of the object to be detected include at least one of category, size and speed.

[0005] Optionally, in a possible implementation manner of the first aspect, the above-mentioned calling a preset neural network model to perform object detection on multiple frames of point cloud data to determine the object detection result of the object to be detected includes:

[0006] Performing multiple downsampling processes on each frame of point cloud data respectively to obtain multiple scales of data to be fused corresponding to each frame of point cloud data;

[0007] Fusing the data to be fused corresponding to each frame of point cloud data according to multiple weights at each scale respectively to determine the fusion processing result at each scale;

[0008] Combining the fusion processing results at each scale to determine the object detection result of the object to be detected.

[0009] Optionally, in another possible implementation manner of the first aspect, before the above-mentioned performing multiple downsampling processes on each frame of point cloud data respectively to obtain multiple scales of data to be fused corresponding to each frame of point cloud data, it further includes:

[0010] Processing each frame of point cloud data by using a columnar feature extraction layer.

[0011] Optionally, in another possible implementation of the first aspect, the data to be fused is a feature map to be fused, and the fusion processing result is a feature fusion result. The above-mentioned fusion processing of the data to be fused corresponding to each frame of point cloud data according to multiple weights at each scale to determine the fusion processing result of each scale includes:

[0012] Sequentially from the lowest scale to the highest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data according to multiple weights to determine the feature fusion result of each scale. Among them, feature fusion at scales other than the lowest scale is based on the feature fusion result of the previous scale.

[0013] Optionally, in another possible implementation of the first aspect, the above-mentioned sequentially from the lowest scale to the highest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data according to multiple weights to determine the feature fusion result of each scale, including:

[0014] For the lowest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data according to multiple weights to determine the corresponding feature fusion result;

[0015] For each other scale except the lowest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data and the upsampling result corresponding to the previous scale according to multiple weights to determine the corresponding feature fusion result, where the upsampling result corresponding to the previous scale is obtained by performing upsampling processing on the feature fusion result corresponding to the previous scale.

[0016] Optionally, in another possible implementation of the first aspect, the above-mentioned for each other scale except the lowest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data and the upsampling result corresponding to the previous scale according to multiple weights, including:

[0017] Connect the feature maps to be fused corresponding to each frame of point cloud data and the upsampling result corresponding to the previous scale to obtain the connection result corresponding to each frame of point cloud data;

[0018] Perform convolution on the connection result corresponding to each frame of point cloud data to obtain the convolution fusion feature corresponding to each frame of point cloud data;

[0019] Perform feature fusion on the convolution fusion features corresponding to each frame of point cloud data according to multiple weights.

[0020] Optionally, in another possible implementation of the first aspect, the above-mentioned sequentially from the lowest scale to the highest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data according to multiple weights to determine the feature fusion result of each scale, including:

[0021] From the lowest scale to the highest scale in sequence, according to multiple weights, use the cross-attention mechanism to perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data, so as to determine the feature fusion result of each scale.

[0022] Optionally, in another possible implementation manner of the first aspect, the above-mentioned multi-frame point cloud data includes the point cloud data of the historical frame and the point cloud data of the current frame, and the multiple weights include the query weight of the historical frame, the key weight of the current frame, and the value weight of the current frame.

[0023] Optionally, in another possible implementation manner of the first aspect, the above-mentioned from the lowest scale to the highest scale in sequence, according to multiple weights, use the cross-attention mechanism to perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data, so as to determine the feature fusion result of each scale, including:

[0024] From the lowest scale to the highest scale in sequence, based on the query weight of the historical frame, the key weight of the current frame, and the value weight of the current frame, determine the query feature matrix of the historical frame, the key feature matrix of the current frame, and the value feature matrix of the current frame;

[0025] Multiply the query feature matrix of the historical frame by the key feature matrix of the current frame to determine the attention matrix;

[0026] Multiply the attention matrix by the value feature matrix of the current frame to determine the feature fusion result of the current scale.

[0027] Optionally, in another possible implementation manner of the first aspect, the above-mentioned combining the fusion processing results of each scale to determine the target detection result of the target to be detected includes:

[0028] Perform deconvolution on the fusion processing result of each scale respectively to obtain the deconvolution image corresponding to the fusion processing result of each scale;

[0029] Connect the deconvolution images corresponding to the fusion processing results of each scale to determine the target detection result of the target to be detected.

[0030] Optionally, in another possible implementation manner of the first aspect, the above-mentioned target detection result of the target to be detected includes the bounding box position information and category information of the target to be detected.

[0031] The second aspect of the embodiments of the present application provides a target detection device, including:

[0032] A data acquisition module, configured to acquire multi-frame point cloud data within a detection area, where the detection area includes a target to be detected;

[0033] A model calling module is used to call a preset neural network model to perform object detection on multiple frames of point cloud data to determine the object detection result of the object to be detected. The preset neural network model includes multiple weights, and each weight is determined based on the attributes of the object to be detected. The attributes of the object to be detected include at least one of category, size, and speed.

[0034] In the third aspect of the embodiments of the present application, a terminal device is provided, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the object detection method in the first aspect is implemented.

[0035] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the object detection method in the first aspect is implemented.

[0036] In the fifth aspect of the embodiments of the present application, a computer program product is provided. When the computer program product runs on a terminal device, the terminal device is enabled to execute the object detection method in the first aspect.

[0037] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: The embodiments of the present application provide an object detection method, device, terminal device, and storage medium. Among them, the method first obtains multiple frames of point cloud data in a detection area, where the detection area includes an object to be detected; then calls a preset neural network model to perform object detection on the multiple frames of point cloud data to determine the object detection result of the object to be detected. The preset neural network model includes multiple weights, and each weight is determined based on the attributes of the object to be detected. The attributes of the object to be detected include at least one of category, size, and speed. Thus, by combining multiple frames of point cloud data, the object detection result of the object to be detected is comprehensively obtained, improving the accuracy of object detection. In addition, the embodiments of the present application also consider the influence of the attributes of the object to be detected on the merging of multiple frames of point cloud data for object detection, and adaptively adjust the weights of the neural network model according to the attributes of the object to be detected to adapt to the detection of objects with different attributes, further improving the accuracy of object detection. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 It is a schematic flowchart of an object detection method provided by an embodiment of the present application;

[0040] Figure 2 is a schematic flowchart of some steps in a target detection method provided by an embodiment of the present application;

[0041] Figure 3 is an example diagram of cross-attention calculation provided by an embodiment of the present application;

[0042] Figure 4 is an example diagram of a target detection method provided by an embodiment of the present application;

[0043] Figure 5 is an example diagram of feature transfer provided by an embodiment of the present application;

[0044] Figure 6 is a schematic structural diagram of a target detection device provided by an embodiment of the present application;

[0045] Figure 7 is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0046] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are set forth in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.

[0047] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0048] It should also be understood that the term "and / or" as used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0049] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detected [the described condition or event]" can be interpreted as meaning "once determined" or "in response to determining" or "once detected [the described condition or event]" or "in response to detecting [the described condition or event]" depending on the context.

[0050] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for differential description and should not be construed as indicating or implying relative importance.

[0051] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other some embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0052] It should be understood that the magnitude of the sequence numbers of the steps in this embodiment does not mean the sequence of execution, and the execution sequence of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0053] In the related art, due to the characteristics of the target and the sparsity of the laser point cloud, the number of point clouds that can hit the target may be small. However, it is very difficult to perform target detection with a small number of point clouds.

[0054] In view of this, the embodiments of the present application provide a target detection method, device, terminal device and storage medium. By combining multi-frame point cloud data to comprehensively obtain the target detection result of the target to be detected, the accuracy of target detection is improved. In addition, the embodiments of the present application also consider the influence of the attributes of the target to be detected on the merging of multi-frame point cloud data for target detection, and adaptively adjust the weights of the neural network model according to the attributes of the target to be detected to adapt to the detection of targets with different attributes, further improving the accuracy of target detection.

[0055] In order to illustrate the technical solution of the present application, the following will be described by specific embodiments.

[0056] Refer to Figure 1 , which shows a schematic flow chart of a target detection method provided by an embodiment of the present application.

[0057] As Figure 1 shown, the target detection method may include the following steps:

[0058] Step 101, obtain multi-frame point cloud data within the detection area.

[0059] Among them, the detection area includes the target to be detected.

[0060] In an embodiment of the present application, a lidar can be used to collect three-dimensional information of all objects in a detection area and represent it as point cloud data. A point cloud is a set composed of a series of three-dimensional coordinate points, and each point represents a surface point in the environment.

[0061] It should be noted that due to the sparsity of the lidar point cloud, the number of point clouds that can hit the target may be small, and it is very difficult to perform target detection with a small number of point clouds. Therefore, multiple frames of point cloud data can be continuously collected to obtain relatively rich point cloud data of the target to be detected and used for subsequent target detection, thereby improving the accuracy of target detection.

[0062] Step 102: Invoke a preset neural network model to perform target detection on multiple frames of point cloud data to determine the target detection result of the target to be detected.

[0063] In the prior art, usually multiple frames of point cloud data are directly merged into one frame, which can effectively increase the density of the target point cloud, and then target detection is performed. However, this method is only applicable to some stationary targets, or targets with a large volume and slow speed in a moving state. For high-speed targets or small-volume targets in a moving state, directly merging multiple frames of point cloud data will cause ghosting of the target point cloud, which instead reduces the accuracy of target detection.

[0064] Therefore, the preset neural network model in the embodiment of the present application includes multiple weights, and each weight is determined based on the attributes of the target to be detected. Among them, the attributes of the target to be detected may include at least one of category, size, and speed. In this way, when performing target detection, the weights in the preset neural network model can be adaptively adjusted to make the multiple frames of features adaptively aligned. This method is applicable to targets in different states, types, and sizes. Especially when the target to be detected is a high-speed target or a small-volume target in a moving state, it can significantly improve the accuracy of target detection.

[0065] Among them, the category of the target to be detected can indirectly reflect the size and speed of the target to be detected. Taking the target detection in an urban road scenario as an example, the categories of the target to be detected may include: vehicles, pedestrians, bicycles, motorcycles, buses, large trucks, buildings, traffic signs, etc. If the target to be detected is a pedestrian, it can be determined that the target to be detected is a low-speed small target in a moving state; if the target to be detected is a motorcycle, it can be determined that the target to be detected is a high-speed small target in a moving state; if the target to be detected is a building, it can be determined that the target to be detected is a stationary target.

[0066] In the embodiments of the present application, the object detection result of the object to be detected includes the bounding box position information and category information of the object to be detected. Specifically, a detection head can be used for regression and classification to obtain the final object detection result. Among them, regression refers to predicting the position of the bounding box corresponding to the object to be detected, such as the upper left corner coordinates and lower right corner coordinates of the bounding box, or the center point coordinates and the width and height of the bounding box. Classification refers to predicting the category of the object to be detected. It should be understood that the category can be set in advance according to the actual application scenario and requirements, and the result of category prediction can be one or more. In the case of predicting multiple categories, the probability of each category can be output at the same time.

[0067] The object detection method disclosed in the above embodiments of the present application first obtains multiple frames of point cloud data in the detection area, where the detection area includes the object to be detected; then calls a preset neural network model to perform object detection on the multiple frames of point cloud data to determine the object detection result of the object to be detected, where the preset neural network model includes multiple weights, and each weight is determined based on the attributes of the object to be detected, and the attributes of the object to be detected include at least one of category, size, and speed. Thus, by combining multiple frames of point cloud data, the object detection result of the object to be detected is comprehensively obtained, improving the accuracy of object detection. In addition, the embodiments of the present application also consider the influence of the attributes of the object to be detected on the object detection when merging multiple frames of point cloud data, and adaptively adjust the weights of the neural network model according to the attributes of the object to be detected to adapt to the detection of objects with different attributes, further improving the accuracy of object detection.

[0068] See Figure 2 , which shows a schematic flowchart of an object detection method provided by an embodiment of the present application. Figure 2 It can be regarded as an example of step 102. As Figure 2 shown, the object detection method may include the following steps:

[0069] Step 201, perform multiple downsampling processes on each frame of point cloud data to obtain multiple scales of data to be fused corresponding to each frame of point cloud data.

[0070] It should be understood that the number of downsampling times can be set according to the actual application scenario and requirements. For example, by performing two downsampling processes, 3 scales of data to be fused corresponding to each frame of point cloud data can be obtained.

[0071] In a possible implementation, feature extraction can be performed based on PointPillars from the Bird's Eye View (BEV) perspective. Among them, PointPillars is a three-dimensional object detection algorithm based on point cloud data, where Point represents the lidar point cloud data, and Pillars represents the cylinders obtained by converting the point cloud data into a two-dimensional image representation. That is, before step 201, the Pillar Feature Extraction (PFE) layer can also be used to process each frame of the point cloud data and convert it into a pseudo-image from the BEV perspective. The bird's eye view is a representation method that projects three-dimensional point cloud data onto a two-dimensional plane. In three-dimensional object detection, converting point cloud data into a bird's eye view can retain the geometric information of the scene while simplifying the data processing and feature extraction processes. In the bird's eye view perspective, the horizontal plane corresponds to the ground plane, and the point cloud data is projected onto this plane to form a two-dimensional dense image.

[0072] Correspondingly, if each frame of the point cloud data is converted into a pseudo-image from the bird's eye view perspective before step 201, then multiple-scale feature maps to be fused corresponding to each frame of the point cloud data can be obtained through step 201.

[0073] Step 202: At each scale, fuse the data to be fused corresponding to each frame of the point cloud data according to multiple weights to determine the fusion result at each scale.

[0074] Among them, when the data to be fused is the feature map to be fused, what is finally obtained in step 202 is the feature fusion result at each scale.

[0075] Furthermore, for multiple frames of feature maps to be fused at the same scale, feature fusion is performed. The fused features can then be gradually passed to a larger scale for feature transfer. In this way, weights can be added as a buffer during the multi-scale transfer, so that the multi-scale features can be better aligned. That is, as a possible implementation manner of the embodiment of the present application, the above step 202 may include: sequentially from the lowest scale to the highest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of the point cloud data according to multiple weights to determine the feature fusion result at each scale, where feature fusion at other scales except the lowest scale is based on the feature fusion result of the previous scale.

[0076] Specifically, for the multi-frame feature maps to be fused at the lowest scale, feature fusion can be performed according to multiple weights to determine the feature fusion result corresponding to the lowest scale. Then, the feature fusion result at the lowest scale can be used for feature transfer to the next higher scale through upsampling (Upsample), and thus the feature fusion of all scales can be completed in sequence. That is, as a possible implementation manner of the embodiments of the present application, for the lowest scale, feature fusion is performed on the feature maps to be fused corresponding to each frame of point cloud data according to multiple weights to determine the corresponding feature fusion result; for each other scale except the lowest scale, according to multiple weights, feature fusion is performed on the feature maps to be fused corresponding to each frame of point cloud data and the upsampling result corresponding to the previous scale to determine the corresponding feature fusion result, where the upsampling result corresponding to the previous scale is obtained by performing upsampling processing on the feature fusion result corresponding to the previous scale.

[0077] In a possible implementation manner, after obtaining the feature fusion result at a certain scale, feature transfer can be performed to a larger scale through upsampling and convolution (Conv). Specifically, for each other scale except the lowest scale, the feature maps to be fused corresponding to each frame of point cloud data and the upsampling result corresponding to the previous scale can be concatenated to obtain the concatenation result corresponding to each frame of point cloud data; convolution is performed on the concatenation result corresponding to each frame of point cloud data to obtain the convolution fusion feature corresponding to each frame of point cloud data; according to multiple weights, feature fusion is performed on the convolution fusion feature corresponding to each frame of point cloud data.

[0078] Among them, the above-mentioned concatenation of the feature maps to be fused corresponding to each frame of point cloud data and the upsampling result corresponding to the previous scale can adopt the Concat operation. Concat refers to the operation of concatenating or splicing multiple features along a certain dimension, which is used to splice features from different sources or different levels to generate a richer feature representation.

[0079] In a possible implementation manner, feature fusion of the multi-frame feature maps at the same scale can be performed through a cross-attention mechanism. That is to say, starting from the lowest scale to the highest scale in sequence, according to multiple weights, the cross-attention mechanism is used to perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data to determine the feature fusion result at each scale.

[0080] It should be noted that the cross-attention mechanism uses two different sequences as query and key-value pairs to process the relationship between these two sequences. In the embodiments of the present application, the multi-frame point cloud data includes the point cloud data of the historical frame and the point cloud data of the current frame. Correspondingly, the multiple weights in the preset neural network model include the query weight of the historical frame, the key weight of the current frame, and the value weight of the current frame. Through the cross-attention mechanism, multi-frame information can be better introduced, and feature fusion between the current frame and the historical frame can be performed.

[0081] The following combines with Figure 3 to illustrate the specific calculation process of cross-attention. As Figure 3 shown, where Frame t represents the feature map to be fused in the current frame, and Frame t-n represents the feature map to be fused in the historical frame. Among them, the historical frame is the first n frames before the current frame, which can be the first three frames or the first two frames. The embodiments of the present application do not limit this. The feature map to be fused in the historical frame Frame t-n corresponds to the query weight Wq in the cross-attention mechanism, and the feature map to be fused in the current frame Frame t corresponds to the key weight Wk and the value weight Wv in the cross-attention mechanism. Starting from the lowest scale to the highest scale in sequence, based on the query weight Wq of the historical frame, the key weight Wk of the current frame, and the value weight Wv of the current frame, determine the query feature matrix Fq of the historical frame, the key feature matrix Fk of the current frame, and the value feature matrix Fv of the current frame; multiply the query feature matrix Fq of the historical frame by the key feature matrix Fk of the current frame to determine the attention matrix Attention Matrix; multiply the attention matrix Attention Matrix by the value feature matrix Fv of the current frame to determine the feature fusion result Attention output at the current scale. Among them, matrix multiplication can be implemented by calling the matmul() function.

[0082] Further, the above-mentioned feature fusion result at the current scale calculated using the cross-attention mechanism can be specifically carried out through the following formula:

[0083]

[0084] where, T represents transpose; is the scaling factor; d k represents the dimension of the key feature matrix; the softmax() function is used for normalization processing.

[0085] Using the cross-attention mechanism can better introduce multi-frame information. The weights can be adaptively adjusted by using a neural network to make the multi-frame features adaptively aligned, which can avoid the negative impacts brought by high-speed targets or small targets, and thus improve the accuracy of object detection.

[0086] Step 203, combine the fusion processing results of each scale to determine the object detection result of the object to be detected.

[0087] It should be noted that after obtaining the fusion processing results of each scale, deconvolution (Deconv) is performed on each fusion processing result respectively, and finally a Concat operation is used for connection to finally obtain the object detection result of the object to be detected. That is, as a possible implementation manner of the embodiment of the present application, step 203 above may include: performing deconvolution on the fusion processing result of each scale respectively to obtain a deconvolution image corresponding to the fusion processing result of each scale; connecting the deconvolution images corresponding to the fusion processing results of each scale to determine the object detection result of the object to be detected.

[0088] For the object detection method disclosed in the above embodiment of the present application, first, multiple downsampling processes are respectively performed on each frame of point cloud data to obtain multiple scales of data to be fused corresponding to each frame of point cloud data; then, at each scale, the data to be fused corresponding to each frame of point cloud data is fused according to multiple weights to determine the fusion processing result of each scale; finally, in combination with the fusion processing results of each scale, the object detection result of the object to be detected is determined. Thus, multiple scales of data to be fused are obtained through downsampling, and then the multi-frame data to be fused of the same scale is fused, and after fusion, it is passed step by step to a larger scale. In this way, weights can be added as a buffer during the multi-scale transmission, so that the multi-scale data can be better aligned, thereby improving the accuracy of object detection.

[0089] Refer to Figure 4 , which shows an example diagram of a complete object detection method provided by an embodiment of the present application. Among them, Figure 4 Taking two downsamplings, 3 scales, the current frame and the previous 3 frames as an example for illustration. As Figure 4 shown, through two downsamplings, 3 scales of feature maps to be fused of each frame of point cloud data are obtained. Then, the multi-frame features of the same scale are fused through a cross-attention mechanism in the multi-frame information fusion module, that is, the Feature Aggregation module. The fused features are passed to a larger scale through upsampling and convolution to obtain the fusion processing result of each scale. Finally, through deconvolution and Concat operations, each scale is combined and regression and classification are performed through a detection head to obtain the final object detection result.

[0090] Refer to Figure 5 , which shows an example diagram of a feature transmission provided by an embodiment of the present application. Figure 5 Taking the current frame and the previous 3 frames as an example for illustration. As Figure 5As shown, after the cross-attention of the previous scale for four consecutive frames, the aggregated feature of the previous scale, Previous Aggregated Feature, is obtained. The result of the aggregation part will pass through an upsampling Upsample to align with the feature size of a larger scale, and then through Concat and convolution operations with the feature of the current scale, Current Frame Feature, to obtain the fused feature Fused Feature. Then, the cross-attention of the current scale is performed.

[0091] See Figure 6 , which shows a schematic structural diagram of an object detection device provided by an embodiment of the present application. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.

[0092] The object detection device may specifically include the following modules:

[0093] The data acquisition module 601 is used to acquire multiple frames of point cloud data within the detection area, where the detection area includes the target to be detected.

[0094] The model calling module 602 is used to call a preset neural network model to perform object detection on the multiple frames of point cloud data to determine the object detection result of the target to be detected. The preset neural network model includes multiple weights, and each weight is determined based on the attributes of the target to be detected. The attributes of the target to be detected include at least one of category, size, and speed.

[0095] The object detection device disclosed in the above embodiments of the present application first acquires multiple frames of point cloud data within the detection area, where the detection area includes the target to be detected; then calls a preset neural network model to perform object detection on the multiple frames of point cloud data to determine the object detection result of the target to be detected. The preset neural network model includes multiple weights, and each weight is determined based on the attributes of the target to be detected. The attributes of the target to be detected include at least one of category, size, and speed. Thus, by combining multiple frames of point cloud data, the object detection result of the target to be detected is comprehensively obtained, improving the accuracy of object detection. In addition, the embodiments of the present application also consider the influence of the attributes of the target to be detected on the merging of multiple frames of point cloud data for object detection, and adaptively adjust the weights of the neural network model according to the attributes of the target to be detected to adapt to the detection of targets with different attributes, further improving the accuracy of object detection.

[0096] In a possible implementation manner of the embodiment of the present application, the above model calling module 602 may specifically include the following sub-modules:

[0097] The first processing sub-module is used to perform multiple downsampling processes on each frame of point cloud data respectively to obtain multiple scales of data to be fused corresponding to each frame of point cloud data.

[0098] A second processing sub-module, configured to perform fusion processing on the data to be fused corresponding to each frame of point cloud data according to multiple weights at each scale, so as to determine the fusion processing result of each scale.

[0099] A third processing sub-module, configured to determine the target detection result of the target to be detected by combining the fusion processing results of each scale.

[0100] In another possible implementation manner of the embodiment of the present application, the above-mentioned model calling module 602 may specifically further include the following sub-modules:

[0101] A fourth processing sub-module, configured to process each frame of point cloud data by using a columnar feature extraction layer.

[0102] In another possible implementation manner of the embodiment of the present application, the above-mentioned second processing sub-module may specifically include the following units:

[0103] A first processing unit, configured to perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data according to multiple weights from the lowest scale to the highest scale in sequence, so as to determine the feature fusion result of each scale, wherein, for scales other than the lowest scale, feature fusion is performed based on the feature fusion result of the previous scale.

[0104] In another possible implementation manner of the embodiment of the present application, the above-mentioned first processing unit is specifically configured to: for the lowest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data according to multiple weights, so as to determine the corresponding feature fusion result; for each other scale except the lowest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data and the upsampling result corresponding to the previous scale according to multiple weights, so as to determine the corresponding feature fusion result, wherein the upsampling result corresponding to the previous scale is obtained by performing upsampling processing on the feature fusion result corresponding to the previous scale.

[0105] In another possible implementation manner of the embodiment of the present application, the above-mentioned first processing unit is specifically configured to: concatenate the feature maps to be fused corresponding to each frame of point cloud data and the upsampling result corresponding to the previous scale, to obtain the concatenation result corresponding to each frame of point cloud data; perform convolution on the concatenation result corresponding to each frame of point cloud data, to obtain the convolution fusion feature corresponding to each frame of point cloud data; perform feature fusion on the convolution fusion feature corresponding to each frame of point cloud data according to multiple weights.

[0106] In another possible implementation manner of the embodiment of the present application, the above-mentioned first processing unit is specifically configured to: perform feature fusion on the feature maps to be fused corresponding to each frame of point cloud data according to multiple weights from the lowest scale to the highest scale in sequence by using a cross-attention mechanism, so as to determine the feature fusion result of each scale.

[0107] In another possible implementation manner of the embodiment of the present application, the above-mentioned multi-frame point cloud data includes the point cloud data of the historical frame and the point cloud data of the current frame, and the multiple weights include the query weight of the historical frame, the key weight of the current frame, and the value weight of the current frame.

[0108] In yet another possible implementation manner of the embodiment of the present application, the above-mentioned first processing unit is specifically configured to: sequentially from the lowest scale to the highest scale, based on the query weight of the historical frame, the key weight of the current frame, and the value weight of the current frame, determine the query feature matrix of the historical frame, the key feature matrix of the current frame, and the value feature matrix of the current frame; multiply the query feature matrix of the historical frame by the key feature matrix of the current frame to determine the attention matrix; multiply the attention matrix by the value feature matrix of the current frame to determine the feature fusion result of the current scale.

[0109] In still another possible implementation manner of the embodiment of the present application, the above-mentioned third processing sub-module may specifically further include the following units:

[0110] A second processing unit, configured to perform deconvolution on the fusion processing result of each scale respectively to obtain a deconvolution image corresponding to the fusion processing result of each scale.

[0111] A third processing unit, configured to concatenate the deconvolution images corresponding to the fusion processing results of each scale to determine the target detection result of the target to be detected.

[0112] In another possible implementation manner of the embodiment of the present application, the target detection result of the target to be detected includes the bounding box position information and the category information of the target to be detected.

[0113] The target detection device disclosed in the above embodiment of the present application first performs multiple downsampling processes on each frame of point cloud data respectively to obtain multiple scales of data to be fused corresponding to each frame of point cloud data; then, at each scale, fuses the data to be fused corresponding to each frame of point cloud data according to multiple weights to determine the fusion processing result of each scale; finally, combines the fusion processing results of each scale to determine the target detection result of the target to be detected. Thus, by downsampling, multiple scales of data to be fused are obtained, then the multi-frame data to be fused of the same scale are fused, and after fusion, it is passed to a larger scale step by step. In this way, weights can be added as a buffer during the multi-scale transfer, so that the multi-scale data can be better aligned, and thus the accuracy of target detection can be improved.

[0114] The target detection device provided in the embodiment of the present application can be applied in the foregoing method embodiment. For details, refer to the description of the foregoing method embodiment, which will not be elaborated here.

[0115] Figure 7It is a schematic structural diagram of the terminal device provided in the fourth embodiment of this application. As Figure 7 shown, the terminal device 700 in this embodiment includes: at least one processor 710 ( Figure 7 only one processor is shown in the figure), a memory 720, and a computer program 721 stored in the memory 720 and executable on the at least one processor 710. When the processor 710 executes the computer program 721, it implements the steps in the embodiment of the above target detection method.

[0116] The terminal device 700 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor 710 and a memory 720. Those skilled in the art can understand that Figure 7 this is only an example of the terminal device 700, and does not constitute a limitation on the terminal device 700. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0117] The so-called processor 710 may be a central processing unit (CPU), and the processor 710 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0118] The memory 720 may be an internal storage unit of the terminal device 700 in some embodiments, such as the hard disk or memory of the terminal device 700. The memory 720 may also be an external storage device of the terminal device 700 in other embodiments, such as a plug-in hard disk equipped on the terminal device 700, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 720 may include both the internal storage unit of the terminal device 700 and the external storage device. The memory 720 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory 720 may also be used to temporarily store data that has been output or will be output.

[0119] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0120] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0121] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0122] In the embodiments provided in the present application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0123] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0124] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0125] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by a computer program instructing the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0126] All or part of the processes in the methods of the above embodiments can also be implemented by a computer program product. When the computer program product runs on a terminal device, the terminal device can execute the steps in the above method embodiments when executed.

[0127] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A target detection method, characterized in that, it includes: Obtain multiple frames of point cloud data within the detection area, where the detection area includes the target to be detected; Call a preset neural network model to perform target detection on the multiple frames of point cloud data to determine the target detection result of the target to be detected, where the preset neural network model includes multiple weights, and each weight is determined based on the attributes of the target to be detected, and the attributes of the target to be detected include at least one of category, size, and speed.

2. The target detection method according to claim 1, characterized in that, The step of calling a preset neural network model to perform target detection on the multiple frames of point cloud data to determine the target detection result of the target to be detected includes: Perform multiple downsampling processes on each frame of the point cloud data respectively to obtain multiple scales of data to be fused corresponding to each frame of the point cloud data; At each scale, perform fusion processing on the data to be fused corresponding to each frame of the point cloud data according to the multiple weights to determine the fusion processing result at each scale; Combine the fusion processing results at each scale to determine the target detection result of the target to be detected.

3. The target detection method according to claim 2, characterized in that, Before the step of performing multiple downsampling processes on each frame of the point cloud data respectively to obtain multiple scales of data to be fused corresponding to each frame of the point cloud data, it further includes: Process each frame of the point cloud data using a columnar feature extraction layer.

4. The target detection method according to claim 2, characterized in that, The data to be fused is a feature map to be fused, and the fusion processing result is a feature fusion result. The step of performing fusion processing on the data to be fused corresponding to each frame of the point cloud data according to the multiple weights at each scale to determine the fusion processing result at each scale includes: Sequentially from the lowest scale to the highest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of the point cloud data according to the multiple weights to determine the feature fusion result at each scale, where at scales other than the lowest scale, feature fusion is performed based on the feature fusion result of the previous scale.

5. The target detection method according to claim 4, characterized in that, The step of sequentially from the lowest scale to the highest scale, performing feature fusion on the feature maps to be fused corresponding to each frame of the point cloud data according to the multiple weights to determine the feature fusion result at each scale includes: For the lowest scale, perform feature fusion on the feature maps to be fused corresponding to each frame of the point cloud data according to the multiple weights to determine the corresponding feature fusion result; For each other scale except the lowest scale, according to the multiple weights, perform feature fusion on the feature maps to be fused corresponding to each frame of the point cloud data and the upsampling result corresponding to the previous scale to determine the corresponding feature fusion result, where the upsampling result corresponding to the previous scale is obtained by performing upsampling processing on the feature fusion result corresponding to the previous scale.

6. The target detection method according to claim 5, characterized in that, For each scale other than the lowest scale, according to the multiple weights, perform feature fusion on the to-be-fused feature maps corresponding to each frame of the point cloud data and the upsampling result corresponding to the previous scale, including: Connect the to-be-fused feature maps corresponding to each frame of the point cloud data and the upsampling result corresponding to the previous scale to obtain the connection result corresponding to each frame of the point cloud data; Perform convolution on the connection result corresponding to each frame of the point cloud data to obtain the convolution fusion feature corresponding to each frame of the point cloud data; According to the multiple weights, perform feature fusion on the convolution fusion features corresponding to each frame of the point cloud data.

7. The object detection method according to claim 4, wherein, The sequentially performing, from the lowest scale to the highest scale, feature fusion on the to-be-fused feature maps corresponding to each frame of the point cloud data according to the multiple weights to determine the feature fusion result of each scale includes: Sequentially from the lowest scale to the highest scale, according to the multiple weights, use the cross-attention mechanism to perform feature fusion on the to-be-fused feature maps corresponding to each frame of the point cloud data to determine the feature fusion result of each scale.

8. The object detection method according to claim 7, wherein, The multi-frame point cloud data includes the point cloud data of the historical frame and the point cloud data of the current frame, and the multiple weights include the query weight of the historical frame, the key weight of the current frame, and the value weight of the current frame.

9. The object detection method according to claim 8, wherein, The sequentially performing, from the lowest scale to the highest scale, feature fusion on the to-be-fused feature maps corresponding to each frame of the point cloud data according to the multiple weights, using the cross-attention mechanism to determine the feature fusion result of each scale includes: Sequentially from the lowest scale to the highest scale, based on the query weight of the historical frame, the key weight of the current frame, and the value weight of the current frame, determine the query feature matrix of the historical frame, the key feature matrix of the current frame, and the value feature matrix of the current frame; Multiply the query feature matrix of the historical frame and the key feature matrix of the current frame to determine the attention matrix; Multiply the attention matrix and the value feature matrix of the current frame to determine the feature fusion result of the current scale.

10. The object detection method according to claim 4, wherein, The combining the fusion processing results of each scale to determine the object detection result of the to-be-detected object includes: Perform deconvolution on the fusion processing result of each scale respectively to obtain the deconvolution image corresponding to the fusion processing result of each scale; Connect the deconvolution images corresponding to the fusion processing results of each scale to determine the object detection result of the to-be-detected object.

11. The object detection method according to any one of claims 1-10, wherein, The object detection result of the to-be-detected object includes the bounding box position information and the category information of the to-be-detected object.

12. An object detection device, wherein, comprising: A data acquisition module for acquiring multi-frame point cloud data within a detection area, wherein the detection area includes an object to be detected; A model calling module, configured to call a preset neural network model to perform object detection on the multi-frame point cloud data, so as to determine the object detection result of the object to be detected, wherein the preset neural network model includes a plurality of weights, and each weight is determined based on the attributes of the object to be detected, and the attributes of the object to be detected include at least one of category, size, and speed.

13. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, the method according to any one of claims 1 to 11 is implemented.

14. A computer-readable storage medium, storing a computer program, wherein, when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.