3D target detection method based on spiking neural network

Through the feature fusion and parsing method based on pulse neural network, the dependence problem of hardware computing power and inference time in 3D target detection is solved, and high-precision detection with low energy consumption and low latency is achieved.

CN120783020APending Publication Date: 2025-10-14CHINA FAW CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510837487.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing 3D object detection methods rely on high hardware computing power and long inference time, resulting in energy consumption and latency issues not being effectively addressed.

Method used

A method based on pulse neural networks is used to obtain features from multi-view images and radar point clouds, perform feature fusion and enhancement processing, and use the detection head to analyze features to obtain 3D detection target information. Combined with an event-driven mechanism and sparse pulse coding, it reduces computing energy consumption and inference latency.

Benefits of technology

It significantly reduces computing energy consumption and inference latency while maintaining high detection accuracy, achieving low-power, low-latency 3D object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783020A_ABST
    Figure CN120783020A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and artificial intelligence, in particular to a pulse neural network-based 3D target detection method, which comprises the following steps of: acquiring a first target feature of a multi-view image under an aerial view view and a second target feature of a radar point cloud under the aerial view view, and performing feature fusion to obtain a first target feature and a second target feature; the method comprises the steps of obtaining initial multi-modal fusion features of a multi-view image, performing feature enhancement processing on the initial fusion features by using a preset feature encoder to obtain final multi-modal fusion features of the multi-view image, and analyzing the final multi-modal fusion features by using a detection head to obtain target information of a 3D detection target. Therefore, the problems of dependence on hardware computing power and reasoning time, task detection energy consumption, time delay and the like caused by a detection method in the related technology are solved, the calculation energy consumption and the reasoning delay are remarkably reduced through an event driving mechanism and sparse pulse coding in combination with multi-modal feature fusion, and meanwhile high detection precision is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computer vision and artificial intelligence technology, and in particular to a 3D target detection method based on a pulse neural network. Background Art

[0002] With the continuous updating and iteration of technology in recent years, novel solution designs and efficient optimization algorithms have alleviated this situation, but the essence of high power consumption and high latency of the detection algorithm has not changed.

[0003] In related technologies, target detection methods generally use artificial neural networks as the network backbone for feature extraction, and improve detection accuracy by continuously increasing the network scale to obtain richer semantic features.

[0004] However, the above methods cause dependence on hardware computing power and inference time, especially in 3D object detection tasks, which require huge energy consumption and time delay, and need to be solved urgently. Summary of the Invention

[0005] This application provides a 3D target detection method based on a pulse neural network to solve the problems of dependence on hardware computing power and inference time, as well as task detection energy consumption and time delay caused by the detection methods of related technologies.

[0006] The first embodiment of the present application provides a 3D object detection method based on a spiking neural network, comprising the following steps: Acquire a first target feature of the multi-view image in a bird's-eye view and a second target feature of the radar point cloud in a bird's-eye view; Performing feature fusion on the first target feature and the second target feature to obtain an initial multimodal fusion feature of the multi-view image, and performing feature enhancement processing on the initial fusion feature using a preset feature encoder to obtain a final multimodal fusion feature of the multi-view image; The final multimodal fusion feature is input to a detection head, so as to utilize the detection head to analyze the final multimodal fusion feature and obtain target information of a 3D detection target.

[0007] Furthermore, in some embodiments, obtaining the first target feature of the multi-view image from a bird's-eye view includes: Extracting image features from the multi-view images using a 2D backbone network based on a preset spiking neural network, and performing dimensionality conversion on the extracted image features to obtain target dimensional image features of the multi-view images; Predicting the discrete depth distribution of each pixel in the target dimension image feature using a preset algorithm to obtain the depth probability corresponding to each pixel; The target dimension image feature is rescaled based on the depth probability corresponding to each pixel point, and the scaled target dimension image feature is pooled to obtain the first target feature of the multi-view image under the bird's-eye view perspective.

[0008] Furthermore, in some embodiments, obtaining the second target feature of the radar point cloud from a bird's-eye view perspective includes: Acquiring point cloud information of the radar point cloud, and performing data processing on the point cloud information based on a preset voxel size to obtain voxelized point cloud information; Extracting point cloud features from the voxelized point cloud information using a 3D backbone network based on a preset spiking neural network, and performing dimension conversion on the extracted point cloud features to obtain target dimension point cloud features of the radar point cloud; The target dimension point cloud features are compressed based on the target axis and projected into the bird's-eye view space to obtain the second target features of the radar point cloud under the bird's-eye view perspective.

[0009] Furthermore, in some embodiments, inputting the final multimodal fusion feature into a detection head, and using the detection head to parse the final multimodal fusion feature to obtain target information of a 3D detection target, includes: Predicting a 3D bounding box of the 3D detection target and a category of the 3D detection target, and outputting a probability distribution of each category of the 3D detection target; Based on the probability distribution of each category of the 3D detection target, the target information of the 3D detection target corresponding to the category with the highest probability distribution is determined through a non-maximum suppression screening strategy.

[0010] Furthermore, in some embodiments, determining the target information of the 3D detection target corresponding to the highest category of the probability distribution by a non-maximum suppression screening strategy includes: Scoring the 3D bounding boxes of each category by category, and sorting them in descending order according to the preset strategy, selecting the optimal 3D bounding box with the highest current score as the benchmark; Calculating the 3D intersection-over-union (IoU) of the optimal 3D bounding box and other 3D bounding boxes, and determining whether the 3D IoU exceeds a preset threshold; If the 3D intersection-and-union ratio exceeds a preset threshold, the 3D bounding box is determined to be a redundant bounding box. After deleting the redundant bounding box, the 3D bounding box with the highest score is selected again from the remaining 3D bounding boxes as a new benchmark, and the 3D bounding box is used as the optimal 3D bounding box. The steps of calculating the 3D intersection-and-union ratios of the optimal 3D bounding box and the other 3D bounding boxes respectively and determining whether the 3D intersection-and-union ratios exceed a preset threshold are continued until the calculation of the other 3D bounding boxes is completed, a final 3D bounding box is obtained, and the final 3D bounding box is used to determine the target information of the 3D detection target.

[0011] According to the 3D target detection method based on a pulse neural network in an embodiment of the present application, the first target feature of the multi-view image in a bird's-eye view and the second target feature of the radar point cloud in a bird's-eye view are obtained, and feature fusion is performed on them to obtain the initial multimodal fusion feature of the multi-view image. The initial fusion feature is enhanced using a preset feature encoder to obtain the final multimodal fusion feature of the multi-view image. The final multimodal fusion feature is parsed using a detection head to obtain the target information of the 3D detected target. As a result, the problems of dependence on hardware computing power and inference time, as well as task detection energy consumption and time delay caused by the detection methods of related technologies are solved. Through an event-driven mechanism and sparse pulse coding, combined with multimodal feature fusion, computing energy consumption and inference delay are significantly reduced while maintaining high detection accuracy.

[0012] A second aspect of the present application provides a computer program product, including a computer program, which is executed to implement the 3D target detection method based on the pulse neural network described in the above embodiment.

[0013] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart of a 3D object detection method based on a spiking neural network according to an embodiment of the present application; Figure 2 This is an overall flow chart of target detection according to one embodiment of the present application; Figure 3 This is a structural diagram of a backbone network based on a pulse neural network according to an embodiment of the present application. DETAILED DESCRIPTION

[0015] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0016] The following describes a 3D target detection method based on a pulse neural network according to an embodiment of the present application with reference to the accompanying drawings. In response to the problems mentioned in the background art above, such as the dependence on hardware computing power and inference time caused by the detection methods of the related technologies, as well as the energy consumption and time delay of task detection, the present application provides a 3D target detection method based on a pulse neural network. In this method, the first target feature and the second target feature of the multi-view image under the bird's-eye view are obtained, and feature fusion is performed on them to obtain the initial multimodal fusion feature of the multi-view image. The initial fusion feature is enhanced by a preset feature encoder to obtain the final multimodal fusion feature of the multi-view image. The final multimodal fusion feature is parsed by the detection head to obtain the target information of the 3D detection target. Thus, the problems of the dependence on hardware computing power and inference time caused by the detection methods of the related technologies, as well as the energy consumption and time delay of task detection, are solved. By combining the event-driven mechanism and sparse pulse coding with multimodal feature fusion, the computing energy consumption and inference delay are significantly reduced while maintaining high detection accuracy.

[0017] Specifically, Figure 1 A schematic flow chart of a 3D target detection method based on a spiking neural network provided in an embodiment of the present application.

[0018] like Figure 1 As shown, the 3D target detection method based on the pulse neural network includes the following steps: In step S101 , a first target feature of a multi-view image in a bird's-eye view and a second target feature of a radar point cloud in a bird's-eye view are acquired.

[0019] Furthermore, in some embodiments, obtaining the first target feature of a multi-perspective image from a bird's-eye view perspective includes: extracting image features of the multi-perspective image based on a 2D backbone network of a preset pulse neural network, and performing dimensionality conversion on the extracted image features to obtain target dimension image features of the multi-perspective image; predicting the discrete depth distribution of each pixel in the target dimension image feature using a preset algorithm to obtain the depth probability corresponding to each pixel; rescaling the target dimension image feature based on the depth probability corresponding to each pixel, and pooling the scaled target dimension image feature to obtain the first target feature of the multi-perspective image from a bird's-eye view perspective.

[0020] Furthermore, in some embodiments, obtaining a second target feature of the radar point cloud from a bird's-eye view perspective includes: obtaining point cloud information of the radar point cloud, and performing data processing on the point cloud information based on a preset voxel size to obtain voxelized point cloud information; extracting point cloud features from the voxelized point cloud information based on a 3D backbone network of a preset pulse neural network, and performing dimensionality conversion on the extracted point cloud features to obtain target dimension point cloud features of the radar point cloud; compressing the target dimension point cloud features based on the target axis and projecting them to the bird's-eye view space to obtain the second target feature of the radar point cloud from a bird's-eye view perspective.

[0021] Among them, the preset spiking neural network, the preset algorithm and the preset voxel size can all be selected by those skilled in the art according to actual target detection requirements and are not specifically limited here.

[0022] Specifically, since the target detection method of the related technology relies on artificial neural networks and aims to improve accuracy by increasing the network scale, it will result in high computing power requirements, high energy consumption and long inference delay. Therefore, in order to solve the above-mentioned problems, the embodiment of the present application uses a pulse neural network to process multi-view images and point cloud information, and uses the event-driven mechanism of pulse neurons to achieve efficient feature extraction, and obtains multimodal BEV (Bird's Eye View) features with structural and semantic information through feature fusion. Finally, the features are decoded to obtain the category and relative position of the detected object, realize target recognition and positioning in the perception environment, and efficiently complete the 3D target detection task.

[0023] Specifically, if Figure 2 As shown, the embodiment of the present application implements the image feature extraction step by the image feature extraction module, wherein the image feature extraction module includes a 2D backbone network module based on a preset pulse neural network and an image to BEV conversion module. The 2D backbone network module is responsible for extracting multi-view image features and performing dimensionality conversion on the extracted image features, that is, mapping high-dimensional data into low-dimensional feature representation to obtain semantically rich target dimensional image features, namely, multi-scale features, wherein, as Figure 3 As shown in (a), the 2D backbone network includes two 2D convolutional layers, two normalization layers, and a spike neuron. The spike neuron is the core part of the preset spike neural network, which is used to convert feature information into a spike signal for propagation in the network. In this application, a leaky integral-trigger model is used as the spike neuron. Because of its good trade-off between biological plausibility and complexity, it is more suitable for constructing the preset spike neural network. Its expression is:

[0024]

[0025] in, is the membrane potential of the i-th neuron in the n+1-th layer at time step t, is the attenuation coefficient (explain what each symbol means), Pulse input The weight of is a step function. When x≥0, ,otherwise is the discharge threshold, X is the input, and V is the membrane position.

[0026] Furthermore, in order to solve the problem that pulse signals cannot be distinguished in back propagation, the embodiment of the present application uses a proxy gradient, which can be expressed as:

[0027] Among them, the introduction It is to ensure that the integral of the gradient is 1, X is the input, V is the membrane position, and in addition, the embodiment of the present application also introduces a residual module to enhance the feature extraction capability of the 2D backbone network to avoid the gradient disappearance or gradient explosion problem in the deep network, while enhancing the feature extraction capability of the deep network.

[0028] Secondly, after obtaining the target dimension image features of the multi-view image, a preset algorithm, such as the LSS (Lift-Splat-Shoot) algorithm, is used to predict the discrete depth distribution of each pixel in the target dimension image features, and obtain the depth probability corresponding to each pixel, so that it only emits pulses when the membrane potential exceeds the threshold, reducing redundant calculations, and rescales the target dimension image features based on the depth probability corresponding to each pixel. Then, the image to BEV conversion module is used to pool the scaled target dimension image features to obtain the first target features of the multi-view image at the bird's-eye view perspective, that is, the image BEV features of the multi-view image at the bird's-eye view perspective, thereby realizing the conversion of the multi-view image from image features to BEV features.

[0029] Furthermore, the point cloud information of the multi-view image is obtained. Since the disordered point cloud cannot be directly input into the neural network, it needs to be converted into structured data. The point cloud information is processed based on the preset voxel size, that is, the discrete point cloud information is converted into a regular voxel representation to obtain voxelized point cloud information. The voxelized point cloud information is then used as input and the voxelized point cloud information is extracted based on the 3D backbone network of the preset spiking neural network. For example, only non-empty voxels are calculated to improve efficiency, such as Figure 3The 3D backbone network adopts a similar structure to the 2D backbone network, the core part is composed of 2 3D sparse convolution layers, 2 normalization layers and 1 pulse neuron, and a residual structure is also used to prevent gradient disappearance and enhance the feature extraction capability of the deep network. At the same time, the extracted point cloud features are converted in dimension to obtain target dimension point cloud features of the multi-view image. Since the target dimension point cloud features need to be fused with the image BEV features in the same space BEV, the target dimension point cloud features need to be compressed based on the target axis (for example, the z-axis) and projected to the bird's eye view space to obtain the second target feature of the multi-view image in the bird's eye view perspective, that is, the point cloud BEV feature of the multi-view image in the bird's eye view perspective.

[0030] In step S102, the first target feature and the second target feature are fused to obtain an initial multi-modal fusion feature of the multi-view image, and a preset feature encoder is used to perform feature enhancement processing on the initial fusion feature to obtain a final multi-modal fusion feature of the multi-view image.

[0031] The preset feature encoder can be selected by a person skilled in the art according to actual target detection requirements, and is not specifically limited here.

[0032] Specifically, in the above feature extraction process, the image BEV feature and the point cloud BEV feature have been converted to the BEV perspective, so that the image BEV feature and the point cloud BEV feature are in the same space. However, due to the inaccuracy of the view converter, the image BEV feature and the point cloud BEV feature still exist spatial misalignment to a certain extent, and simple feature linking will cause feature information confusion. Therefore, an embodiment of the present application designs a convolution-based feature encoder as a feature fusion module, which is responsible for fusing the image BEV feature and the point cloud BEV feature into a multi-modal fusion BEV feature with structure and semantic information.

[0033] Specifically, in the embodiment of the present application, a structure similar to the 2D backbone network is adopted, the image BEV feature and the point cloud BEV feature are spliced as input, the image BEV feature and the point cloud BEV feature are fused to obtain an initial multi-modal fusion feature of the multi-view image, and then a preset feature encoder is used to perform feature enhancement processing on the initial fusion feature, for example, through channel splicing + convolution dimension reduction, to generate a unified multi-modal BEV feature, that is, to obtain a final multi-modal fusion feature of the multi-view image.

[0034] In step S103, the final multi-modal fusion feature is input to the detection head to parse the final multi-modal fusion feature by using the detection head to obtain target information of the 3D detection target.

[0035] Further, in some embodiments, the final multi-modal fusion feature is input into the detection head to parse the final multi-modal fusion feature by the detection head to obtain target information of the 3D detection target, including predicting a 3D bounding box of the 3D detection target and a category of the 3D detection target, and outputting a probability distribution of each category of the 3D detection target; and determining target information of the 3D detection target corresponding to the category with the highest probability distribution based on the probability distribution of each category of the 3D detection target through a non-maximum suppression screening strategy.

[0036] Further, in some embodiments, the target information of the 3D detection target corresponding to the category with the highest probability distribution is determined through the non-maximum suppression screening strategy, including scoring the 3D bounding box of each category by category, and selecting the optimal 3D bounding box with the highest current score as a reference according to a preset descending order sorting strategy; calculating the 3D intersection over union of the optimal 3D bounding box with other 3D bounding boxes, and determining whether the 3D intersection over union exceeds a preset threshold; if the 3D intersection over union exceeds the preset threshold, determining that the 3D bounding box is a redundant bounding box, and after deleting the redundant bounding box, selecting the 3D bounding box with the highest score from the remaining 3D bounding boxes as a new reference, and taking the 3D bounding box as the optimal 3D bounding box, continuing to perform the steps of calculating the 3D intersection over union of the optimal 3D bounding box with other 3D bounding boxes, and determining whether the 3D intersection over union exceeds the preset threshold, until all other 3D bounding boxes are calculated, obtaining the final 3D bounding box, and determining the target information of the 3D detection target based on the final 3D bounding box.

[0037] The preset descending order sorting strategy and the preset threshold can both be selected by a person skilled in the art according to actual target detection requirements, and are not specifically limited herein.

[0038] Specifically, after obtaining the final multi-modal fusion feature of the multi-view image, the final multi-modal fusion feature is input into the detection head, which is responsible for converting the extracted high-level features into specific detection results (such as object category, position, size, etc.), so as to parse the final multi-modal fusion feature by the detection head to obtain target information of the 3D detection target, including position information, category information, and circumscribed rectangular bounding box information of all 3D detection targets, and predict the position of the center of the 3D detection target by the detection head, regress the related information of the 3D detection target, and screen the final detection result through non-maximum suppression.

[0039] Specifically, first, the final multi-modal fusion feature is input into the detection head to predict a 3D bounding box of the 3D detection target and a category of the 3D detection target, such as a vehicle or a pedestrian, then predict the center coordinates of the 3D detection target, and output the probability distribution of each category of the 3D detection target, and then determine the target information of the 3D detection target corresponding to the category with the highest probability distribution based on the probability distribution of each category of the 3D detection target. Secondly, through the non-maximum suppression screening strategy, redundant bounding boxes are filtered, and the most reliable detection results are retained. For example, the scores of 3D bounding boxes of each class can be ranked in descending order according to a preset ranking strategy, the optimal 3D bounding box with the highest score is selected as a reference, the 3D intersection over union of the optimal 3D bounding box and other 3D bounding boxes is calculated, and when the 3D intersection over union exceeds a preset threshold, the 3D bounding box is determined as a redundant bounding box and is deleted. After deleting the redundant bounding box, the 3D bounding box with the highest score is selected again from the remaining 3D bounding boxes as a new reference, and the 3D bounding box is taken as the optimal 3D bounding box. The step of calculating the 3D intersection over union of the optimal 3D bounding box and other 3D bounding boxes and determining whether the 3D intersection over union exceeds the preset threshold is continued until the optimal 3D bounding box of each round is calculated with other 3D bounding boxes, and the final 3D bounding box is obtained. The target information of the 3D detection target is determined based on the final 3D bounding box, wherein the target information can include the position, size information and orientation information of the 3D detection target.

[0040] Therefore, the embodiments of the present application replace the traditional continuous activation neuron calculation with the event-driven mechanism of the spiking neural network, reduce the hardware resource consumption caused by redundant feature extraction, realize the spiking sparse coding based on the dynamic characteristics of biological neurons, trigger the processing of only the key spatio-temporal features in high-dimensional data such as 3D point clouds, and reduce the computational complexity. Finally, the present application significantly reduces the hardware system power consumption and shortens the end-to-end inference delay in the 3D target detection task, while maintaining the feature information sensitivity through the spiking threshold firing value mechanism, breaking through the collaborative optimization bottleneck of energy efficiency and real-time performance in the detection task.

[0041] In summary, based on the specific discussion of the above embodiments, the embodiments of the present application can achieve the following beneficial effects: (1) A 2D and 3D backbone network based on a spiking neural network is proposed, which considers the design of the network structure as a whole, combines the characteristics of spiking neurons, reduces the additional resource consumption caused by redundant information, realizes efficient information transmission through the sparse representation of membrane potential dynamics, can efficiently extract features from image and point cloud information, and realizes a low-power and low-delay perception solution; (2) A method for multi-modal fusion in the BEV perspective is proposed, including BEV feature extraction of image information and point cloud information, designing an effective feature fusion module for feature alignment, realizing feature fusion of the two in the BEV perspective, obtaining multi-modal features with structure and semantic information, and improving detection accuracy; (3) The application takes the pulse neural network as the basic backbone network for feature extraction, the unique event-driven mechanism replaces the traditional continuous activation neuron calculation, reduces the hardware resource consumption caused by redundant feature extraction, and overcomes the problem of hardware computing power dependence caused by large energy consumption in the existing target detection method; (4) In the implementation of the target detection reasoning process, the application transmits the features in the form of pulse sparse coding, only triggers the key spatiotemporal features in high-dimensional data such as 3D point cloud, reduces the computational complexity, and greatly alleviates the reasoning delay problem.

[0042] According to the 3D target detection method based on the pulse neural network, the first target feature of the multi-view image under the bird's eye view and the second target feature of the radar point cloud under the bird's eye view are obtained, and the features are fused to obtain the initial multi-modal fusion feature of the multi-view image. The preset feature encoder is used for feature enhancement processing on the initial fusion feature to obtain the final multi-modal fusion feature of the multi-view image. The detection head is used to analyze the final multi-modal fusion feature to obtain the target information of the 3D detection target. Therefore, the dependence on hardware computing power and reasoning time caused by the detection method of the related technology, as well as the task detection energy consumption and time delay problems are solved. Through the event-driven mechanism and sparse pulse coding, combined with multi-modal feature fusion, the calculation energy consumption and reasoning delay are significantly reduced, while the high detection accuracy is maintained.

[0043] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, the skilled in the art can combine and combine the different embodiments or examples described in the present application and the features of the different embodiments or examples without contradiction.

[0044] In addition, the terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically limited.

[0045] Any processes or methods described in the flowcharts or otherwise described herein can be understood as representing code modules, segments, or portions of code which include one or more executable instructions for implementing specific logic functions (or steps) of the application, and alternate implementations are possible. The various steps or functions described in the flowcharts or otherwise described herein can be implemented as program instructions (i.e., as one or more modules of computer program code) in any of various forms, including but not limited to program components, applets, threads of execution, procedures, functions, etc., whether implemented in hardware, software, firmware, or their combination. It will be understood that the various steps or functions described in the flowcharts or otherwise described herein can be implemented by any number of hardware devices or software modules, including but not limited to processors, hardwired circuitry, software modules, firmware modules, etc. Thus, the steps or functions of the flowcharts or other descriptions herein can be embodied in any of a wide variety of forms, including but not limited to software code, hardware logic, software code and hardware logic, firmware code, etc. The various forms of the application can be implemented in any of a variety of ways, including as software code, firmware code, hardware logic, etc. The software code, firmware code, and / or hardware logic can be stored, for example, on a computer-readable medium, which can be any medium readable by a computer or other instruction execution system, including but not limited to any of various types of memory devices, including volatile memory, non-volatile memory, etc. The software code, firmware code, and / or hardware logic can be executed by an instruction execution system, which can be any system that can fetch, decode, and execute instructions, including but not limited to a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from a computer-readable medium.

[0046] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a list of executable instructions for implementing the logic function, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any medium that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be any one of many physical media, including but not limited to electronic, magnetic, optical, electromagnetic, infrared, and semiconductor systems (or apparatuses) including a single processor or multiple processors. The processor may, for example, be constituted by one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any other devices, or combinations thereof, that can fetch, interpret, and / or execute software instructions.

[0047] It should be understood that aspects of the application can be implemented in hardware, software, firmware, or combinations thereof. In the above embodiments, the steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. As such, in hardware implementations, any of the following technologies, or combinations thereof, can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and / or the like.

[0048] Those skilled in the art of the present technology can understand that all or part of the steps carried out by the above-mentioned embodiment method can be completed by programs instructing related hardware, and the programs can be stored in a computer readable storage medium. When the program is executed, it includes one of the steps of the method embodiment or a combination thereof.

[0049] In addition, each functional unit in each embodiment of the present application can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The integrated module can be realized in the form of hardware or in the form of a software functional module. The integrated module, if realized in the form of a software functional module and sold or used as an independent product, can also be stored in a computer readable storage medium.

[0050] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it should be understood that the above-mentioned embodiments are exemplary and cannot be understood as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above-mentioned embodiments within the scope of the present application.

Claims

1. A 3D object detection method based on a spiking neural network, characterized in that: The following steps are involved: Acquire a first target feature of the multi-view image in a bird's-eye view and a second target feature of the radar point cloud in a bird's-eye view; Performing feature fusion on the first target feature and the second target feature to obtain an initial multimodal fusion feature of the multi-view image, and performing feature enhancement processing on the initial fusion feature using a preset feature encoder to obtain a final multimodal fusion feature of the multi-view image; The final multimodal fusion feature is input to a detection head, so as to utilize the detection head to analyze the final multimodal fusion feature and obtain target information of a 3D detection target.

2. The method according to claim 1, characterized in that The obtaining of the first target feature of the multi-view image in a bird's-eye view includes: Extracting image features from the multi-view images using a 2D backbone network based on a preset spiking neural network, and performing dimensionality conversion on the extracted image features to obtain target dimensional image features of the multi-view images; Predicting the discrete depth distribution of each pixel in the target dimension image feature using a preset algorithm to obtain the depth probability corresponding to each pixel; The target dimension image feature is rescaled based on the depth probability corresponding to each pixel point, and the scaled target dimension image feature is pooled to obtain the first target feature of the multi-view image under the bird's-eye view perspective.

3. The method according to claim 1, characterized in that The obtaining of the second target feature of the radar point cloud from a bird's-eye view perspective includes: Acquiring point cloud information of the radar point cloud, and performing data processing on the point cloud information based on a preset voxel size to obtain voxelized point cloud information; Extracting point cloud features from the voxelized point cloud information using a 3D backbone network based on a preset spiking neural network, and performing dimension conversion on the extracted point cloud features to obtain target dimension point cloud features of the radar point cloud; The target dimension point cloud features are compressed based on the target axis and projected into the bird's-eye view space to obtain the second target features of the radar point cloud under the bird's-eye view perspective.

4. The method according to claim 1, wherein Inputting the final multimodal fusion feature into a detection head, and using the detection head to parse the final multimodal fusion feature to obtain target information of a 3D detection target, includes: Predicting a 3D bounding box of the 3D detection target and a category of the 3D detection target, and outputting a probability distribution of each category of the 3D detection target; Based on the probability distribution of each category of the 3D detection target, the target information of the 3D detection target corresponding to the category with the highest probability distribution is determined through a non-maximum suppression screening strategy.

5. The method according to claim 4, characterized in that The determining, by the non-maximum suppression screening strategy, target information of the 3D detection target corresponding to the highest category of the probability distribution includes: Scoring the 3D bounding boxes of each category by category, and sorting them in descending order according to the preset strategy, selecting the optimal 3D bounding box with the highest current score as the benchmark; Calculating the 3D intersection-over-union (IoU) of the optimal 3D bounding box and other 3D bounding boxes, and determining whether the 3D IoU exceeds a preset threshold; If the 3D intersection-and-union ratio exceeds a preset threshold, the 3D bounding box is determined to be a redundant bounding box. After deleting the redundant bounding box, the 3D bounding box with the highest score is selected again from the remaining 3D bounding boxes as a new benchmark, and the 3D bounding box is used as the optimal 3D bounding box. The steps of calculating the 3D intersection-and-union ratios of the optimal 3D bounding box and the other 3D bounding boxes respectively and determining whether the 3D intersection-and-union ratios exceed a preset threshold are continued until the calculation of the other 3D bounding boxes is completed, a final 3D bounding box is obtained, and the final 3D bounding box is used to determine the target information of the 3D detection target.

Citation Information

Cited By

  • Few-sample point cloud classification method based on multi-modal pulse fusion neurons

    CN121330386A