3D target detection method and system for jointly updating scene and query point

Through the joint update method of scene and query point, the features are updated dynamically synchronously, which solves the problems of query point feature fixity and background point interference in three-dimensional object detection, and achieves more efficient object detection and recognition.

CN120495630APending Publication Date: 2025-08-15DEEP SPACE EXPLORATION LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510589355.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing three-dimensional indoor object detection method has performance bottlenecks in querying point feature fixity and background point interference issues, which limits the detection capability and accuracy of the model.

Method used

The method of joint update of scene and query point is adopted to generate initial features through feature extractors, and feature updates are performed using the decoder of the interactive state space model. Combined with Hilber feature cloud serialization, bidirectional scanning, state attention and gated feedforward network, query point and scene point features are dynamically synchronized to enhance spatial relationship modeling and suppress background point interference.

Benefits of technology

It improves the accuracy and efficiency of three-dimensional object detection, can more accurately detect object instances and identify object categories, and improves detection performance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495630A_ABST
    Figure CN120495630A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D target detection method and system for jointly updating a scene and a query point, and relates to the technical field of three-dimensional computer vision. Point cloud data are received, the point cloud data are input into a pre-established feature extractor, point cloud features are coded, and initial features of the scene point and the query point are obtained; inputting the initial features of the scene points and the query points into a decoder of a pre-established interactive state space model, and outputting to obtain updated query point features and scene point features; and inputting the updated query point features and scene point features into a detection head, and outputting a 3D target detection result of the object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional computer vision technology, and in particular to a 3D target detection method and system for jointly updating scenes and query points. Background Art

[0002] With the widespread adoption of LiDAR and depth cameras, 3D point cloud data has become increasingly accessible. This data provides rich geometric information for 3D scene understanding in fields such as autonomous driving, robotics, and augmented reality. As a fundamental task in 3D scene understanding, 3D indoor object detection has attracted considerable attention from both academia and industry. Current 3D indoor object detection methods can be broadly categorized into three categories: voting-based methods, extension-based methods, and DETR-based methods. Voting-based methods utilize a voting mechanism to shift surface points to the center of an object and generate query points through clustering. While these methods have achieved significant results in 3D object detection, their voting mechanism operates in a class-agnostic manner, which can easily lead to grouping adjacent shifted points belonging to different classes, limiting the model's detection capabilities. Extension-based methods employ a generative sparse decoder to generate high-quality proposals based on voxel features with the same semantic predictions on the object surface. Compared to voting-based methods, these methods consider the consistency of intra-voxel semantics and achieve superior performance. However, their reliance on a proposal generation module and the extensive manual thresholding required limit the model's versatility and scalability. DETR-based methods have demonstrated strong performance in 3D object detection. This type of method selects a small number of points from the point cloud or voxel as the initial query, and gradually refines these query points using scene point features. The design of the query refinement module preserves the original geometric structure of the 3D point cloud and significantly improves the accuracy of indoor object detection. However, DETR-based methods still face a key performance bottleneck: although their multi-layer transformer decoders can update query point features layer by layer, each layer uses the same fixed scene point features. This feature fixity results in only marginal benefits for improvements in subsequent decoder layers, which limits further improvements in model performance. Summary of the Invention

[0003] In order to address the deficiencies mentioned in the above background technology, the object of the present invention is to provide a 3D object detection method and system for jointly updating scenes and query points.

[0004] In a first aspect, the purpose of the present invention can be achieved by the following technical solution: a 3D object detection method for jointly updating scenes and query points, the method comprising the following steps:

[0005] Receive point cloud data, input the point cloud data into a pre-established feature extractor, encode the point cloud features, and obtain the initial features of the scene points and query points;

[0006] Input the initial features of the scene point and the query point into the decoder of the pre-established interactive state space model, and output the updated query point features and scene point features;

[0007] Based on the updated query point features and scene point features, the features are input into the detection head and the 3D target detection results of the object are output.

[0008] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: inputting the point cloud data into a pre-established feature extractor to encode point cloud features:

[0009] The point cloud data is extracted through the feature extraction network PointNet++ and combined with the query point sampling module to generate scene points P seed and query point P prop spatial location and characteristics.

[0010] In combination with the first aspect, in some implementations of the first aspect, the method further includes: obtaining a scene point P by sampling the point cloud data at the farthest point seed , the scene point uses k-nearest neighbor and multi-layer perceptron to aggregate the features of surrounding points, which is repeated twice as the input point cloud of the next stage;

[0011] Use the last output scene point P seed The features of are used to predict the probability of the foreground object point, and the scene point P with the largest foreground object probability is selected. seed , as the query point P prop .

[0012] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: the pre-established interactive state-space model describes the evolution of the system state h(t), and predicts the future state h(t) and the system output y(t) based on the system input x(t), wherein the system is defined as follows:

[0013] h(t)′=Ah(t)+Bx(t)

[0014] y(t)=Ch(t)′

[0015] Among them, A represents the state transition matrix describing how the system state evolves, B represents the control matrix describing the impact of the system input on the system state, and C represents the observation matrix describing the impact of the system state on the system output.

[0016] In combination with the first aspect, in some implementations of the first aspect, the method further includes: in the pre-established interactive state space model, for each state point Predict the rotated 3D bounding box and calculate the relative offset ΔP between the scene point and the 8 vertices of the bounding boxi , use a multilayer perceptron to map positional relationships to parameters in the state space model:

[0017] Δ t =Linear(S i ),B=Linear(S i ),C=Linear(S i )

[0018] Among them, Linear is a linear layer that uses scene point features to adjust the state space model parameters:

[0019] Δ t =Linear(S i )+Linear t (x),

[0020] B=Linear(S i )+Linear b (x),

[0021] C=Linear(S i )+Linear c (x)

[0022] The background point interference problem is handled by the explicit delay kernel below:

[0023]

[0024] in Status point The bounding sphere radius of the predicted bounding box, P x is the location of the scene point, α is a learnable parameter, and B, C, state points select corresponding scene points for update, and scene points simultaneously obtain surrounding structural information from state points.

[0025] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: an operation flow of the decoder of the pre-established interactive state space model includes:

[0026] Hilbert point cloud serialization strategy, bidirectional scanning, state attention, and gated feedforward networks.

[0027] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: the Hilbert point cloud serialization strategy is as follows:

[0028] Let P i =(x i ,y i ,z i) represents the position of the i-th scene point, and these points are mapped to one-dimensional space through the Hilbert curve. The priority of each mapping method is:

[0029] index=H xyz (P i )

[0030] Among them H xyz It is a Hilbert curve mapping function based on the priority order of x, y and z axes;

[0031] The bidirectional scanning is as follows:

[0032] The scene points are input into the pre-order and post-order state space models in positive and reverse order respectively. The state points are used as system states for feature update, and the pre-order and post-order outputs are fused through a linear layer to generate the final scene point and state point features.

[0033] The state attention is as follows: the query self-attention mechanism is introduced to model the relationship between objects, and the state point h′ i The features of are weighted updated:

[0034]

[0035] Among them, Q i ,K i ,V i Denote the matrix of query, key and value respectively, d k is the feature dimension, and the final output is h′ i is the updated feature of the state point;

[0036] The gated feedforward network is as follows:

[0037] The calculation of the gated linear unit is as follows:

[0038] g i =σ(W g h i +b g )

[0039] Where σ is the sigmoid activation function, W g and b g is the learning parameter, the output gate value g i To control the activation level of each feature, the final feedforward network output is:

[0040]

[0041] MLP stands for Multi-Layer Perceptron.

[0042] In a second aspect, in order to achieve the above-mentioned object, the present invention discloses a 3D object detection system with joint scene and query point update, comprising:

[0043] The feature extraction module is used to receive point cloud data, input the point cloud data into a pre-established feature extractor, encode the point cloud features, and obtain the initial features of the scene points and query points;

[0044] A feature processing module is used to input the initial features of the scene point and the query point into the decoder of the pre-established interactive state space model, and output updated query point features and scene point features;

[0045] The target detection module is used to input the updated query point features and scene point features into the detection head and output the 3D target detection results of the object.

[0046] In another aspect of the present invention, in order to achieve the above-mentioned purpose, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, it adopts the 3D target detection method of jointly updating the scene and query point as described above.

[0047] In another aspect of the present invention, in order to achieve the above-mentioned purpose, a computer-readable storage medium is disclosed, in which a computer program is stored. When the computer program is loaded and executed by a processor, the 3D target detection method of jointly updating the scene and query points as described above is adopted.

[0048] Beneficial effects of the present invention:

[0049] This invention largely solves the problems of insufficient fusion of query point and scene point features and background point interference in 3D target detection, and ultimately helps the model accurately and efficiently detect object instances and identify object categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0051] Figure 1 It is a schematic flow chart of the method of the present invention;

[0052] Figure 2 Schematic diagram of visual comparison between the present invention and the GF3D method on the ScanNet V2 dataset;

[0053] Figure 3 It is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0055] Example 1:

[0056] like Figure 1 As shown, a 3D object detection method for jointly updating scenes and query points includes the following steps:

[0057] S101: receiving point cloud data, inputting the point cloud data into a pre-established feature extractor, encoding point cloud features, and obtaining initial features of scene points and query points;

[0058] Specifically, the input point cloud is passed through the feature extraction network PointNet++ and combined with the query point sampling module to generate the scene point P seed and query point P prop spatial location and characteristics;

[0059] First, the input point cloud is sampled by the farthest point to obtain the scene point P seed , the scene point uses k nearest neighbor and multi-layer perceptron to aggregate the features of surrounding points, which is repeated twice as the input point cloud of the next stage;

[0060] Use the last output scene point P seed The features of the scene point P are used to predict the probability of it being a foreground object point. seed , which we call the query point P prop .

[0061] S102: Inputting the initial features of the scene point and the query point into a decoder of a pre-established interactive state space model, and outputting updated query point features and scene point features;

[0062] The pre-established interactive state space model is used to describe the evolution of the system state h(t) and predict the future state h(t) and system output y(t) based on the system input x(t). The system is defined as follows:

[0063] h(t)′=Ah(t)+Bx(t)

[0064] y(t)=Ch(t)′

[0065] Among them, A represents the state transition matrix describing how the system state evolves, B represents the control matrix describing the impact of the system input on the system state, and C represents the observation matrix describing the impact of the system state on the system output.

[0066] In 3D object detection tasks, query point features are modeled as system states, and scene point features serve as system inputs at different time steps, thereby simultaneously obtaining updated query point features (system states) and updated scene point features (system outputs). To address the complex state update problem of 3D point clouds, this paper designs a spatial correlation module to encode the spatial relationship between scene points x and the initial state point h0 to generate the parameters of the state-space model.

[0067] Specifically, for each state point First, a rotated 3D bounding box is predicted, and then the relative offset ΔP between the scene point and the 8 vertices of the bounding box is calculated. i Finally, a multilayer perceptron is used to map these positional relationships to parameters in the state space model:

[0068] Δ t =Linear(S i ),B=Linear(S i ),C=Linear(S i )

[0069] Among them, Linear is a linear layer. To suppress background point interference, scene point features are further used to adjust the state space model parameters. In addition, an explicit delay kernel is designed to deal with the background point interference problem:

[0070]

[0071] in Status point The bounding sphere radius of the predicted bounding box, P x is the location of the scene point, and α is a learnable parameter. The state point can select the appropriate scene point to update, and the scene point can obtain the surrounding structure information from the state point at the same time.

[0072] The decoder of the pre-established interactive state-space model is designed with four modules for feature modeling of scene points and query points: Hilbert point cloud serialization strategy, bidirectional scanning module, state attention module, and gated feedforward network. The specific implementation steps are as follows:

[0073] Hilbert Point Cloud Serialization Strategy: Point cloud data is disordered and irregular. To model scene points as system inputs for a state-space model, this paper designs a serialization strategy based on the Hilbert curve. Six different serialization methods are generated based on the priority order of the x, y, and z axes, providing a multi-view observation of scene points. A different serialization strategy is applied at each decoder layer to ensure that the decoder fully captures the characteristics of scene points.

[0074] Let P i =(x i ,y i ,z i ) represents the position of the i-th scene point, and these points are mapped to one-dimensional space through the Hilbert curve. The priority of each mapping method is:

[0075] index=H xyz (P i )

[0076] Among them H xyz It is a Hilbert curve mapping function based on the priority order of the x, y, and z axes.

[0077] State attention module: In the 3D object detection task, there is a certain correlation between objects, such as tables and chairs often appear in adjacent areas. This paper introduces the query self-attention mechanism to model the correlation between objects, and designs the inter-state attention module to solve the problem of lack of state point interaction modeling in the state space model. Specifically, the state point h i The features of can be updated weightedly:

[0078]

[0079] Among them, Q i ,K i ,V i Denote the matrix of query, key and value respectively, d k is the feature dimension, and the final output is h′ i is the updated feature of the state point. This module significantly enhances the feature extraction capability of the state point, especially when dealing with objects with blurred boundaries or difficult to distinguish from the background, thus improving the detection performance.

[0080] Bidirectional Scanning Module: To enhance feature interaction between scene points, a bidirectional scanning module was designed. Specifically, scene points are input into the pre-order and post-order state space models in forward and reverse order, respectively. The state points are used as system states for feature updates, and the pre-order and post-order outputs are fused via a linear layer to generate the final scene point and state point features. Furthermore, deep convolution is introduced into the module to enhance local feature extraction.

[0081] Gated Feedforward Network: To further enhance feature modeling capabilities, a gated feedforward network containing gated linear units was designed. In each layer of the feedforward network, a gating mechanism was introduced to dynamically select activation paths. Specifically, the gated linear unit is calculated as follows:

[0082] g i =σ(W g h i +b g )

[0083] Where σ is the sigmoid activation function, W g and b g is the learning parameter, the output gate value g i Used to control the activation level of each feature. The final feedforward network output is:

[0084]

[0085] MLP stands for Multi-Layer Perceptron. This module improves the ability to capture complex patterns, thereby optimizing the overall detection performance.

[0086] S103: Based on the updated query point features and scene point features, the features are input into the detection head and a 3D target detection result of the object is obtained as an output.

[0087] Specifically, the present invention can be widely applied to systems in fields such as autonomous driving, robotic grasping, and augmented reality, accurately locating and identifying objects in point cloud scenes. In practice, the system can be installed as software on front-end devices, robots, and autonomous vehicles to provide real-time object detection. It can also be installed on back-end servers to provide large-scale object localization and recognition results for 3D point cloud scenes. Beneficial Effects: This patent proposes a 3D object detection method, system, and device that jointly updates scene and query points based on a state-space model. To address the issue of insufficient updating of query point and scene point features in traditional 3D object detection methods, a decoder based on a state-space model is designed. By dynamically and synchronously updating query and scene point features, object boundaries and background information can be more accurately captured. Furthermore, a spatial correlation module is designed to enhance the spatial relationship modeling between state points and scene points, thereby addressing background interference and improving object detection accuracy. This patented design significantly addresses issues such as insufficient fusion of query and scene point features and background interference in 3D object detection, ultimately enabling the model to accurately and efficiently detect object instances and identify object categories.

[0088] This invention can be widely applied to systems in areas such as autonomous driving, robotic grasping, and augmented reality, accurately locating and identifying objects in point cloud scenes. In practice, it can be installed as software on front-end devices, robots, and autonomous vehicles to provide real-time object detection; it can also be installed on back-end servers to provide large-scale object positioning and recognition results for 3D point cloud scenes.

[0089] Table 1: Comparison of experimental results on ScanNetV2 and SUNRGB-D datasets

[0090]

[0091] As shown in Table 1, we compared our method with the best performing method GF3D on the ScanNetV2 and SUNRGB-D datasets. The results show that our method achieves the best performance in both mAP@50 and mAP@25, verifying the effectiveness of our method.

[0092] exist Figure 2 In

[15] , we compare the detection results of the GF3D method and the method proposed in this patent on the ScanNetV2 dataset. The first column shows the ground truth labels, which annotate the actual object categories and their bounding boxes in the scene. The second column shows the detection results of the GF3D method, and the third column shows the detection results of the method proposed in this patent. As can be seen from the visualization, the method proposed in this patent is closer to the ground truth labels in object localization and bounding box prediction, and has a significant advantage in complex scenes and objects with blurred boundaries.

[0093] Example 2: In order to achieve the above purpose, Figure 3 As shown, the present invention discloses a 3D object detection system for jointly updating scenes and query points, comprising:

[0094] The feature extraction module 11 is used to receive point cloud data, input the point cloud data into a pre-established feature extractor, encode the point cloud features, and obtain the initial features of the scene points and query points;

[0095] A feature processing module 12 is used to input the initial features of the scene point and the query point into a decoder of a pre-established interactive state space model, and output updated query point features and scene point features;

[0096] The target detection module 13 is configured to input the updated query point features and scene point features into the detection head and output a 3D target detection result of the object.

[0097] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.

[0098] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium having a computer program stored thereon, which executes the above method when executed by a processor. The storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.

[0099] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present disclosure. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0100] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present disclosure. Various changes and improvements may be made to the present disclosure without departing from the spirit and scope of the present disclosure, and such changes and improvements shall fall within the scope of the present disclosure.

Claims

1. A 3D object detection method with joint scene and query point update, characterized in that: The method comprises the following steps: Receive point cloud data, input the point cloud data into a pre-established feature extractor, encode the point cloud features, and obtain the initial features of the scene points and query points; Input the initial features of the scene point and the query point into the decoder of the pre-established interactive state space model, and output the updated query point features and scene point features; Based on the updated query point features and scene point features, the features are input into the detection head and the 3D target detection results of the object are output.

2. The 3D object detection method for jointly updating scenes and query points according to claim 1, characterized in that: The process of inputting point cloud data into a pre-established feature extractor to encode point cloud features: The point cloud data is extracted through the feature extraction network PointNet++ and combined with the query point sampling module to generate scene points P seed and query point P prop spatial location and characteristics.

3. The 3D object detection method for jointly updating scenes and query points according to claim 2, characterized in that: The point cloud data is sampled by the farthest point to obtain the scene point P seed , the scene point uses k-nearest neighbor and multi-layer perceptron to aggregate the features of surrounding points, which is repeated twice as the input point cloud of the next stage; Use the last output scene point P seed The features of are used to predict the probability of the foreground object point, and the scene point P with the largest foreground object probability is selected. seed , as the query point P prop .

4. The 3D object detection method for jointly updating scenes and query points according to claim 1, characterized in that: The pre-established interactive state space model describes the evolution of the system state h(t) and predicts the future state h(t) and system output y(t) based on the system input x(t). The system is defined as follows: h(t)′=A h(t)+Bx(t) y(t)=Ch(t)′ Among them, A represents the state transition matrix describing how the system state evolves, B represents the control matrix describing the impact of the system input on the system state, and C represents the observation matrix describing the impact of the system state on the system output.

5. The 3D object detection method with joint scene and query point update according to claim 4, characterized in that: In the pre-established interactive state space model, for each state point Predict the rotated 3D bounding box and calculate the relative offset ΔP between the scene point and the 8 vertices of the bounding box i , use a multilayer perceptron to map positional relationships to parameters in the state space model: Δ t =Linear(S i ),B=Linear(S i ),C=Linear(S i ) Among them, Linear is a linear layer that uses scene point features to adjust the state space model parameters: Δ t =Linear(S i )+Linear t (x), B=Linear(S i )+Linear b (x), C=Linear(S i )+Linear c (x) The background point interference problem is handled by the explicit delay kernel below: in Status point The bounding sphere radius of the predicted bounding box, P x is the location of the scene point, α is a learnable parameter, and B, C, state points select corresponding scene points for update, and scene points simultaneously obtain surrounding structural information from state points.

6. The 3D object detection method for jointly updating scenes and query points according to claim 1, characterized in that: The operation flow of the decoder of the pre-established interactive state space model includes: Hilbert point cloud serialization strategy, bidirectional scanning, state attention, and gated feedforward networks.

7. The 3D object detection method for jointly updating scenes and query points according to claim 6, characterized in that: The Hilbert point cloud serialization strategy is as follows: Let P i =(x i ,y i ,z i ) represents the position of the i-th scene point, and these points are mapped to one-dimensional space through the Hilbert curve. The priority of each mapping method is: index=H xyz (P i ) Among them H xyz It is a Hilbert curve mapping function based on the priority order of x, y and z axes; The bidirectional scanning is as follows: The scene points are input into the pre-order and post-order state space models in positive and reverse order respectively. The state points are used as system states for feature update, and the pre-order and post-order outputs are fused through a linear layer to generate the final scene point and state point features. The state attention is as follows: the query self-attention mechanism is introduced to model the relationship between objects, and the state point h′ i The features of are weighted updated: Among them, Q i ,K i ,V i Denote the matrix of query, key and value respectively, d k is the feature dimension, and the final output is h′ i is the updated feature of the state point; The gated feedforward network is as follows: The calculation of the gated linear unit is as follows: g i =σ(W g h i +b g ) Where σ is the sigmoid activation function, W g and b g is the learning parameter, the output gate value g i To control the activation level of each feature, the final feedforward network output is: MLP stands for Multi-Layer Perceptron.

8. A 3D object detection system with joint scene and query point updates, characterized in that: include: The feature extraction module is used to receive point cloud data, input the point cloud data into a pre-established feature extractor, encode the point cloud features, and obtain the initial features of the scene points and query points; A feature processing module is used to input the initial features of the scene point and the query point into the decoder of the pre-established interactive state space model, and output updated query point features and scene point features; The target detection module is used to input the updated query point features and scene point features into the detection head and output the 3D target detection results of the object.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, the 3D object detection method for jointly updating scenes and query points according to any one of claims 1 to 7 is adopted.

10. A computer-readable storage medium storing a computer program, wherein: When the computer program is loaded and executed by a processor, the 3D object detection method for jointly updating scenes and query points according to any one of claims 1 to 7 is adopted.