A three-dimensional target detection method based on spatio-temporal adaptive representation regulation

By introducing spatiotemporal adaptive representation control in 3D target detection, the geometric features of radar point cloud data are filtered and the reliability of historical queries is evaluated. This solves the problem of information accumulation in existing methods and improves detection accuracy and stability, especially the consistency of bounding boxes and spatial perception capabilities in dynamic scenes.

CN122493065APending Publication Date: 2026-07-31BEIJING UNIV OF CHEM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF CHEM TECH
Filing Date
2026-05-13
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing temporal multimodal 3D target detection methods lack quality screening of the geometric marker structure of radar point cloud data in the spatial dimension, and lack dynamic evaluation and selective propagation of the reliability of historical queries in the temporal dimension, resulting in the accumulation of unreliable information and affecting the stability and robustness of the detection system.

Method used

By employing a spatiotemporally adaptive representation control method, features are extracted from images and point cloud backbone networks, voxel features are selected and query sets are generated, and then processed using a cross-modal interactive attention model. This achieves adaptive reliability control of the multimodal query set, suppresses the propagation of unreliable information, and improves detection stability.

Benefits of technology

It improves the overall accuracy and stability of multimodal 3D target detection, enhances the cross-modal fusion quality of point cloud query and image query, suppresses error accumulation in time-series propagation, and improves long-term time-series modeling capabilities and spatial perception capabilities of medium- and long-distance targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493065A_ABST
    Figure CN122493065A_ABST
Patent Text Reader

Abstract

This application provides a 3D target detection method based on spatiotemporal adaptive representation control, comprising: extracting features from multi-view images using an image backbone network to obtain semantic features; processing the semantic features using an image query generator to obtain an image query set; extracting features from point cloud data using a point cloud backbone network to obtain multiple voxel features; processing the multiple voxel features using a point cloud query generator to obtain a first point cloud query set; determining a second point cloud query set based on the filtered voxel features and the first point cloud query set; combining the image query set and the second point cloud query set to obtain a multimodal query set at the current time; and processing the multimodal query set at the current time and the historical multimodal query set using a cross-modal interactive attention model to obtain the 3D target detection result at the current time. This application improves the overall accuracy and stability of multimodal 3D target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of three-dimensional target detection technology, and in particular to a three-dimensional target detection method based on spatiotemporal adaptive representation control. Background Technology

[0002] Existing temporal multimodal 3D target detection methods (especially target-level methods) tend to focus on "information augmentation," which improves performance through more complex attention mechanisms and richer feature representations. However, these methods generally neglect another equally important dimension: information regulation. Specifically, existing methods lack proactive management mechanisms for the reliability of representations during propagation: in the spatial dimension, they lack structural quality screening of the geometric markers (voxel features) of radar point cloud data; in the temporal dimension, they lack dynamic evaluation and selective propagation of the reliability of historical queries. This "emphasis on augmentation, neglect of regulation" design tendency allows unreliable information to accumulate continuously in cross-modal interactions and cross-frame propagation, ultimately limiting the stability and robustness of the detection system.

[0003] Therefore, there is an urgent need for a technical solution that can adaptively regulate the reliability of object representation during propagation from both spatial and temporal perspectives. Summary of the Invention

[0004] In view of this, this application provides a three-dimensional target detection method based on spatiotemporal adaptive representation control to solve the above-mentioned technical problems.

[0005] In a first aspect, embodiments of this application provide a three-dimensional target detection method based on spatiotemporal adaptive representation control, comprising: Acquire multi-view images and point cloud data at the current moment; Image backbone network is used to extract features from multi-view images to obtain semantic features; image query generator is used to process the semantic features to obtain image query set; Feature extraction is performed on point cloud data using a point cloud backbone network to obtain multiple voxel features; the point cloud query generator is then used to process the multiple voxel features to obtain the first point cloud query set. Based on the filtered voxel features and the first point cloud query set, the second point cloud query set is determined. The image query set and the second point cloud query set are combined to obtain the multimodal query set at the current moment; The current multimodal query set and the historical multimodal query set are processed using a cross-modal interactive attention model to obtain the current three-dimensional target detection result. The historical multimodal query set includes the multimodal query sets of consecutive frames before the current time.

[0006] Secondly, embodiments of this application provide a three-dimensional target detection device based on spatiotemporal adaptive representation control, comprising: The acquisition unit is used to acquire multi-view images and point cloud data at the current moment; The image processing unit is used to extract features from multi-view images using an image backbone network to obtain semantic features; and to process the semantic features using an image query generator to obtain an image query set. The point cloud processing unit is used to extract features from point cloud data using the point cloud backbone network to obtain multiple voxel features; and to process the multiple voxel features using the point cloud query generator to obtain the first point cloud query set. The determination unit is used to determine the second point cloud query set based on the filtered voxel features and the first point cloud query set; The combination unit is used to combine the image query set and the second point cloud query set to obtain the multimodal query set at the current moment; The processing unit is used to process the current multimodal query set and the historical multimodal query set using a cross-modal interactive attention model to obtain the current three-dimensional target detection result. The historical multimodal query set includes the multimodal query sets of multiple consecutive frames before the current time.

[0007] Thirdly, embodiments of this application provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method of embodiments of this application.

[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the methods of embodiments of this application.

[0009] Fifthly, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the method of embodiments of this application.

[0010] This application improves the overall accuracy and stability of multimodal 3D target detection. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0012] Figure 1 A flowchart of a three-dimensional target detection method based on spatiotemporal adaptive representation control provided in an embodiment of this application; Figure 2 A functional structure diagram of a three-dimensional target detection device based on spatiotemporal adaptive representation control provided in an embodiment of this application; Figure 3 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0014] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0015] First, a brief introduction to the design concept of the embodiments of this application will be given.

[0016] In the field of 3D perception for autonomous driving, temporal multimodal 3D object detection is one of the core technologies. This type of method integrates onboard cameras and LiDAR, utilizing image semantic information and point cloud geometry, and combining temporal context between consecutive frames to improve detection stability and accuracy. Based on different fusion granularities, existing temporal multimodal 3D detection methods can be mainly divided into the following two paradigms.

[0017] Feature-level fusion methods, typically represented by the BEVFusion series, work by unifying data from different modalities into a single feature space (e.g., the BEV space in a bird's-eye view) for fusion, followed by the introduction of temporal information through multi-frame feature aggregation. However, these methods suffer from the following drawbacks: feature alignment weakens the modality-specific structural information, especially the fine geometric details of point clouds, which are difficult to preserve completely; dense historical feature aggregation makes it difficult to explicitly model object motion, limiting its ability to model the temporal consistency of dynamic targets and easily leading to motion blur or positioning drift.

[0018] Object-level fusion methods: These methods construct explicit object representations (such as object queries) using object instances as basic units, and achieve cross-modal alignment and cross-frame propagation through Transformers. Typical examples include query-based architectures such as MV2DFusion and SparseLIF. The underlying principle is to generate LiDAR target queries and image target queries using point clouds and image backbone networks respectively, achieve fusion through cross-modal attention, and then propagate historical frame queries to the current frame for updates through cross-frame attention or a historical query queue.

[0019] Despite significant progress in detection accuracy, these methods still suffer from two key shortcomings in practical applications.

[0020] At the spatial level: Radar target queries lack geometric structure filtering. In existing object-level methods, radar target queries are typically generated directly from single-frame point cloud features or obtained through sparse sampling. These geometric tokens then directly participate in cross-modal interactions. However, the original point cloud contains a large amount of redundant, noisy, or unstable geometric information (such as ground points and uniform points inside objects). Existing methods do not perform explicit structure filtering and quality assessment on these geometric tokens, resulting in low-quality tokens contaminating cross-modal interactions, weakening the spatial structural integrity of LiDAR queries, and consequently affecting the stability and positioning accuracy of subsequent temporal propagation.

[0021] At the temporal level: Historical query propagation lacks reliability assessment. In temporal modeling, existing object-level methods generally assume smooth object motion, directly integrating target queries from historical frames into the current frame through an attention mechanism. This design implicitly assumes that historical information is always reliable. However, real-world scenarios contain numerous non-smooth movements (such as sharp turns and emergency braking) and observational anomalies (such as occlusion and sensor noise). Existing methods lack explicit assessment of the reliability of historical information and selective forgetting mechanisms, leading to the unconditional propagation and accumulation of unreliable historical representations. As the number of frames increases, the cumulative error effect manifests as tracking loss, trajectory deviation, and detection delays after occlusion recovery.

[0022] To this end, this application provides a 3D target detection method based on spatiotemporal adaptive representation control, comprising: acquiring multi-view images and point cloud data at the current moment; extracting features from the multi-view images using an image backbone network to obtain semantic features; processing the semantic features using an image query generator to obtain an image query set; extracting features from the point cloud data using a point cloud backbone network to obtain multiple voxel features; processing the multiple voxel features using a point cloud query generator to obtain a first point cloud query set; determining a second point cloud query set based on the filtered voxel features and the first point cloud query set; combining the image query set and the second point cloud query set to obtain a multimodal query set at the current moment; and processing the multimodal query set at the current moment and the historical multimodal query set using a cross-modal interactive attention model to obtain the 3D target detection result at the current moment, wherein the historical multimodal query set includes multimodal query sets of consecutive frames prior to the current moment.

[0023] The technical advantages of this application include: 1. Improve the overall accuracy and stability of multimodal 3D target detection. Geometrically perceptual feature selection and structural enhancement are performed on LiDAR point cloud queries to generate structurally complete and semantically focused point cloud representations, improving the matching quality of point cloud queries and image queries in the cross-modal fusion stage; motion consistency constraints are used to suppress the propagation of unreliable historical information, further stabilizing the detection results.

[0024] 2. Effectively suppresses error accumulation during time-series propagation. Short-term motion consistency constraints Long-term velocity drift suppression constraints The synergistic effect of these mechanisms allows for adaptive control of unreliable historical information during cross-frame propagation. Among these, Multidimensional supervision is performed on the changes in velocity direction, velocity amplitude and confidence between adjacent frames. When motion anomalies are detected, the constraint weight of historical information is adaptively reduced in a smooth gating manner, thereby blocking the continuous propagation of unreliable historical information. Then, soft boundary constraints are applied to velocity drift during long-term propagation through category-level velocity priors, and... This creates short- and long-term complementarity. Ablation experiments show that... Introducing it alone can stably improve detection performance; on the basis of providing better geometric characterization, further introducing... This approach achieves optimal overall performance. Compared to existing methods that directly introduce historical representations, this application maintains the stability and consistency of bounding boxes in dynamic scenes (such as sharp turns and occlusions), effectively reducing missed detections.

[0025] 3. Enhance long-term time-series modeling capabilities and improve velocity estimation accuracy.

[0026] By imposing multi-scale constraints on the motion consistency of propagation queries during the training phase, the model maintains selective use of reliable historical information throughout the entire time-series propagation process.

[0027] 4. Improve spatial perception of targets at medium and long distances.

[0028] By using geometric saliency scoring (integrating three orthogonal dimensions: local density, semantic consistency, and local variation) to perform structured screening of point cloud tokens, the detection system pays more attention to the geometric information of target boundaries and structurally stable regions, thereby enhancing its spatial perception capability for sparse or unevenly distributed point cloud scenes.

[0029] After introducing the application scenarios and design concepts of the embodiments of this application, the technical solutions provided by the embodiments of this application will be described below.

[0030] like Figure 1 As shown, this application provides a three-dimensional target detection method based on spatiotemporal adaptive representation control, including: Step 101: Obtain multi-view images and point cloud data at the current moment; Step 102: Use the image backbone network to extract features from multi-view images to obtain semantic features; use a pre-trained image query generator to process the semantic features to obtain an image query set; Step 103: Use the point cloud backbone network to extract features from the point cloud data to obtain multiple voxel features; use the pre-trained point cloud query generator to process the multiple voxel features to obtain the first point cloud query set. Step 104: Based on the filtered voxel features and the first point cloud query set, determine the second point cloud query set; Step 105: Combine the image query set and the second point cloud query set to obtain the multimodal query set at the current moment; Step 106: Use the pre-trained cross-modal interactive attention model to process the current multimodal query set and the historical multimodal query set to obtain the current 3D target detection result. The historical multimodal query set includes the multimodal query sets of multiple consecutive frames before the current time.

[0031] Specifically, the 3D target detection results include: target category, target velocity (acceleration and angular velocity), confidence level, and target bounding box. The target bounding box includes the center point, size, and orientation.

[0032] The three-dimensional target detection method in this embodiment explicitly controls the propagation reliability of lidar target query from both spatial and temporal dimensions, solving the problem of decreased multimodal fusion effect and accumulation of temporal errors caused by insufficient geometric representation quality and lack of evaluation of historical information reliability in existing methods.

[0033] In some embodiments, a second point cloud query set is determined based on the filtered voxel features and the first point cloud query set; including: Multiple voxel features are filtered based on a geometric perception scoring filtering strategy to obtain the filtered voxel features. Combine the second point cloud query sets of multiple consecutive frames before the current moment into a historical point cloud query set; The selected voxel features, the first point cloud query set, and the historical point cloud query set are processed using a point cloud refinement attention model to obtain the second point cloud query set at the current time.

[0034] In some embodiments, multiple voxel features are filtered based on a geometric perception scoring filtering strategy to obtain filtered voxel features; including: Construct the first Nearest neighbor set of individual feature Nearest neighbor set Include Individual characteristics; Calculate the first Local density of individual characteristics : in, It is the numerical stability constant; For the first Three-dimensional coordinates of individual features; For the first Three-dimensional coordinates of individual features; Calculate the first Semantic consistency of individual features : in, For the first Feature embedding vector of individual primitive features; For the first Feature embedding vector of individual primitive features; Calculate the first Local feature variation of individual phenotypic characteristics : Calculate the first Geometric significance score of individual primitive features : in, , and All are learnable weights; All voxel features are sorted in descending order based on geometric significance scores to obtain a voxel feature sequence; the top voxel features are then selected from the voxel feature sequence. Individual voxel features are used as the voxel features after screening. This represents the total number of voxel features selected.

[0035] This embodiment addresses the sparse and unevenly distributed characteristics of point clouds by abandoning global evaluation and constructing an evaluation context using local K-nearest neighbors. For the first time, it combines three orthogonal dimensions—"local density (reflecting compactness)," "semantic consistency (reflecting the internal stability of the target)," and "local feature variation (reflecting the salience of structural boundaries)"—and calculates the geometric salience score using learnable weights. The combination logic of these three dimensions has a clear physical meaning: positive weighting of density and consistency encourages the selection of structural regions with dense point clouds and stable semantic representations; negative weighting of variation suppresses the preferential selection of boundary noise points with unstable features, thereby accurately distinguishing effective structural points of the target from background noise or redundant points.

[0036] In some embodiments, the point cloud refinement attention model includes a self-attention layer, a cross-attention layer, and a feedforward network layer connected in sequence; The selected voxel features, the first point cloud query set, and the historical point cloud query set are processed using a point cloud refinement attention model to obtain the second point cloud query set at the current time, including: The current moment First point cloud query set With historical query set Combine them to obtain the query set. ; Using a self-attention layer to query the first point cloud set and query set Processing is performed to obtain the time-series context query. : Query by time series context For the query, the filtered voxel features The key and value are input to a cross-attention layer for processing, resulting in a query set. : Using feedforward network layers for query sets Processing yields the second cloud query set at the current moment. .

[0037] In some embodiments, the method further includes: Establish a training set, which includes sample frames of multiple consecutive frames. Each sample frame includes: multi-view image samples with labeled target bounding boxes and point cloud samples. The semantic features of the current sample frame are obtained by extracting features from the multi-view images of the current sample frame using an image backbone network; the semantic features of the current sample frame are then processed by an image query generator to obtain an image query set for the current sample frame. The voxel features of the current sample frame are obtained by using the point cloud backbone network to extract features from the point cloud samples of the current sample frame; the voxel features of the current sample frame are processed by the point cloud query generator to obtain the first point cloud query set of the current sample frame. Based on the voxel features after filtering the current sample frame and the first point cloud query set, determine the second point cloud query set; The image query set and the second point cloud query set of the current sample frame are combined to obtain the multimodal query set of the current sample frame; The target prediction box is obtained by processing the multimodal query set of the current sample frame and the multimodal query set of the previous consecutive frames using a cross-modal interactive attention model. Based on the target predicted bounding box and the target ground truth bounding box, determine the classification loss value and the predicted bounding box regression loss value; Based on the velocity vector and classification confidence of each query in the multimodal query set of the current sample frame and the multimodal query set of the previous sample frame, the short-term motion consistency interruption loss value is determined. Based on the velocity of the category to which each query belongs in the multimodal query set of the current sample frame, determine the long-term velocity drift suppression loss value; The parameters of the image query generator, point cloud query generator, point cloud refinement attention model, and cross-modal interactive attention model are updated based on the weighted sum of classification loss, predicted box regression loss, short-term motion consistency interruption loss, and long-term velocity drift suppression loss.

[0038] In some embodiments, a second point cloud query set is determined based on the filtered voxel features of the current sample frame and the first point cloud query set; including: A filtering strategy is determined based on the current number of training iterations to filter multiple voxel features of the current sample frame; the filtering strategy is a confidence-guided filtering strategy, a geometry-aware scoring filtering strategy, or a hybrid strategy; the hybrid strategy includes a confidence-guided filtering strategy and a geometry-aware scoring filtering strategy. Based on the filtering strategy, multiple voxel features of the current sample frame are filtered to obtain the filtered voxel features of the current sample frame. The point cloud refinement attention model is used to process the voxel features after filtering the current sample frame, the first point cloud query set of the current sample frame, and the second point cloud query set of the previous multiple consecutive frames to obtain the second point cloud query set of the current sample frame.

[0039] In this embodiment, instead of a single selection strategy, a gradual transition mechanism between a confidence-guided selection strategy and a geometric perception scoring selection strategy is designed during model training. In the early stages of model training, a confidence-guided selection strategy with strong supervision signals (classification confidence, regression boxes) ensures localization reliability. In the later stages of training, as feature representations stabilize, the strategy gradually transitions to a geometric perception scoring selection strategy that relies on local geometric saliency (density, consistency, variability) to ensure the integrity of structural coverage. Its innovation lies in the fact that no manual strategy switching is required; the transition coefficient formula, driven by the number of training iterations, automatically completes the smooth transition from strong supervision guidance to geometric intrinsic characteristic dominance, which is significantly different from existing fixed-strategy or single-strategy methods.

[0040] In some embodiments, the screening strategy is a confidence-guided screening strategy; Based on a filtering strategy, multiple voxel features of the current sample frame are filtered to obtain the filtered voxel features of the current sample frame; including: Step S1: Obtain the classification confidence and target box center coordinates of multiple targets in the current sample frame output by the point cloud backbone network; Step S2: Sort the multiple targets in descending order based on classification confidence to obtain the target sequence; Step S3: Using the first [number]th ... The origin is used as the center of the target bounding box for each target to define a spatial neighborhood with radius r. ; Step S4: In the spatial neighborhood sampling Individual pixel feature samples, where, when Larger than spatial neighborhood The number of voxel feature samples in the data will then be... Set as spatial neighborhood The number of voxel feature samples in the sample; Step S5: Determine whether the cumulative number of sampled voxel features has reached K, where K is the total number of selected voxel features; if yes, end; otherwise, proceed. Updated to Proceed to step S3.

[0041] The confidence-guided screening strategy in this embodiment directly utilizes the strong supervision signals (classification confidence and regression box center) of the detection head to locate target-related regions. In scenarios with uneven point cloud distribution or significant target scale differences, it prioritizes effective coverage of the geometric structures surrounding high-confidence targets, thereby concentrating the collected geometric markers in target-related structured regions and avoiding the allocation of substantial computational resources to background or noisy areas. Furthermore, since both the classification confidence and regression center coordinates undergo end-to-end supervised training, their localization reliability is high. This makes it suitable for providing stable and reliable structural localization basis for geometric marker screening in the early stages of training when the network representation is not yet stable and the geometric saliency score is not yet reliable.

[0042] In some embodiments, when the filtering strategy is a hybrid strategy, the method further includes: Calculate the current training iteration number Transition coefficient : in, This is the preset number of transition start iterations. This is the length of the transition window; Calculate the number of voxel feature samples selected by executing the geometry-aware scoring screening strategy. : ; in, The total number of voxel features selected; Calculate the number of voxel feature samples selected by executing the confidence-guided screening strategy. : .

[0043] Specifically, in the total number of iterations for a single update of the model's parameters, the confidence-guided screening strategy can be implemented for the first 50% of the iterations, and a hybrid strategy can be implemented for 50% to 70% of the iterations. The number of voxel feature samples selected by the confidence-guided screening strategy and the number of voxel feature samples selected by the geometric perception scoring screening strategy can be determined based on the number of iterations. The sum of the two is the total number of voxel features selected. The geometric perception scoring screening strategy is implemented for the last 30% of the iterations.

[0044] In this embodiment, the confidence-guided selection strategy accounts for a large proportion in the early stage of training to ensure convergence stability, while the proportion of the geometric perception scoring selection strategy increases linearly with the number of iterations to enhance geometric perception capability.

[0045] In some embodiments, a short-term motion consistency disruption loss value is determined based on the velocity vector and classification confidence of each query in the multimodal query set of the current sample frame and the multimodal query set of the previous sample frame; including: Get the first multimodal query set of the current sample frame The query is in the same frame as the previous one. velocity vector and classification confidence Current sample frame velocity vector and classification confidence , The sequence number of the current sample frame; Calculate the first The query is in the same frame as the previous one. velocity amplitude and the current sample frame velocity amplitude : Calculate the first The query is in the same frame as the previous one. unit direction vector and the current sample frame unit direction vector : in, It is the numerical stability constant; Calculate the first The directional difference in the speed direction of each query: Here, clip(·, -1, 1) truncates the dot product to the range [-1, 1] to prevent numerical overflow; when the velocity directions of the two frames are completely consistent... =0; when the direction is opposite =2; the value range is [0,2], and the larger the value, the more violent the directional deflection.

[0046] Calculate the first Outliers in the directional difference of each query : in, For threshold parameters, For temperature parameters; A smoothing function for the direction; Calculate the first The speed difference of each query : Positive values ​​indicate acceleration, while negative values ​​indicate deceleration. The larger the absolute value of the difference, the more drastic the change in speed.

[0047] Calculate the first Outliers in the speed difference of individual queries : in, These are the threshold parameters for each dimension; For the smoothing function of velocity; Calculate the first The confidence difference of each query : Confidence difference measures the degree to which the detection confidence of the current frame decreases relative to historical frames; Calculate the first Outliers with poor confidence scores for each query : in, For threshold parameters, is a smoothing function for confidence levels; Calculate the comprehensive anomaly score for the i-th query. : in, , and All are learnable weights; Based on the aggregated anomaly scores, calculate the credibility weight for each query. : Among them, the credibility weight The physical meaning is intuitive: When approaching 0 (stable motion), If the value approaches 1, the query is subject to strong consistency constraints. When approaching 1 (abnormal motion height), As the value approaches zero, the constraints automatically weaken, allowing the model to make larger state updates in the current frame, thus adapting to real motion changes.

[0048] The short-term motion consistency interruption loss is calculated according to the following formula. : in, The balance coefficient for the direction consistency loss term. The balance coefficient for the velocity amplitude consistency loss term. The number of queries in the multimodal sample query set. This is a multimodal sample query set.

[0049] It should be noted that: The target movement is not required to be exactly the same as the historical movement; instead, it is determined by credibility weighting. The adjustment strikes a balance between constraints and permissible variation. For targets whose motion does change (such as normal turning or acceleration), their outlier scores will increase, but their weights will not drop sharply to zero. The model will still be subject to moderate consistency supervision, thus maintaining trajectory continuity during incremental updates.

[0050] Short-term motion consistency interruption loss in this embodiment It can perform propagation queries on the same target associated between adjacent frames, calculate three motion anomaly indicators: velocity direction deviation, velocity amplitude change, and confidence abnormal decrease, and convert them into adaptive gating weights. The consistency constraint strength is automatically reduced for motion anomaly queries.

[0051] The goal of the short-term motion consistency interruption loss is to establish motion consistency supervision between adjacent frames, so that the model maintains strong temporal dependence when the target motion is stable, and automatically reduces dependence on unreliable historical information when the target motion undergoes abnormal changes.

[0052] This embodiment proposes an explicit quantitative evaluation of the historical dependency reliability of inter-frame propagation queries using three complementary dimensions: velocity direction, velocity amplitude, and detection confidence. Existing temporal detection methods generally implicitly fuse historical queries through attention mechanisms, lacking explicit judgment on the reliability of historical information. This embodiment designs corresponding quantitative indicators for different causes of motion anomalies in real-world driving scenarios: velocity direction difference... Used to capture trajectory abrupt changes (such as sharp turns), and velocity amplitude differences. Used to detect acceleration / deceleration anomalies (such as emergency braking), with poor confidence level. It is used to capture perceptual quality degradation anomalies (such as occlusion and sensor noise). The three dimensions are independent and complementary to each other. No single dimension can fully cover all types of motion anomalies in real-world scenes. Together, they constitute a complete judgment system for historical reliability.

[0053] In some embodiments, a long-term velocity drift suppression loss value is determined based on the velocity of the category to which each query belongs in the multimodal sample query set of the current sample frame; including: Determine the first multimodal sample query set of the current sample frame The category to which each query belongs Reasonable upper limit of speed : in, These are the confidence interval coefficients; For category average speed For category The speed standard deviation; where, This corresponds to a 3σ confidence interval, meaning that approximately 99.7% of normal velocity distributions are allowed to pass the constraint, with penalties imposed only on extreme abnormal velocities. The value can be flexibly adjusted according to the actual scenario: The larger the value, the looser the constraints (suitable for scenarios with drastic speed changes); The smaller the value, the stricter the constraint (suitable for scenarios with relatively stable speed).

[0054] Calculate the first Speed ​​overflow of a single query : in, This represents the allowable upper limit tolerance margin for speed; The long-term velocity drift suppression loss value is calculated using the following formula. : in, For the first The weight of each query. This is the smoothed L1 loss function.

[0055] Long-term velocity drift suppression loss in this embodiment During training, the upper bound of the velocity distribution of each target category is statistically analyzed online using exponential moving average. SmoothL1 penalty is applied to velocity predictions that exceed the reasonable range to prevent abnormal drift in velocity predictions during long-term propagation.

[0056] When calculating the temporal consistency loss between adjacent frames, a smooth gating weight is introduced. By calculating the velocity direction difference, velocity amplitude difference, and confidence difference separately, these are mapped to smooth anomaly scores and aggregated into confidence weights. When motion anomalies are detected, the strength of the consistency constraint is automatically reduced (rather than hard truncation).

[0057] The decoupled design uses "confidence difference" only as an adjustment factor for the gating weights, without directly participating in the final loss value calculation (to avoid double penalty); and the differentiated smoothing mapping method uses the Sigmoid function to process continuous motion signals and the exponential activation function to process confidence-decreasing signals.

[0058] Throughout the training cycle, a velocity-motion prior for each target category is maintained, and a soft penalty is applied to velocity predictions that exceed the reasonable range of the prior to prevent velocity estimates from drifting during long-term temporal propagation.

[0059] For each detection category c (such as pedestrians, vehicles, motorcycles, etc.), maintain a separate velocity amplitude statistics library, which includes the following statistics: (1) Category speed mean: reflects the average speed of the target in this category; (2) Category velocity standard deviation: reflects the distribution width of the velocity of the target category; (3) Cumulative sample count: Record the number of speed samples that have been observed in this category.

[0060] The above statistics are updated online using the exponential moving average (EMA), with a decay factor of 1. For each newly observed velocity amplitude sample The update rules for statistics of each category are as follows: in, This is the attenuation factor, with a value of 0.99. and Categories The mean and standard deviation of the velocity are estimated. The above updates do not require storing all historical data; the statistics adaptively adjust during the training process.

[0061] To address the velocity drift problem in long-term propagation, instead of relying on manually set fixed thresholds, the exponential moving average (EMA) is used to statistically analyze the mean and standard deviation of velocity for each category online throughout the entire training period, dynamically constructing an upper bound for velocity with a tolerance margin (soft boundary).

[0062] This method combines categorical speed distribution statistics (mean + K standard deviations) with confidence filtering conditions to apply a long-term constraint of SmoothL1 penalty only to predictions with high confidence that exceed the dynamic upper bound. This method can effectively prevent extreme speed drift without inhibiting normal maneuvers (such as high-speed driving).

[0063] When calculating the weighted sum of the classification loss, the predicted bounding box regression loss, the short-term motion consistency disruption loss, and the long-term velocity drift suppression loss, the weight of the short-term motion consistency disruption loss can be set to 0.1, and the weight of the long-term velocity drift suppression loss can be set to 0.05.

[0064] Based on the same inventive concept, embodiments of this application provide a three-dimensional target detection device based on spatiotemporal adaptive representation control, see reference. Figure 2As shown, the three-dimensional target detection device 200 based on spatiotemporal adaptive representation control provided in this application embodiment includes at least: The acquisition unit 201 is used to acquire multi-view images and point cloud data at the current moment; The image processing unit 202 is used to extract features from multi-view images using an image backbone network to obtain semantic features; and to process the semantic features using an image query generator to obtain an image query set. The point cloud processing unit 203 is used to extract features from point cloud data using the point cloud backbone network to obtain multiple voxel features; and to process the multiple voxel features using the point cloud query generator to obtain the first point cloud query set. Determining unit 204 is used to determine the second point cloud query set based on the filtered voxel features and the first point cloud query set; The combination unit 205 is used to combine the image query set and the second point cloud query set to obtain the multimodal query set at the current time. The processing unit 206 is used to process the current multimodal query set and the historical multimodal query set using a cross-modal interactive attention model to obtain the current three-dimensional target detection result. The historical multimodal query set includes the multimodal query sets of multiple consecutive frames before the current time.

[0065] It should be noted that the principle of the three-dimensional target detection device 200 based on spatiotemporal adaptive representation control provided in this application embodiment to solve the technical problem is similar to the method provided in this application embodiment. Therefore, the implementation of the three-dimensional target detection device 200 based on spatiotemporal adaptive representation control provided in this application embodiment can refer to the implementation of the method provided in this application embodiment, and the repeated parts will not be described again.

[0066] Based on the same inventive concept, embodiments of this application also provide an electronic device, such as... Figure 3 As shown, it includes a memory and a processor. The memory stores an executable program, and the processor executes the executable program to implement the steps of the three-dimensional target detection method based on spatiotemporal adaptive representation regulation provided in the above embodiments.

[0067] The aforementioned processor can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0068] Since the electronic device described in this application embodiment is an electronic device equipped with a memory for implementing the three-dimensional target detection method based on spatiotemporal adaptive representation control disclosed in this application embodiment, those skilled in the art can understand the structure and variations of the electronic device described in this application embodiment based on the three-dimensional target detection method based on spatiotemporal adaptive representation control disclosed in this application embodiment, and therefore will not be described again here.

[0069] This application also provides a computer-readable storage medium storing a computer program thereon. When the computer program is run by a processor, it implements the steps of the three-dimensional target detection method based on spatiotemporal adaptive representation control provided in the above embodiments.

[0070] The storage medium in this embodiment may be included in an electronic device; or it may exist independently and not be assembled into an electronic device. The storage medium carries one or more computer programs, which, when executed, implement the steps of the three-dimensional target detection method based on spatiotemporal adaptive representation control provided in the above embodiments.

[0071] It should be understood that the various solutions in this embodiment have the same technical effects as those in the above method embodiments, and will not be repeated here.

[0072] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. Optionally, specific examples in this embodiment can refer to the examples described in any embodiment of this application, which will not be repeated here. Obviously, those skilled in the art should understand that the various modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular hardware and software combination.

[0073] This application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the three-dimensional target detection method based on spatiotemporal adaptive representation modulation provided in the above embodiments.

[0074] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions targeted in the blocks may occur in a different order than those targeted in the drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0075] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

[0076] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

Claims

1. A three-dimensional target detection method based on spatiotemporal adaptive representation control, characterized in that, include: Acquire multi-view images and point cloud data at the current moment; Semantic features are obtained by extracting features from multi-view images using an image backbone network. The semantic features are processed using an image query generator to obtain an image query set; Feature extraction is performed on point cloud data using a point cloud backbone network to obtain multiple voxel features; the point cloud query generator is then used to process the multiple voxel features to obtain the first point cloud query set. Based on the filtered voxel features and the first point cloud query set, the second point cloud query set is determined. The image query set and the second point cloud query set are combined to obtain the multimodal query set at the current moment; The current multimodal query set and the historical multimodal query set are processed using a cross-modal interactive attention model to obtain the current three-dimensional target detection result. The historical multimodal query set includes the multimodal query sets of consecutive frames before the current time.

2. The method according to claim 1, characterized in that, Based on the filtered voxel features and the first point cloud query set, a second point cloud query set is determined; including: Multiple voxel features are filtered based on a geometric perception scoring filtering strategy to obtain the filtered voxel features. Combine the second point cloud query sets of multiple consecutive frames before the current moment into a historical point cloud query set; The selected voxel features, the first point cloud query set, and the historical point cloud query set are processed using a point cloud refinement attention model to obtain the second point cloud query set at the current time.

3. The method according to claim 2, characterized in that, Multiple voxel features are filtered using a geometry-based perceptual scoring strategy to obtain the filtered voxel features, including: Construct the first Nearest neighbor set of individual feature Nearest neighbor set Include Individual characteristics; Calculate the first Local density of individual characteristics : in, It is the numerical stability constant; For the first Three-dimensional coordinates of individual features; For the first Three-dimensional coordinates of individual features; Calculate the first Semantic consistency of individual features : in, For the first Feature embedding vector of individual primitive features; For the first Feature embedding vector of individual primitive features; Calculate the first Local feature variation of individual phenotypic characteristics : Calculate the first Geometric significance score of individual primitive features : in, , and All are learnable weights; All voxel features are sorted in descending order based on geometric significance scores to obtain a voxel feature sequence; the top voxel features are then selected from the voxel feature sequence. Individual voxel features are used as the voxel features after screening. This represents the total number of voxel features selected.

4. The method according to claim 2, characterized in that, The point cloud refinement attention model includes a self-attention layer, a cross-attention layer, and a feedforward network layer connected in sequence. The selected voxel features, the first point cloud query set, and the historical point cloud query set are processed using a point cloud refinement attention model to obtain the second point cloud query set at the current time, including: The current moment First point cloud query set With historical query set Combine them to obtain the query set. ; Using a self-attention layer to query the first point cloud set and query set Processing is performed to obtain the time-series context query. : Query by time series context For the query, the filtered voxel features The key and value are input to a cross-attention layer for processing, resulting in a query set. : Using feedforward network layers for query sets Processing yields the second cloud query set at the current moment. .

5. The method according to claim 2, characterized in that, The method further includes: Establish a training set, which includes sample frames of multiple consecutive frames. Each sample frame includes: multi-view image samples with labeled target bounding boxes and point cloud samples. The semantic features of the current sample frame are obtained by extracting features from the multi-view images of the current sample frame using an image backbone network; the semantic features of the current sample frame are then processed by an image query generator to obtain an image query set for the current sample frame. The voxel features of the current sample frame are obtained by using the point cloud backbone network to extract features from the point cloud samples of the current sample frame; the voxel features of the current sample frame are processed by the point cloud query generator to obtain the first point cloud query set of the current sample frame. Based on the voxel features after filtering the current sample frame and the first point cloud query set, determine the second point cloud query set; The image query set and the second point cloud query set of the current sample frame are combined to obtain the multimodal query set of the current sample frame; The target prediction box is obtained by processing the multimodal query set of the current sample frame and the multimodal query set of the previous consecutive frames using a cross-modal interactive attention model. Based on the target predicted bounding box and the target ground truth bounding box, determine the classification loss value and the predicted bounding box regression loss value; Based on the velocity vector and classification confidence of each query in the multimodal query set of the current sample frame and the multimodal query set of the previous sample frame, the short-term motion consistency interruption loss value is determined. Based on the velocity of the category to which each query belongs in the multimodal query set of the current sample frame, determine the long-term velocity drift suppression loss value; The parameters of the image query generator, point cloud query generator, point cloud refinement attention model, and cross-modal interactive attention model are updated based on the weighted sum of classification loss, predicted box regression loss, short-term motion consistency interruption loss, and long-term velocity drift suppression loss.

6. The method according to claim 5, characterized in that, Based on the voxel features filtered from the current sample frame and the first point cloud query set, a second point cloud query set is determined; including: A filtering strategy is determined based on the current number of training iterations to filter multiple voxel features of the current sample frame; the filtering strategy is a confidence-guided filtering strategy, a geometry-aware scoring filtering strategy, or a hybrid strategy; the hybrid strategy includes a confidence-guided filtering strategy and a geometry-aware scoring filtering strategy. Based on the filtering strategy, multiple voxel features of the current sample frame are filtered to obtain the filtered voxel features of the current sample frame. The point cloud refinement attention model is used to process the voxel features after filtering the current sample frame, the first point cloud query set of the current sample frame, and the second point cloud query set of the previous multiple consecutive frames to obtain the second point cloud query set of the current sample frame.

7. The method according to claim 6, characterized in that, When the screening strategy is a confidence-guided screening strategy; Based on a filtering strategy, multiple voxel features of the current sample frame are filtered to obtain the filtered voxel features of the current sample frame; including: Step S1: Obtain the classification confidence and target box center coordinates of multiple targets in the current sample frame output by the point cloud backbone network; Step S2: Sort the multiple targets in descending order based on classification confidence to obtain the target sequence; Step S3: Using the first [number]th ... The origin is used as the center of the target bounding box for each target to define a spatial neighborhood with radius r. ; Step S4: In the spatial neighborhood sampling Individual pixel feature samples, where, when Larger than spatial neighborhood The number of voxel feature samples in the data will then be... Set as spatial neighborhood The number of voxel feature samples in the sample; Step S5: Determine whether the cumulative number of sampled voxel features has reached K, where K is the total number of selected voxel features; if yes, end; otherwise, proceed. Updated to Proceed to step S3.

8. The method according to claim 6, characterized in that, When the filtering strategy is a hybrid strategy, the method further includes: Calculate the current training iteration number Transition coefficient : in, This is the preset number of transition start iterations. This is the length of the transition window; Calculate the number of voxel feature samples selected by executing the geometry-aware scoring screening strategy. : ; in, The total number of voxel features selected; Calculate the number of voxel feature samples selected by executing the confidence-guided screening strategy. : 。 9. The method according to claim 5, characterized in that, Based on the velocity vector and classification confidence of each query in the multimodal query set of the current sample frame and the multimodal query set of the previous sample frame, the short-term motion consistency interruption loss value is determined. include: Get the first multimodal query set of the current sample frame The query is in the same frame as the previous one. velocity vector and classification confidence Current sample frame velocity vector and classification confidence , The sequence number of the current sample frame; Calculate the first The query is in the same frame as the previous one. velocity amplitude and the current sample frame velocity amplitude : Calculate the first The query is in the same frame as the previous one. unit direction vector and the current sample frame unit direction vector : in, It is the numerical stability constant; Calculate the first The directional difference in the speed direction of each query: The clip(·, -1, 1) function truncates the dot product to the range [-1, 1] to prevent numerical overflow. Calculate the first Outliers in the directional difference of each query : in, For threshold parameters, For temperature parameters; A smoothing function for the direction; Calculate the first The speed difference of each query : Calculate the first Outliers in the speed difference of individual queries : in, These are the threshold parameters for each dimension; For the smoothing function of velocity; Calculate the first The confidence difference of each query : Calculate the first Outliers with poor confidence scores for each query : in, For threshold parameters, is a smoothing function for confidence levels; Calculate the first Comprehensive anomaly score for each query : in, , and All are learnable weights; Based on the aggregated anomaly score, calculate the first... The credibility weight of each query : The short-term motion consistency interruption loss is calculated according to the following formula. : in, The balance coefficient for the direction consistency loss term. The balance coefficient for the velocity amplitude consistency loss term. The number of queries in the multimodal sample query set. This is a multimodal sample query set.

10. The method according to claim 9, characterized in that, Based on the velocity of each query in the multimodal sample query set of the current sample frame, determine the long-term velocity drift suppression loss value; include: Determine the first multimodal sample query set of the current sample frame The category to which each query belongs Reasonable upper limit of speed : in, These are the confidence interval coefficients; For category average speed For category The speed standard deviation; Calculate the first Speed ​​overflow of a single query : in, This represents the allowable upper limit tolerance margin for speed; The long-term velocity drift suppression loss value is calculated using the following formula. : in, For the first The weight of each query. This is the smoothed L1 loss function.