Cooperative 3D target detection and tracking method based on deep learning
By constructing a collaborative 3D object detection and tracking method based on deep learning, using the spatial information priority map to filter key areas and combining with the entropy encoder for feature compression, the perception inaccuracy of the bicycle perception system is solved, and efficient perception and object detection under limited communication resources are achieved.
Patent Information
- Application Number
- CN202510400676.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-08
AI Technical Summary
In complex dynamic driving environments, the bicycle perception system is inaccurate due to problems such as occlusion, limited detection range and sparse sensor observation, which affects the safety and reliability of autonomous driving. The existing collaborative perception methods fail to effectively utilize the time dimension characteristics, resulting in limited performance improvement.
A collaborative 3D object detection and tracking method based on deep learning is constructed, feature maps are extracted through feature encoder, key areas are screened using spatial information priority maps, feature compression is performed in combination with entropy encoder, and targets are tracked through Kalman filters to achieve spatiotemporal fusion to optimize information transmission and perception accuracy.
Under limited communication resources, perception ability and communication efficiency are improved, target detection effect is enhanced, data transmission is reduced, and detection accuracy is maintained when facing positioning errors.
Smart Images

Figure CN120279253A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of collaborative perception, and in particular, to a collaborative 3D object detection and tracking method based on deep learning. Background Art
[0002] In a complex and dynamic driving environment, the perception ability of intelligent vehicles is crucial for ensuring road safety and improving traffic efficiency. With the progress of deep learning, the performance of single-vehicle perception systems in autonomous driving has been significantly improved, showing reliable results in tasks such as instance segmentation and object detection. However, under actual driving conditions, single-vehicle perception systems face significant limitations, such as occlusion, limited detection range, and sparse sensor observations. These limitations often lead to inaccurate perception of complex driving scenarios, thus affecting the safety and reliability of autonomous driving.
[0003] To address these issues, multi-agent collaborative perception has gradually received attention as an emerging and promising solution. Through vehicle-to-vehicle (V2V) or vehicle-to-everything (V2X) communication, collaborative vehicles can share perception information, thereby expanding the perception range of a single vehicle and improving environmental understanding.
[0004] Collaborative perception integrates the data of collaborative agents into the autonomous vehicle through fusion methods, which can be classified into three categories: early fusion, late fusion, and intermediate fusion. In early fusion, the raw data of multiple collaborative agents are fused on the autonomous vehicle to form a unified representation for subsequent processing. This method has high detection accuracy, but requires a large amount of data transmission and is highly sensitive to noise and latency. Given the current limitations of V2X communication bandwidth and stability, implementing reliable early fusion is challenging. In late fusion, each collaborative agent generates detection results through a deep learning model and exchanges them. This approach reduces the communication burden by reducing the amount of data, but the improvement in perception accuracy is limited. The intermediate fusion method, by sharing the feature maps generated by the deep learning model, significantly reduces the communication cost while ensuring high accuracy, and is therefore considered a more promising solution.
[0005] To achieve efficient collaborative perception, some intermediate fusion-based methods have been recently proposed, such as V2VNet, V2X-ViT, When2Com, and Where2Com. Among them, the literature [R. Xu, H. Xiang, Z. Tu, X. Xia, M.-H. Yang, and J. Ma, “V2x-vit: Vehicle-to-everything cooperative perception with vision transformer,” in European conference on computer vision, pp. 107–124, Springer, 2022.] uses a heterogeneous multi-agent self-attention mechanism to address the device heterogeneity problem; the literature [Y. Hu, S. Fang, Z. Lei, Y. Zhong, and S. Chen, “Where2comm: Communication-efficient collaborative perception via spatial confidence maps,” Advances in Neural Information Processing Systems, vol. 35, pp. 4874–4886, 2022.] minimizes the bandwidth requirement through a spatial confidence perception system. These above-mentioned studies utilize techniques such as handshake communication mechanisms, agent selection, message selection, and feature compression. However, these studies mainly focus on the spatial dimension of semantic information and ignore the selection of temporal dimension features. Due to the heterogeneity of sensors among agents and the differences in the time required for information transmission, the shared semantic information in the actual scenario will inevitably be delayed. If the temporal dimension is ignored, the performance improvement in collaborative perception may be limited or even degraded. For example, relying on outdated information that is no longer in the area of interest may confuse the autonomous vehicle and mislead the perception result. Therefore, in collaborative perception, it is crucial to comprehensively consider the spatial and temporal dimensions of semantic information. Temporal feature selection can not only reduce communication costs but also effectively alleviate transmission and queuing delays. In addition, the autonomous vehicle can utilize historical collaborative information to optimize the current semantic features, thereby further enhancing the perception ability. Besides the selection-based methods, another type of collaborative perception method is feature compression. Feature compression methods attempt to transform the original input data into an intermediate representation while retaining the information required for downstream tasks. Convolutional neural networks are usually used as encoders to reduce the feature dimension, and then the features are quantized and compressed through entropy coding. These methods need to find a balance between the amount of data transmitted (measured in average bytes) and the performance of downstream tasks.
[0006] Recent research, such as the literature [K.Yang, D.Yang, J.Zhang, H.Wang, P.Sun, and L.Song, “What2comm: Towards Communication-efficient Collaborative Perception via Feature Decoupling,” in Proceedings of the 31st ACM International Conference on Multimedia, (Ottawa ON Canada), pp.7686–7695, ACM, Oct. 2023.], has explored methods that combine feature selection and feature compression, aiming to further improve efficiency and task performance. However, in these methods, these two stages are usually designed separately, where feature selection is applied to the spatial dimension, while feature compression is applied to the channel dimension. This independent design has inherent limitations. For example, compression in the channel dimension cannot identify key foreground regions, thus unable to achieve the optimal bitrate allocation for detection tasks. Summary of the Invention
[0007] To solve the problems existing in the prior art, the present invention provides a collaborative 3D object detection and tracking method based on deep learning, which can more efficiently utilize and encode spatio-temporal semantic information.
[0008] The present invention can address the perception limitations of agents in object detection, such as the fact that the field of view is easily blocked by other objects and the perception accuracy decreases with increasing distance. The agents can solve the inherent limitations of current single-agent perception through data sharing and fusion. At the same time, through end-to-end optimized rate distortion, neural compression can guide bit allocation according to the importance of semantic information during the feature selection stage, thus efficiently combining feature selection with quantization and entropy coding to achieve a better trade-off between average communication bits and perception accuracy.
[0009] The technical solution adopted by the present invention to solve the technical problems is as follows:
[0010] A collaborative 3D object detection and tracking method based on deep learning provided by the present invention includes the following steps:
[0011] Step 1: Construct a collaborative perception framework based on feature extraction;
[0012] The collaborative perception framework includes: a feature encoder, a feature map filter, an entropy encoder, a semantic information fusion module, an object detection decoder, and a tracker;
[0013] Step 2: Construct an optimization problem;
[0014] Assume that at time frame t, there are a total of N vehicles, 1 autonomous vehicle i, and N - 1 cooperative agents;
[0015] The objective function of the perception task is defined as:
[0016]
[0017] where T represents the number of frames within a period of time, represents the sparse feature map, represents the set of collaborators of the i-th cooperative agent at time t, φ θ represents the neural network Φ with parameter θ, f eva (·) represents the evaluation function of collaborative perception. In the experiment, the average precision of object detection is adopted. The max operation is used to maximize the gain G, and represent the observation value and the ground truth value of the autonomous vehicle i at time frame t, respectively;
[0018] The objective function of the collaborative perception framework is defined as:
[0019]
[0020] where λ is used to balance the relationship between the gain and the feature encoding rate. E represents taking the expectation, q j->i represents the quantization value of the semantic feature, represents the probability of the quantization value of the semantic feature occurring;
[0021] Step 3: Perception feature extraction;
[0022] The cooperative agent j uses the feature encoder to extract features from the original sensor data to obtain the feature map The autonomous vehicle i uses the feature encoder to extract features from the original sensor data to obtain the feature map
[0023] Step 4: Transmission feature screening;
[0024] The feature map is converted into a spatial information priority map through the spatial information priority generator in the feature map filter, and the dynamic change index of the current time frame is calculated. By excluding the regions with relatively small dynamic changes through time selection, the cooperative agent j uses the feature map filter to identify the regions beneficial to the autonomous vehicle i from its original feature map to obtain the sparse feature map
[0025] Step 5: Compression, encoding, and decoding of perception features;
[0026] The sending end first quantizes the sparse feature map and then converts it into a bitstream through an entropy encoder for transmission;
[0027] Step 6: Reception and fusion of perceptual features;
[0028] Autonomous vehicle i receives the bitstream transmitted by neighboring intelligent agent j, decodes the bitstream using an entropy decoder, and extracts the decompressed sparse feature map The decompressed sparse feature map is fused with the local feature map F i t to perform vehicle - to - vehicle feature fusion and obtain a spatially perception - enhanced feature map The spatially perception - enhanced feature map is fused with the time - fused feature map containing historical information to perform fusion in the time dimension and obtain the spatio - temporal fusion feature map of the current time frame
[0029] Step 7: Object detection and tracking;
[0030] The spatio - temporal fusion feature map is converted into 3D object detection results through an object detection decoder :
[0031]
[0032] where, Φ decoder represents the object detection decoder, represents the 3D object detection result, i.e., the ground truth of autonomous vehicle i at time frame t;
[0033] At each time frame, a tracker is initialized for each newly detected object. The tracker associates the 3D object detection results of each time frame to form a trajectory for each detection bounding box; based on the historical trajectory, the Kalman filter is used to predict the position of the detection bounding box in the current time frame; the trajectory of the detection bounding box is a multi - dimensional variable, including the speed, position, and orientation of the object:
[0034]
[0035] where, φ predict represents the tracker, represents the detection bounding box of autonomous vehicle i at time t, represents the detection bounding box of autonomous vehicle i at time t - 1;
[0036] The 3D object detection results are optimized using the trajectory of the detection bounding box and the following formula:
[0037]
[0038] Among them, f box fusion represents the fusion of the detected bounding box, and represents the final 3D object detection output result;
[0039] The Kalman filter is used to update the trajectory of the detected bounding box using the final 3D object detection output result :
[0040]
[0041] Among them, Φ update represents the update of the Kalman filter.
[0042] Furthermore, the collaborative perception framework accepts single-modal input and multi-modal input. For RGB image input, after passing through the feature encoder, a transformation function that converts the front view to BEV is used to obtain a feature map from the BEV perspective For 3D point cloud data input, it is discretely converted to BEV, and then through the feature encoder, a feature map from the BEV perspective is obtained from the 3D point cloud data using the anchor-based PointPillar method
[0043] Furthermore, the specific implementation process of step four is as follows:
[0044] S4.1: Obtain the spatial information priority map;
[0045] The spatial information priority generator consists of the target decoder Φ priority-decoder and generates the spatial information priority map priority-decoder from the target decoder Φ
[0046]
[0047] S4.2: Information entropy calculation;
[0048] Each element in the spatial information priority map represents the probability of containing a target. The object detection problem is regarded as a binary classification task, where each class represents whether there is an object at the current position; according to the definition of information theory, use to represent the information entropy map of the j-th collaborative agent, which represents the information entropy map of the j-th collaborative agent at the position (h, w):
[0049]
[0050] Among them, the loop variable k represents the class, Indicates the information priority map for the k-th category at coordinates (h, w); at each coordinate (h, w), it satisfies:
[0051] P (h,w,0) +P (h,w,1) = 1
[0052] Estimate the dynamic change level at each position by comparing the information entropy differences between adjacent frames of feature maps;
[0053] S4.3: Spatial sampling;
[0054] Filter out the spatial regions worth transmitting from the spatial information priority map through spatial sampling operations and generate a binary selection matrix, so as to only transmit the spatial regions worth transmitting and omit the background regions and static regions that are not worth transmitting;
[0055] S4.4: Dynamic characteristic evaluation;
[0056] Feature dynamic change index Is expressed as:
[0057]
[0058] Where, Represents the information entropy map of the j-th cooperative agent at time t - 1;
[0059] For the cooperative agent j, its dynamic change index is defined as:
[0060]
[0061] Where, Φ map (·) is used to convert the difference of the trajectory into the corresponding coordinates in the BEV map, Represents the detection bounding box of the cooperative agent j at time t, ) represents the detection bounding box of the cooperative agent j at time t - 1;
[0062] S4.5: Temporal selection process;
[0063] According to the feature dynamic change index Whether it exceeds the threshold to judge whether the current time frame is a key frame, and evaluate the positions with significant temporal dynamic changes compared with the previous time frame in the spatial regions selected by spatial sampling, and retain the elements at these positions;
[0064] For key frames, only apply spatial sampling without temporal selection; for non-key frames, perform spatio-temporal sampling in sequence; by integrating these two cases, the binary selection matrix is defined as Represents the spatial selection matrix, It represents a time selection matrix, and ⊙ is element-wise multiplication;
[0065] S4.6: Obtain a sparse feature map;
[0066] Select a feature map through a feature map filter according to a binary selection matrix Select the region in to obtain a sparse feature map
[0067]
[0068] where ⊙ is element-wise multiplication.
[0069] Furthermore, in step S4.3, the process of spatial sampling is expressed as:
[0070]
[0071] where represents a binary selection matrix. The regions with a value of 1 in the binary selection matrix are the spatial regions worth transmitting, and the regions with a value of 0 are the spatial regions not worth transmitting; S select (·) is a selection function. This selection function preferentially sets the spatial regions worth transmitting to 1 and the spatial regions not worth transmitting to 0 until the information of the spatial regions worth transmitting reaches the critical value of the limited communication volume V.
[0072] Furthermore, in step S4.5, the time selection process can be expressed as:
[0073]
[0074] where represents the time selection matrix of the j-th user at the coordinates (h, w) at time t; represents an indicator function, which takes the value of 1 when the condition holds, and 0 otherwise; represents the degree of temporal variation; δ and τ represent selection thresholds.
[0075] Furthermore, the specific implementation process of step five is as follows:
[0076] (a) Perform scalar quantization on the selected sparse feature map to generate a discrete vector
[0077]
[0078] where round(·) represents the operation of rounding each element to the nearest integer;
[0079] (b) Use the arithmetic coding method to perform lossless compression on the discrete vector through an entropy encoder ;
[0080] (c) Use the probability mass function to encode the discrete vector into a bitstream through an entropy encoder, and its transmission code rate is denotes the prior probability model of the discrete vector , that is, the entropy model, and ω represents the prior probability model parameter; the probability mass function can be calculated through the prior probability model ;
[0081] (d) In the arithmetic decoding process, first perform entropy decoding on the discrete vector through the bitstream, and then input the decoded sparse feature map into the inter-vehicle feature fusion module;
[0082] Furthermore, the calculation formula for the transmission code rate R of the sparse feature map is:
[0083]
[0084] where denotes the estimated distribution function of the sparse feature map , denotes the discrete vector generated by scalar quantization of the sparse feature map , and l represents the position of the element in ;
[0085] The loss function of the optimization problem is:
[0086]
[0087] where L det (·) represents the detection loss between the prediction target and the true value, and L rate represents the loss function of the transmission code rate R, represents the prediction target, represents the true value;
[0088] The calculation formula for the loss function L rate is:
[0089]
[0090] where denotes the occurrence probability of the selected feature map weighted by quality; m j is the quality map, which is used to represent the importance level of each element in the cooperative agent j, and its importance is determined by the spatial information priority map And the information entropy map is measured as follows:
[0091]
[0092] Furthermore, the feature map with enhanced spatial perception has the following mathematical expression:
[0093]
[0094] where Φ spatial-fusion represents the spatial fusion function;
[0095] The spatio-temporal fusion feature map has the following mathematical expression:
[0096]
[0097] where the max operation is used to select the maximum value from the feature maps of neighboring agents and autonomous vehicle i at each position; represents the temporal fusion feature map containing historical information, and its mathematical expression is:
[0098]
[0099] where (h, w) represents the position coordinates in the optical flow map, and (v x , v y ) represents the velocity vector corresponding to each element at the position (h, w), represents the fusion feature map of the previous time frame.
[0100] Furthermore, in step seven, the tracker associates the target tracked by the Kalman filter with the current target detection result through an intersection over union threshold; after the tracking box is associated with the detection bounding box, the 3D IOU value is used as the matching score; the detection bounding box b is optimized by the tracking box t , and the coordinates of the detection bounding box b t are obtained by weighted calculation from the tracking score s t and the detection score c t , and its specific calculation formula is as follows:
[0101]
[0102] where γ represents the decay factor.
[0103] Furthermore, it also includes the training step of the collaborative perception framework based on feature extraction;
[0104] In the first stage, the parameters of the feature map filter are optimized;
[0105] For the selection function S select (·) and τ, set lower thresholds to filter out most of the background regions that do not contain the target while retaining all the features that contribute to collaborative perception;
[0106] In the second stage, using the discrete vector and the quality map m j as guidance, calculate the loss function of the transmission code rate R Then calculate the loss function of the optimization problem; through the optimization formula perform end-to-end training on the entire collaborative perception framework including the prior probability model of the entropy encoder, where L det (·) represents the detection loss between the predicted target and the ground truth, and L rate represents the loss function of the transmission code rate R, represents the predicted target, represents the ground truth.
[0107] The beneficial effects of the present invention are:
[0108] The present invention provides a collaborative 3D object detection and tracking method based on deep learning, aiming to solve the problem of how to select the most critical information for transmission under limited communication resources, thereby improving the perception ability and communication efficiency of the entire system. The core idea of the present invention is to use the spatial information priority graph to characterize the density distribution of valuable information in the space, and then use the temporal information screening to further remove the temporal redundant information, thereby optimizing the collaborative perception process of the multi-agent system. At the same time, the present invention also tracks and predicts the movement of the target through the Kalman filter, and adds a time adjustment factor based on the spatial information priority graph, making full use of the temporal correlation of the detection results, and further improving the perception ability of the system under limited bandwidth.
[0109] In addition, the present invention uses a factored hyperprior model to encode the extracted features, convert them into latent representations, and then quantize and encode them into bitstreams. By combining this process with semantic information extraction, the present invention realizes bitrate allocation based on the importance of semantic information, thereby further reducing the amount of data transmission. At the autonomous vehicle end, the present invention uses the Kalman filter to track the targets in the historical frames and predict their current positions. These prediction results are fused with the current detection bounding boxes to enhance the object detection effect.
[0110] To verify the performance of the method, the present invention conducts simulation tests on the cooperative perception datasets OPV2V and Dair-V2X, and compares with models such as V2VNet, V2X-ViT, and Where2Comm. The results show that the method of the present invention has achieved significant improvements in the trade-off between communication and perception performance, obtained effective performance gains, and verified the perception performance advantage of the method of the present invention under limited communication bandwidth. Description of the Drawings
[0111] Figure 1 It is a schematic diagram of a collaborative perception framework based on feature extraction constructed by the present invention.
[0112] Figure 2 It is a specific implementation flowchart of the feature map filter.
[0113] Figure 3 It is a specific implementation flowchart of the entropy encoding and decoding.
[0114] Figure 4 It is a comparison of the trade-off between perception performance and communication bandwidth for different cooperation algorithms.
[0115] Figure 5 It is a comparison chart of the target detection performance of different cooperation algorithms under different positioning error conditions. Detailed Description of the Invention
[0116] The following further describes the present invention in detail with reference to the drawings.
[0117] A collaborative 3D target detection and tracking method based on deep learning provided by the present invention has the following specific implementation process:
[0118] Step 1: Construct a collaborative perception framework based on feature extraction;
[0119] The present invention constructs a collaborative perception framework based on feature extraction, which is used to extract sparse but key spatio-temporal features from the original LiDAR perception data in a bandwidth-limited environment, improve the perception ability of the autonomous vehicle by integrating the information of multiple agents, and thus achieve efficient 3D target detection.
[0120] As Figure 1 shown, the collaborative perception framework based on feature extraction mainly includes the following parts: a feature encoder, a feature map filter, an entropy encoder, a semantic information fusion module, a target detection decoder, and a tracker.
[0121] Assume that there are N vehicles in a certain time frame t, where 1 vehicle is selected as the autonomous vehicle, and the remaining vehicles are cooperative agents, that is Figure 1Cooperative vehicles therein. Each cooperative agent extracts semantic features from its observations and transmits them to the autonomous vehicle. The autonomous vehicle aims to maximize the perception accuracy by integrating these shared semantic features, while the cooperative agents focus on minimizing the communication bandwidth.
[0122] Step 2: Construct the optimization problem;
[0123] First, let and Y i t represent the observation value and the ground truth of autonomous vehicle i at time frame t, respectively. The present invention uses the evaluation function f eva (·) to evaluate the perception performance of autonomous vehicle i. Therefore, the objective function of the perception task is expressed as:
[0124]
[0125] where T represents the number of frames within a period of time, represents the sparse feature map, represents the set of collaborators of the i-th cooperative agent at time t, φ θ represents the neural network Φ with parameter θ, f eva (·) represents the evaluation function of cooperative perception. In the experiment, the average precision of object detection is adopted. The max operation is used to maximize the gain G, and represent the observation value and the ground truth of autonomous vehicle i at time frame t, respectively.
[0126] In addition, based on the objective function of the perception task, the objective function of the entire system framework is further considered, which is to minimize the transmission rate R of the cooperative agents while maximizing the perception accuracy. Since reasonable design of entropy coding can make the achievable transmission rate close to the entropy. The present invention estimates the transmission rate R through entropy and defines the objective function of the entire system framework as:
[0127]
[0128] where λ is used to balance the relationship between the gain and the feature coding rate. E represents taking the expectation. q j->i represents the quantization value of the semantic feature, represents the probability of the quantization value of the semantic feature occurring; the mathematical expression of the quantization value q j->i of the semantic feature is as follows:
[0129]
[0130] where round(·) represents the operation of rounding each element to the nearest integer.
[0131] Step 3: Perceptual feature extraction;
[0132] At each time frame t, given the original sensor data, i.e., the observation of cooperative agent j at time frame t Cooperative agent j uses a feature encoder to extract features from the original sensor data X j t to obtain a feature map This process can be expressed as:
[0133]
[0134] where H, W, and C represent the height, width, and number of channels of the feature respectively, and f enc (·) represents the feature encoder.
[0135] At each time frame t, autonomous vehicle i uses a feature encoder to extract features from the original sensor data to obtain a feature map In this process, the data is also transformed to a consistent bird's-eye view (BEV) perspective so that the perceptual information of different agents can be represented in the same global coordinate system.
[0136] The collaborative perception framework constructed in the present invention accepts single / multi-modal inputs, such as RGB images and 3D point clouds. For RGB image input, after passing through the feature encoder, a transformation function that converts the front view to BEV is used to obtain a feature map in the BEV perspective For 3D point cloud data input, it is discretely transformed to BEV, and then a feature encoder uses the anchor-based PointPillar method to obtain a feature map in the BEV perspective from the 3D point cloud
[0137] Step 4: Transmission feature screening;
[0138] S4.1: Obtain the spatial information priority map;
[0139] As Figure 2 shown, the spatial information priority generator in the feature map filter is responsible for converting the feature map into a spatial information priority map and calculating the dynamic change index for the current time frame. The spatial information priority map reflects the criticality of the perceptual information in different spatial regions.
[0140] Specifically, for the object detection task of the present invention, the regions containing objects are more important than the background regions without objects. During the collaboration process, the region information with objects can help other agents obtain the object information that cannot be perceived due to limited views, so as to detect the objects that cannot be detected in single-agent object detection.
[0141] Among them, the spatial information priority generator consists of an object decoder Φ priority-decoder and is composed of the object decoder Φ priority-decoder to generate a spatial information priority map This process can be expressed as:
[0142]
[0143] S4.2: Information entropy calculation;
[0144] The spatial information priority map can identify the spatial regions worthy of sharing and provide information gain for the autonomous vehicle. However, only using this spatial selection metric is not sufficient to judge the importance of the selected feature map regions in terms of time. To determine which positions in the feature map have undergone significant dynamic changes and thus contribute to the real-time information gain, the present invention evaluates its dynamics by comparing the information entropy differences between two adjacent frames of feature maps.
[0145] Each element in the spatial information priority map represents the probability of containing an object. Accordingly, the object detection problem is regarded as a binary classification task, where each class represents whether there is an object at the current position. According to the definition of information theory, use to represent the information entropy map of the jth collaborative agent, represent the value of the information entropy map of the jth collaborative agent at the position (h, w), and the specific calculation formula of this value is:
[0146]
[0147] where the loop variable k represents the class, represents the information priority map of the kth class at the coordinates (h, w), and at each coordinate (h, w) it satisfies:
[0148] P (h,w,0) +P (h,w,1) =1
[0149] By comparing the information entropy differences between two adjacent frames of feature maps, the dynamic change level of each position can be estimated. Its intuitive meaning is that the objects in the same scene have unique dynamic change degrees. By assigning different sampling frequencies to different objects, the communication bandwidth can be utilized more efficiently.
[0150] S4.3: Spatial sampling;
[0151] Through the spatial sampling operation, the spatial regions worth transmitting can be screened out from the spatial information priority map, and a binary selection matrix can be generated, so as to only transmit the spatial regions worth transmitting and omit the background regions and static regions not worth transmitting, saving bandwidth while improving the perception ability.
[0152] Specifically, the process of spatial sampling can be expressed as:
[0153]
[0154] Among them, represents the binary selection matrix. In the binary selection matrix the regions with a value of 1 are the spatial regions worth transmitting, and the regions with a value of 0 are the spatial regions not worth transmitting. S select S(·) is a selection function. This selection function preferentially sets the regions where the target appears with high probability, that is, the spatial regions worth transmitting, to 1, and then sets the spatial regions not worth transmitting to 0 until the information of the spatial regions worth transmitting reaches the critical value of the limited communication volume V.
[0155] S4.4: Dynamic characteristic evaluation;
[0156] Feature dynamic change index can be expressed as:
[0157]
[0158] Among them, represents the information entropy map of the j-th cooperative agent at time t-1.
[0159] For the cooperative agent j, its dynamic change index is defined as:
[0160]
[0161] Among them, Φ map (·) is used to convert the difference of the trajectory into the corresponding coordinates in the BEV map, represents the detection bounding box of the cooperative agent j at time t, represents the detection bounding box of the cooperative agent j at time t-1.
[0162] S4.5: Time selection process;
[0163] Next, the present invention evaluates which positions in the selected region have significant temporal dynamic changes compared to the previous time frame and retains only the elements at these positions. The temporal selection process aims to exclude regions with less dynamic changes because the targets in these regions can be tracked locally by the Kalman filter. Specifically, according to the change in the information entropy of the current time frame and the information entropy of the previous time frame, i.e., the feature dynamic change index Q j t judges whether the current time frame is a key frame based on whether it exceeds a threshold, and evaluates the positions in the spatially sampled spatial region that have significant temporal dynamic changes compared to the previous time frame, and retains the elements at these positions.
[0164] In the present invention, the temporal selection process can be expressed as:
[0165]
[0166] where represents the temporal selection matrix of the j-th user at the coordinate (h, w) at time t; represents an indicator function that takes the value of 1 when the condition is satisfied and 0 otherwise; represents the degree of temporal change; δ and τ represent the selection thresholds.
[0167] For the key frame, the present invention only applies spatial sampling without further temporal selection; while for the non-key frame, the present invention sequentially performs spatio-temporal sampling. By integrating these two cases, the binary selection matrix is defined as represents the spatial selection matrix, represents the temporal selection matrix, and ⊙ is the element-wise multiplication.
[0168] S4.6: Obtain the sparse feature map;
[0169] For the extracted perceptual features, the feature map is processed by compression to generate the sparse feature map and transmitted to other cooperative agents, while receiving the sparse feature maps of other cooperative agents. The information transmitted from the cooperative agent j to the autonomous vehicle i aims to enhance its perception ability. In addition, in these regions, the regions with rapid dynamic changes are given higher priorities. Therefore, the cooperative agent j uses the feature map filter to identify the regions beneficial to the autonomous vehicle i from its original feature map to obtain the sparse feature map This process can be expressed as:
[0170]
[0171] Among them, ⊙ represents element multiplication.
[0172] In the present invention, the feature map filter selects the feature map according to the binary selection matrix for the regions therein to obtain a sparse feature map The sparse feature map is characterized in that generally the information is sparsely distributed in space, but the key information density is high in the non-zero regions.
[0173] Step Five: Compression, encoding, and decoding of perceptual features;
[0174] For the compression and encoding of perceptual features, a method based on information importance is used to allocate the transmission rate R. At the sending end, the sparse feature map is first quantized and then converted into bits through an entropy encoder for transmission, thereby achieving optimal perception-communication performance.
[0175] Specifically, when autonomous vehicle i communicates, it assigns different transmission rates R to different regions according to the spatial information priority map and encodes the sparse feature map into bits for transmission to the other N - 1 cooperative intelligent agents. At the same time, it receives the sparse feature maps and the spatial information priority maps from the other N - 1 cooperative intelligent agents. It is assumed that the results of each round of detection are independent of each other, and the constructed communication graph is always a fully connected graph.
[0176] As Figure 3 shown, the specific implementation process of entropy encoding and decoding is as follows:
[0177] (a) Perform scalar quantization on the selected sparse feature map to generate a discrete vector
[0178]
[0179] where round(·) represents the operation of rounding each element to the nearest integer.
[0180] (b) Use an entropy encoding method (optionally, such as arithmetic coding) to perform lossless compression on the discrete vector through an entropy encoder;
[0181] (c) Use the probability mass function (PMF) to encode the discrete vector into a bit stream through an entropy encoder, and its transmission rate is
[0182] where represents the discrete vector The prior probability model is the entropy model, and ω represents the prior probability model parameters; through the prior probability model the PMF can be calculated, and the parameters of the PMF are known to both the entropy encoder and the entropy decoder.
[0183] (d) During the arithmetic decoding process, first, the discrete vector is entropy decoded through the bitstream, and then the decoded sparse feature map is input into the inter-vehicle feature fusion module.
[0184] Based on the entropy encoder, the present invention can more clearly define the loss function to be optimized. The present invention first reformulates maximizing the gain as minimizing the loss function of object detection. Then, a trade-off optimization is performed between the loss function L and the transmission rate R. Specifically, the loss function L is defined as:
[0185]
[0186] where L det (·) represents the detection loss between the predicted target and the ground truth, and L rate represents the loss function of the transmission rate R, represents the predicted target, represents the ground truth.
[0187] The present invention does not directly optimize the transmission rate of the sparse feature map but proposes an alternative loss function. Its core idea is to better utilize the spatio-temporal heterogeneity of the feature map by selecting the sparse feature map with the highest spatio-temporal gain. This method can more accurately model the probability distribution of the selected feature map, thereby reducing the number of bits actually transmitted.
[0188] Furthermore, the present invention will explain in detail the implementation process of the loss function L rate of the transmission rate R.
[0189] The above-mentioned quantized sparse feature map is the input of the entropy encoder, and the quantization is achieved by rounding each element in the sparse feature map to the nearest integer. However, the rounding function is non-differentiable, which makes it difficult to apply in end-to-end training. To solve this problem, the present invention replaces the rounding function with a differentiable additive uniform noise, as follows:
[0190]
[0191] where u represents the uniform distribution noise centered at 0.5 with a range of 1, represents the rounding output The approximation value. As shown in the literature [J.Balle, V.Laparra, and E.P.Simoncelli, “End-to-end Optimized Image Compression,” Mar. 2017.], this method has been proven to effectively solve the quantization problem.
[0192] Assume that each element in the sparse feature map is independently and identically distributed. In the present invention, P is used to represent the estimated distribution function of the sparse feature map . The transmission rate is equal to its expected coding length. Under the assumption that there is an efficient entropy encoder, the entropy of the discrete probability distribution H[P of the quantized discrete vector q guarantees the lower bound of the transmission rate of the quantization value q j→i of the semantic feature. Therefore, the transmission rate R of the sparse feature map is represented by the entropy calculated by the following formula:
[0193]
[0194] where, represents the estimated distribution function of the sparse feature map , represents the discrete vector generated by scalar quantization of the sparse feature map , and l represents the position of the element in .
[0195] The technique of using a binary mask matrix to select important regions has also been applied to neural image compression. The literature [C.Cai, L.Chen, X.Zhang, and Z.Gao, “End-to-End Optimized ROI Image Compression,” IEEE Transactions on Image Processing, vol. 29, pp. 3442–3457, 2020.] attempts to allocate a higher coding rate to the region of interest (ROI) rather than the background region, so as to retain information-rich regions and reduce redundancy. The ROI map is a binary mask matrix with the same dimension as the feature map, which is used to indicate whether a certain pixel is selected for entropy coding. However, these methods only use the binary mask matrix to select the foreground region and do not fully utilize the importance of individual pixels to guide the coding process. In the present invention, m j is defined as the quality map, which is used to represent the importance level of each element in the cooperative agent j, and its importance is measured by the spatial information priority map and the information entropy map as follows:
[0196]
[0197] It should be noted that for the quality map m j time selection is not considered because it is mainly used to generate a probability distribution for each value in the feature map. In this case, the probability distribution of the sparse feature map can be modeled only through spatial selection, and it guides bit allocation.
[0198] Furthermore, the present invention uses the above-mentioned feature map selection method to optimize the loss function of the transmission bit rate R. Given a discrete vector and the quality map m j , the loss function of the transmission bit rate R can be defined as follows:
[0199]
[0200] wherein, represents the occurrence probability of the selected feature map weighted by quality.
[0201] The bit rate loss is calculated based on the selected discrete vector . During the transmission process, only the selected features are encoded. Therefore, the calculation of the transmission bit rate during the training process only focuses on these selected values, so that the prior probability model of the entropy encoder can be established more precisely. The quality map m j represents re-weighting of these features, reflecting the spatial confidence difference between the selected feature values. When dynamically adjusting the bandwidth limit, the spatial confidence is sorted numerically, so that higher regions are more likely to be selected. As a weighting factor, the quality map m j assigns a higher bit rate loss to the regions more likely to be selected, thereby optimizing the prior probability model of the entropy encoder.
[0202] Step Six: Receiving and fusing perceptual features;
[0203] Furthermore, after the autonomous vehicle i receives information from the cooperative agent j, the autonomous vehicle i needs to integrate all available information to supplement the observation, so as to achieve a more comprehensive perception of the environment. At time frame t, the autonomous vehicle i enhances its observation by aggregating semantic information from other cooperative agents. In addition, considering that the historical information also contains rich details about the positions and motion patterns of the detected objects, the present invention designs a temporal fusion module to enhance the features of spatial fusion.
[0204] The semantic information fusion module is designed to help the autonomous vehicle i utilize the transmitted information from neighboring agents. The specific operation steps include two parts: entropy decoding and feature map fusion. Specifically, when the autonomous vehicle i receives the bitstream transmitted by the neighboring agent j, it uses an entropy decoder to decode the bitstream and extract the decompressed sparse feature map Subsequently, the decompressed sparse feature map With the local feature map F i t Perform feature fusion between vehicles. During the feature fusion process, the key point is to combine the detection confidence. Intuitively, the detection confidence can clearly indicate the importance of features, thus achieving element-wise feature fusion. The present invention uses the element-wise maximum fusion method to fuse the feature map F i t of the autonomous vehicle i and the sparse feature map of the agent j. The fused feature in the spatial dimension can be expressed as:
[0205]
[0206] where, represents the spatially aware enhanced feature map, and Φ spatial-fusion represents the spatial fusion function.
[0207] It should be noted that the present invention assumes that (the selected feature map) will also be transmitted to the autonomous vehicle i, and its size can be ignored compared to the feature map. Because the channel dimension of the feature map is C (usually much larger than 1), while the mask spatial information priority map has only a single channel.
[0208] Furthermore, once the spatially aware enhanced feature map is obtained, it is also necessary to further fuse in the time dimension. This process enables the autonomous vehicle to utilize the historical perception feature sequence and the optical flow map of the current scene to map the past features to the current time frame, thereby facilitating subsequent detection and tracking tasks. Specifically, the autonomous vehicle i first uses a local Kalman filter to track the detected targets and obtain the motion information of each bounding box, including speed and direction. The specific calculation formula is:
[0209]
[0210] where, represents the detection bounding box of the autonomous vehicle i at time t, represents the detection bounding box of the autonomous vehicle i at time t - 1; represents and the difference.
[0211] The motion vector of the bounding box is used to generate an optical flow map in the BEV coordinate system, which represents the motion vector at each position in the scene:
[0212]
[0213] where flow represents the optical flow map, and Φ map represents the conversion from the optical flow map to the bird's-eye view.
[0214] In the optical flow map, each element at position (h, w) is a velocity vector, denoted as (v x , v y ), representing the projection of the velocity vector in the BEV coordinate system.
[0215] Then, the fused feature map of the previous time frame is propagated to the current time frame through element-wise mapping, and denotes the temporal fused feature map containing historical information, and its specific calculation formula is:
[0216]
[0217] Finally, denotes the temporal fused feature map at position (h, w). The temporal fused feature map containing historical information is fused with the spatially aware enhanced feature map of the current time frame to obtain the spatio-temporal fused feature map of the current time frame:
[0218]
[0219] where the max operation is used to select the maximum value from the feature maps of neighboring agents and ego vehicle i at each position.
[0220] Step Seven: Object Detection and Tracking;
[0221] The object detection decoder is responsible for converting the spatio-temporal fused feature map into a 3D object detection result. This process can be expressed as:
[0222]
[0223] where Φ update represents the object detection decoder, represents the 3D object detection result.
[0224] The information of each object contains 7 dimensions, namely classification confidence, position, size, and angle, used to characterize the information of the detection bounding box at that position. In each time frame, a tracker is initialized for each newly detected object, and for the occluded or disappeared objects, their corresponding trackers are terminated.
[0225] The tracker associates the 3D object detection results of each time frame to form a trajectory for each detection bounding box. Based on the historical trajectory, the Kalman filter is used to predict the position of the detection bounding box in the current time frame. Among them, at time frame t, the detection bounding box trajectory predicted by the Kalman filter is a multi-dimensional variable, including the velocity, position, and orientation of the object:
[0226]
[0227] Among them, φ predict represents the tracker, represents the detection bounding box of the autonomous vehicle i at time t, represents the detection bounding box of the autonomous vehicle i at time t-1.
[0228] Subsequently, the 3D object detection result is optimized by the following formula:
[0229]
[0230] Among them, f box fusion represents the fusion of the detection bounding boxes, represents the 3D object detection output result.
[0231] The Kalman filter uses the final 3D object detection output result to update its trajectory:
[0232]
[0233] Among them, Φ update represents the update of the Kalman filter.
[0234] In the present invention, the tracker associates the target tracked by the Kalman filter with the current object detection result through an intersection over union (IOU) threshold. Once the tracking box is associated with the detection bounding box, the present invention uses the 3D IOU value as the matching score. The tracker can provide a tracking box and a tracking score s t , representing the matching degree between the current box and the detected target and the confidence level of the prediction result respectively.
[0235] Next, the present invention proposes a simple and efficient optimization module for optimizing the detection bounding box b through the tracking box t . Since the Kalman filter prediction may introduce errors, the tracking score s t is multiplied by a decay factor γ. The coordinates of the detection bounding box b t are obtained by weighted calculation from the tracking score s t and the detection score c t , and its specific calculation formula is as follows:
[0236]
[0237] Step Eight: Training of the collaborative perception framework based on feature extraction;
[0238] To train the entire system framework, the present invention utilizes the optimization formula Optimize the loss function. Among them, the parameters to be optimized include the object detection decoder, the feature map filter (sharing parameters with the object detection decoder), and the prior probability model of the entropy encoder. When trained with sufficient labeled data and computing resources, this collaborative perception framework can achieve efficient end-to-end optimization. However, due to task complexity and the risk of unstable training process, the training is divided into two stages.
[0239] In the first-stage training, the present invention first optimizes the parameters of the feature map filter. This is because the entropy encoder depends on the sparse feature map generated by the feature map filter, and preferentially optimizing the feature map filter can ensure that the dependencies between components are handled. In the first-stage training, the present invention sets lower thresholds for the selection function S select (·) and τ to filter out most of the background regions that do not contain objects, while retaining all features that contribute to collaborative perception.
[0240] In the second-stage training, using the discrete vector and the guidance of the quality map m j the loss function of the transmission code rate R can be calculated Subsequently, calculate the loss function of the optimization problem Perform end-to-end training on the entire collaborative perception framework including the prior probability model of the entropy encoder through the above optimization formula.
[0241] Some existing collaborative perception methods mainly attempt to minimize the consumption of communication bandwidth resources by these information during the communication process on the premise of processing and compressing the information in each spatial region without distinction. Although many existing works have effectively improved the performance of the collaborative perception model in the scenario of limited communication bandwidth, they all ignore the characteristic that the density of valuable information in different spatial regions is different, thus wasting some communication bandwidth resources.
[0242] In view of this, the present invention aims to propose a collaborative 3D object detection and tracking method based on deep learning, which can screen out the information in key spatial regions for transmission and efficiently fuse the information from other collaborative agents.
[0243] In order to verify the effectiveness of a collaborative 3D object detection and tracking method based on deep learning of the present invention, the following verification and comparison experiments were carried out.
[0244] 1. Evaluate the perception performance of the method of the present invention on the OPV2V test dataset;
[0245] In the present invention, by adjusting the threshold in the spatial information priority map, different numbers of elements in the feature map are selected to adapt to different communication bandwidths. Figure 4Shows the comparison of the trade - off between sensing performance and communication bandwidth for different collaborative algorithms. In the case of transmitting the complete feature map, the method of the present invention shows higher performance among the intermediate fusion methods with fixed communication volume, such as v2x - vit and v2vNet. Compared with the early fusion method, even though there is a loss of information accuracy due to feature selection and quantization, the method of the present invention has improved by 0.8% and 1.3% respectively in terms of AP50 and AP70 metrics, indicating that the method of the present invention can effectively extract semantic features from the original data and further improve the detection accuracy through the temporal fusion process.
[0246] In the case of achieving the same sensing accuracy as the early fusion method, the amount of data to be transmitted by the method of the present invention is reduced by about 758 times. This shows that the feature map filter still retains rich feature information after filtering out elements with lower transmission priority. In addition, by quantizing the transmitted features from float32 type to integer type, although there is a loss of precision, the impact on the final sensing performance is small, further verifying the effectiveness of the present invention.
[0247] 2. Robustness verification of positioning error;
[0248] In practical applications, due to the accuracy limitation of spatial observation, there may be positioning errors. To evaluate the robustness of the method of the present invention to positioning errors, different degrees of Gaussian noise positioning errors are simulated on the OPV2V dataset. Specifically, positioning noise sampled from a Gaussian distribution is added to the test data, and the standard deviations of the position error and the orientation error are σ pos ∈[0m, 0.5m] and σ heading ∈[0°, 0.5°].
[0249] As Figure 5 shown, the present invention uses the no - fusion method as the baseline model. Since the baseline model only uses its own local sensing data and is not affected by noise, its performance is shown as a horizontal line in the figure. Under ideal conditions, the performance of the no - fusion method is the worst. However, in the presence of positioning errors, due to the inaccuracy of the shared information, the performance of all collaborative schemes decreases, and when the error is large, it is even lower than the performance of the baseline model.
[0250] Furthermore, the present invention is compared with other intermediate fusion methods in terms of robustness. As Figure 5 shown, as the error increases, the overall performance of collaborative sensing shows a downward trend. When the standard deviations of the position and orientation errors are maintained within the normal range (σ pos ≤[0, 0.2] or σ heading≤[0°, 0.5°]), the performance of the method of the present invention decreases, but it is still higher than the baseline model. In addition, the perception performance of other intermediate fusion methods (v2vnet, v2x-vit, proposed, where2comm) drops below that of the non-fusion method when the error increases. In the case of large errors, the method of the present invention still maintains a detection accuracy of about 80%, while the performance of other methods drops significantly, further verifying the robustness of the method of the present invention to positioning errors.
[0251] The above are only examples of this specification and are not used to limit this specification. For those skilled in the art, this specification can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification. In addition, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.
Claims
1. A collaborative 3D object detection and tracking method based on deep learning, characterized in that, It includes the following steps: Step 1: Construct a collaborative perception framework based on feature extraction; The collaborative perception framework includes: a feature encoder, a feature map filter, an entropy encoder, a semantic information fusion module, an object detection decoder, and a tracker; Step 2: Construct an optimization problem; Assume that at time frame t, there are a total of N vehicles, 1 autonomous vehicle i, and N - 1 cooperative agents; The objective function of the perception task is defined as: Among them, T represents the number of frames within a period of time, represents the sparse feature map, represents the set of collaborators of the i-th collaborative agent at time t, φ θ represents the neural network Φ with parameter θ, f eva (·) represents the evaluation function of collaborative perception, and the max operation is used to maximize the gain G, and Y i t respectively represent the observation value and the true value of the autonomous vehicle i at time frame t; The objective function of the collaborative perception framework is defined as: Among them, λ is used to balance the relationship between the two items of gain and feature coding rate, E represents taking the expectation, and q j->i represents the quantization value of the semantic feature, represents the probability of the quantization value of the semantic feature occurring; Step 3: Perception feature extraction; The collaborative agent j uses a feature encoder to extract features from the raw sensor data to obtain a feature map The autonomous vehicle i uses a feature encoder to extract features from the raw sensor data to obtain a feature map Step 4: Transmission feature screening; The feature map is converted into a spatial information priority map by a spatial information prioritizer in the feature map filter and the dynamic change index of the current time frame is calculated; by time selection, regions with less dynamic change are excluded, and the cooperative agent j uses the feature map filter to identify regions beneficial to the autonomous vehicle i from its original feature map to obtain a sparse feature map from its original feature map Step 5: Compression, encoding, and decoding of perception features; The sender first quantizes the sparse feature map and then converts it into a bitstream through an entropy encoder for transmission; Step 6: Reception and fusion of perception features; The autonomous vehicle i receives the bitstream transmitted by the neighboring agent j, decodes the bitstream using an entropy decoder, and extracts the decompressed sparse feature map. The decompressed sparse feature map is fused with the local feature map F i t to perform inter-vehicle feature fusion and obtain a spatially aware enhanced feature map. The spatially aware enhanced feature map is fused with the temporal fusion feature map containing historical information to perform fusion in the temporal dimension and obtain the spatio-temporal fusion feature map of the current time frame. Step 7: Object detection and tracking; Convert the spatio-temporal fusion feature map through the target detection decoder into 3D target detection results: Among them, Φ decoder represents the target detection decoder, represents the 3D target detection result, that is, the ground truth of the autonomous vehicle i at time frame t; Initialize a tracker for each newly detected object in each time frame. The tracker associates the 3D object detection results of each time frame to form a trajectory for each detection bounding box; based on the historical trajectory, use the Kalman filter to predict the position of the detection bounding box in the current time frame; the trajectory of the detection bounding box is a multi-dimensional variable, including the speed, position, and orientation of the object: Among them, φ predict represents the tracker, represents the detection bounding box of autonomous vehicle i at time t, represents the detection bounding box of autonomous vehicle i at time t-1; Optimize the 3D object detection results using the trajectory of the detection bounding box and the following formula: Among them, f box fusion represents the fusion of the detection bounding boxes, and represents the final 3D object detection output result; Using the Kalman filter to use the final 3D object detection output result Update the trajectory of the detection bounding box: Among them, Φ update represents the update of the Kalman filter.
2. The method for collaborative 3D object detection and tracking based on deep learning according to claim 1, characterized in that The collaborative perception framework accepts single-modal input and multi-modal input. For RGB image input, after passing through the feature encoder, a transformation function that converts the front view to BEV is used to obtain a feature map from the BEV perspective. For 3D point cloud data input, it is discretely converted to BEV, and then through the feature encoder, a BEV perspective feature map is obtained from the 3D point cloud data using the anchor-based PointPillar method.
3. A collaborative 3D object detection and tracking method based on deep learning according to claim 1, characterized in that, The specific implementation process of Step 4 is as follows: S4.1: Obtain the spatial information priority map; The spatial information priority generator consists of the target decoder Φ priority-decoder which consists of the target decoder Φ priority-decoder to generate a spatial information priority map S4.2: Information entropy calculation; Each element in the spatial information priority diagram represents the probability of containing the target. The target detection problem is regarded as a binary classification task, where each category represents whether there is a target at the current position. According to the definition of information theory, use to represent the information entropy diagram of the j-th collaborative agent, which represents the information entropy diagram of the j-th collaborative agent The value at the position (h, w) is: where the loop variable k represents the category, represents the information priority graph of the k-th category at the coordinate (h, w); at each coordinate (h, w), it satisfies: P (h,w,0) +P (h,w,1) = 1 Estimate the dynamic change level of each position by comparing the information entropy differences between adjacent two-frame feature maps; S4.3: Spatial sampling; Select the spatial regions worth transmitting from the spatial information priority map through spatial sampling operations and generate a binary selection matrix, so as to only transmit the spatial regions worth transmitting and omit the background regions and static regions that are not worth transmitting; S4.4: Dynamic characteristic evaluation; Characteristic dynamic change index Expressed as: Among them, represents the information entropy graph of the j-th collaborative agent at time t-1; For cooperative agent j, its dynamic change index is defined as: where, Φ map (·) is used to convert the difference of the trajectory into the corresponding coordinates in the BEV map, represents the detection bounding box of cooperative agent j at time t, ) represents the detection bounding box of cooperative agent j at time t-1; S4.5: Time selection process; Dynamically varying metrics based on features Determine whether the current time frame is a key frame by checking if a feature exceeds a threshold, evaluate positions in the spatially sampled spatial region that exhibit significant temporal dynamic changes compared to the previous time frame, and retain the elements at these positions; For key frames, only spatial sampling is applied without temporal selection; for non-key frames, spatial-temporal sampling is performed sequentially; by integrating these two cases, the binary selection matrix is defined as denotes the spatial selection matrix, denotes the temporal selection matrix, and ⊙ is the element-wise multiplication; S4.6: Obtain the sparse feature map; Select feature maps according to a binary selection matrix through a feature map filter Select regions in the feature maps to obtain sparse feature maps where ⊙ is element-wise multiplication.
4. A collaborative 3D object detection and tracking method based on deep learning according to claim 3, characterized in that, In Step S4.3, the process of the spatial sampling is expressed as: Among them, represents a binary selection matrix, and the area with a value of 1 in the binary selection matrix is the spatial area worthy of transmission, and the area with a value of 0 is the spatial area not worthy of transmission; S select (·) is a selection function, which preferentially sets the spatial area worthy of transmission to 1 and the spatial area not worthy of transmission to 0 until the information of the spatial area worthy of transmission reaches the critical value of the limited communication volume V.
5. The method for collaborative 3D object detection and tracking based on deep learning according to claim 3, characterized in that In Step S4.5, the time selection process can be expressed as: Among them, represents the time selection matrix of the j-th user at the coordinate (h, w) at the t-th moment; represents an indicator function that takes the value of 1 when the condition holds, and 0 otherwise; represents the degree of temporal variation; δ and τ represent selection thresholds.
6. The collaborative 3D object detection and tracking method based on deep learning according to claim 1, characterized in that, The specific implementation process of Step 5 is as follows: (a) Perform scalar quantization on the selected sparse feature map to generate a discrete vector where round(·) represents the operation of rounding each element to the closest integer; (b) Using the arithmetic coding method, the discrete vector is losslessly compressed by an entropy encoder ; (c) Using a probability mass function, encode the discrete vector through an entropy encoder into a bitstream with a transmission code rate of representing the prior probability model, i.e., the entropy model, of the discrete vector where ω represents the prior probability model parameters; the probability mass function can be calculated through the prior probability model ; (d) During the arithmetic decoding process, first perform entropy decoding on the discrete vector and then input the decoded sparse feature map into the inter-vehicle feature fusion module.
7. A collaborative 3D object detection and tracking method based on deep learning according to claim 1, characterized in that, The sparse feature map The calculation formula for the transmission code rate R is as follows: Among them, represents the estimated distribution function of the sparse feature map , represents the discrete vector generated by scalar quantization of the sparse feature map , and l represents the position of the element in The loss function of the optimization problem is: Among them, L det (·) represents the detection loss between the predicted target and the true value, L rate represents the loss function of the transmission code rate R, represents the predicted target, represents the true value; The loss function L rate has the following calculation formula: Among them, represents the occurrence probability of the selected feature map weighted by quality; m j is the quality map, which is used to represent the importance level of each element in the collaborative agent j, and its importance is measured by the spatial information priority map and the information entropy map as follows:
8. A collaborative 3D object detection and tracking method based on deep learning according to claim 1, characterized in that The feature map with enhanced spatial perception has the following mathematical expression: Among them, Φ spatial-fusion represents a spatial fusion function; The spatio-temporal fusion feature map has the following mathematical expression: Among them, the max operation is used to select the maximum value from the feature maps of neighboring agents and the autonomous vehicle i at each position; represents the temporal fusion feature map containing historical information, and its mathematical expression is: Among them, (h, w) represents the position coordinates in the optical flow diagram, and (v x , v y ) represents the velocity vector corresponding to each element at the position (h, w), represents the fused feature map of the previous time frame.
9. A collaborative 3D object detection and tracking method based on deep learning according to claim 1, characterized in that, In step seven, the tracker associates the target tracked by the Kalman filter with the current target detection result through an intersection over union threshold; after the tracking box is associated with the detection bounding box, the 3D IOU value is used as the matching score; through the tracking box optimizes the detection bounding box b t and the coordinates of the detection bounding box b t are obtained by weighted calculation from the tracking score s t and the detection score c t The specific calculation formula is as follows: where γ represents the decay factor.
10. A collaborative 3D object detection and tracking method based on deep learning according to claim 1, characterized in that, It also includes the training steps of the collaborative perception framework based on feature extraction; In the first stage, optimize the parameters of the feature map filter; For the selection function S select (·) and τ are set with lower thresholds to filter out most of the background regions that do not contain the target while retaining all the features helpful for cooperative perception; In the second stage, the discrete vector and the quality map m j are used to guide the calculation of the loss function of the transmission code rate R and then the loss function of the optimization problem is calculated; the entire cooperative perception framework including the prior probability model of the entropy encoder is trained end-to-end through the optimization formula , where L det (·) represents the detection loss between the predicted target and the true value, and L rate represents the loss function of the transmission code rate R represents the predicted target represents the true value
Citation Information
Cited By
Multi-agent collaborative sensing method and system based on spatio-temporal joint selection
CN121982288A
A multi-agent collaborative perception method and system based on spatio-temporal joint selection
CN121982288B