Aerial view vehicle instance and lane topology prediction method and system
Patent Information
- Application Number
- CN202611193978.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0009]本发明用于解决以下技术问题:动态车辆预测和静态道路拓扑解析分别设置前端感知网络而产生重复处理;本车运动导致历史帧特征空间错位;原始鸟瞰特征对车辆运动区域的响应不足;随机初始化和无约束坐标更新导致车道几何预测不稳定;对不同道路元素采用相同拓扑信息传播强度容易引入拓扑噪声
[0025]采用上述技术方案具有以下优点:
Smart Images

Figure CN122821187A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving vehicle-mounted visual perception technology, and in particular to a method and system for predicting vehicle instances and lane topology from a bird's-eye view. Background Technology
[0002] With the development of autonomous driving technology, vehicles need to identify road structures, lane connectivity, and surrounding dynamic targets based on environmental images captured by camera equipment. Bird's-eye view representation can convert image information from different perspectives into a unified overhead space, facilitating vehicle road structure analysis and target motion judgment within the same spatial coordinate frame, and is therefore being applied to in-vehicle visual perception systems.
[0003] Chinese invention patent application CN117037095A, published on November 10, 2023, discloses a high-precision map road topology prediction scheme based on Transformer. This scheme acquires image or point cloud perception data from multiple directions around a vehicle, generates bird's-eye view features through feature extraction and feature fusion, identifies lane centerlines and traffic lights, and then predicts the topological relationships between lane centerlines and between lane centerlines and traffic lights. This scheme primarily processes static road elements and their connectivity relationships.
[0004] Chinese invention patent application CN120526426A, published on August 22, 2025, discloses a scene topology understanding scheme. This scheme generates bird's-eye view features using multi-view environmental images from the same moment. Based on vehicle motion parameters, it spatiotemporally aligns and fuses the bird's-eye view features of the current frame with those of historical frames. It then corrects the bird's-eye view features using a map prior model and generates a lane topology map through a topology decoder.
[0005] The above scheme can achieve road topology prediction or multi-frame road feature fusion, but it does not simultaneously generate multi-frame future vehicle instance prediction results based on a unified and shared bird's-eye view temporal representation. With separate dynamic vehicle prediction networks and static road topology networks, the two tasks may perform image feature extraction and bird's-eye view spatial mapping respectively, resulting in redundant processing.
[0006] Changes in continuous images include both those caused by the movement of other vehicles and spatial displacements caused by the movement of the vehicle itself. If historical features are not aligned with uniform coordinates, inter-frame change information may be mixed with spurious differences caused by the vehicle's displacement. Processing temporal features in a frame-by-frame loop also limits the parallel processing of features from different times. Directly using un-enhanced bird's-eye view features also makes it difficult to specifically highlight areas of vehicle movement.
[0007] Regarding lane topology parsing, when initializing lane queries using a random method, lane polylines lack road geometric references in the early stages of iteration, potentially leading to significant offsets during coordinate updates. Applying the same topology information propagation intensity to all road elements can cause relatively isolated road elements, such as zebra crossings, to participate in unnecessary connectivity aggregation, thus interfering with the lane topology results.
[0008] Therefore, there is a need for a prediction method and system that can process dynamic vehicle temporal information and static road topology information in a unified bird's-eye view space, and improve the spatial consistency of historical features, the ability to express vehicle motion features, and the stability of lane geometry and connectivity. Summary of the Invention
[0009] This invention addresses the following technical problems: redundant processing occurs when dynamic vehicle prediction and static road topology analysis are performed using separate front-end perception networks; vehicle movement causes spatial misalignment of historical frame features; insufficient response of original bird's-eye view features to vehicle movement areas; unstable lane geometry prediction due to random initialization and unconstrained coordinate updates; and topological noise is easily introduced when the same topological information propagation intensity is used for different road elements.
[0010] To achieve the above objectives, this invention proposes a method for predicting vehicle instances and lane topology from a bird's-eye view, comprising: The system acquires continuous temporal images of the vehicle from multiple perspectives and the vehicle's positioning attitude information. It then performs multi-scale feature extraction and fusion on the images from each perspective. Based on the vehicle's positioning attitude information, it determines the coordinate transformation matrix of the vehicle coordinate system from the historical frame to the current time and uses the coordinate transformation matrix to align the features of the fused historical frame images. Spatial self-attention processing is performed on the learnable bird's-eye view query vector, and cross-attention information interaction is performed between the processed bird's-eye view query vector and the aligned fused image features to generate original bird's-eye view features. The original bird's-eye view features are then refined through a sparse coding network to obtain continuous frame bird's-eye view features. Weighted pixel-by-pixel difference calculation is performed on the bird's-eye view features of adjacent frames located in the same bird's-eye view coordinate system. Temporal attention operation, spatial attention operation and gated linear processing are performed sequentially through a non-cyclic gated transform encoder to obtain temporally enhanced shared bird's-eye view features. The bird's-eye view space is divided into grid partitions. The number of lane prototypes is allocated according to the lane sample density in each grid partition. Lane polyline prototypes are generated by clustering in each grid partition and lane queries are initialized with the lane polyline prototypes. The temporal enhanced shared bird's-eye view features are input into the lane topology decoding branch and the vehicle instance prediction branch, respectively. The lane topology decoding branch updates the lane polyline coordinates layer by layer based on the coordinate correction amount compressed by the hyperbolic tangent function and scaled by the coordinate scaling factor. It generates a topology gating value based on the query features of each lane, and uses the topology gating value to weight the graph convolutional connectivity features obtained based on the soft connectivity adjacency matrix between lanes to obtain the lane geometry and topology prediction results. The vehicle instance prediction branch generates semantic segmentation maps and reverse motion flow graphs for each future frame, and transmits vehicle instance identifiers based on the reverse motion flow graph to obtain multi-frame future vehicle instance prediction results.
[0011] Preferably, the multi-scale feature extraction and fusion includes: determining the corresponding scale alignment coefficient based on the downsampling ratio of each feature level, upsampling the corresponding feature layer using the scale alignment coefficient, and fusing the upsampled feature layers sequentially through a convolution unit, a normalization unit, and an activation unit to obtain multi-scale fused image features; The spatial alignment relationship of frame history image features satisfies ,in, For the first Spatially aligned fused image features For the first Multi-scale fused image features before frame space alignment For the first The coordinate transformation matrix from the frame to the vehicle coordinate system at the current moment.
[0012] Preferably, the bird's-eye view query vector after spatial self-attention processing is interacted with the aligned multi-scale fused image features through cross-attention information exchange to obtain the original bird's-eye view features, wherein the original bird's-eye view features satisfy... ,in, The original bird's-eye view features, For the learnable bird's-eye view query vector, For the aligned multi-scale fused image features, For cross-attention information aggregation operations, the sparse coding network performs multi-scale spatial refinement of the original bird's-eye view features through step-by-step downsampling, step-by-step upsampling, and skip connections between corresponding levels.
[0013] Preferably, the difference enhancement feature obtained by the weighted pixel-by-pixel difference operation satisfies ,in, For the difference enhancement feature, Increase the weighting coefficient for the region of change. A detailed aerial view of the current moment. The bird's-eye view features are refined from the previous time step; the temporally enhanced shared bird's-eye view features satisfy... ,in, For the time-series enhanced shared bird's-eye view features, It is a loop-free gated spatiotemporal coding operation consisting of a temporal attention unit, a spatial attention unit, and a gated linear unit, wherein the temporal attention unit is executed before the spatial attention unit.
[0014] Preferably, the bird's-eye view space is uniformly divided into grid partitions with a preset number of rows and columns, and the density allocation weight is determined according to the number of lane samples in each grid partition. The density allocation weights of each grid partition satisfy... ,in, For the first Density weights are assigned to each grid partition. For the first The number of lane samples within each grid partition; the number of lane prototypes in each grid partition is allocated according to the density allocation weight, and the lane samples are clustered within each grid partition to obtain the lane polyline prototype.
[0015] Preferably, the normalized coordinates of the lane polyline prototype are used as the initial lane polyline coordinates. The lane polyline coordinates are based on Perform iterative updates layer by layer, where... For the lane topology decoding branch number Normalized coordinates of lane polylines after layer iteration For the first Normalized coordinates of lane polylines after layer iteration This is the coordinate scaling factor. For the first The original coordinate correction amount for layer prediction.
[0016] Preferably, the first The topology gating value corresponding to the lane group query satisfies ,in, For the first The topology gating value corresponding to the lane group query. For the gated weight vector, For the first Lane query features corresponding to group lane queries This is the gating bias. A nonlinear gating function is used to map the linear calculation results to topological gating values; based on the lane query features transformed by the multilayer perceptron and the soft connectivity adjacency matrix between lanes, graph convolutional connectivity features are obtained, and the first... The topological gating value corresponding to the lane query group is used as the weighting coefficient of the graph convolutional connectivity feature corresponding to the lane query group, and the weighted graph convolutional connectivity feature is fused with the corresponding original lane query feature and multilayer perceptron transform feature.
[0017] Preferably, the vehicle instance prediction branch includes a multi-layer coding structure and two parallel output channels. The multi-layer coding structure generates feature representations corresponding to different spatial scales based on the temporal enhanced shared bird's-eye view features, and the two parallel output channels respectively generate the first... Semantic segmentation graphs and reverse motion flow graphs for each future time step, satisfying... ,in, For the first A semantic segmentation map for a future moment, wherein the semantic segmentation map represents the area occupied by vehicles in a bird's-eye view space. For the first A reverse motion flow graph for each future time step, wherein the reverse motion flow graph records the spatial offset of the corresponding vehicle instance at the previous time step. This is for multi-scale instance prediction computation.
[0018] Preferably, before performing the prediction, the network used to perform the bird's-eye view vehicle instance and lane topology prediction method is further trained in two stages: In the first stage, only the image feature extraction module and the bird's-eye view spatial mapping module are iterated, and the temporal enhancement module, lane topology decoding branch and vehicle instance prediction branch are not started, and the current frame bird's-eye view segmentation task is used as the training target; In the second stage, the parameters of the image feature extraction module and the bird's-eye view spatial mapping module trained in the first stage are fixed, and the parameters of the temporal enhancement module, lane topology decoding branch and vehicle instance prediction branch are updated.
[0019] Preferably, the second stage employs a multi-task loss method including lane topology regression loss and vehicle instance prediction branch dynamic weighting loss, wherein the vehicle instance prediction branch dynamic weighting loss satisfies ,in, Predict the dynamic weighted loss of the branch for the vehicle instance. The total number of future frames to be predicted. For future frame numbers, For semantic segmentation loss, For reverse flow loss, For pixel offset loss, For target centrality loss, , , and These are the dynamic weights corresponding to the loss, and the dynamic weights are adaptively adjusted with each training round.
[0020] Preferably, the method further includes: aligning the road topology vector information in the lane geometry and topology prediction results with the multi-frame future vehicle instance prediction results, and encapsulating them into a scene-aware data packet according to a unified data format; sending the scene-aware data packet to the vehicle's planning and control module, wherein the planning and control module determines the passable area based on the road topology vector information, and generates at least one of following control information, avoidance control information, and intersection turning control information based on the multi-frame future vehicle instance prediction results.
[0021] Preferably, it further includes: continuously acquiring synchronous multi-view original images and the actual vehicle motion trajectories at each time point during the model running phase, and obtaining road topology ground truth annotations corresponding to the synchronous multi-view original images to form an incremental training sample set; according to Calculate the cumulative loss corresponding to the incremental training sample set, where, The cumulative loss corresponding to the incremental training sample set, For a single set of incremental training samples, The incremental training sample set. The multi-task comprehensive loss is calculated for a single set of incremental training samples. Based on the accumulated loss, the parameters of the image feature extraction module, temporal enhancement module, lane topology decoding branch and vehicle instance prediction branch are simultaneously fine-tuned, and the network weights after parameter fine-tuning are saved.
[0022] This application also discloses a bird's-eye view vehicle instance and lane topology prediction system, including a multi-view image acquisition part, a vehicle positioning attitude information acquisition part, and a computing device; the multi-view image acquisition part includes multiple camera devices arranged around the vehicle for acquiring continuous temporal images of the vehicle from multiple perspectives; the vehicle positioning attitude information acquisition part is used to acquire vehicle positioning attitude information corresponding to each acquisition time; the computing device includes a processor and a memory, the memory storing an image feature extraction program, a bird's-eye view spatial mapping program, a temporal enhancement program, a lane topology decoding program, a vehicle instance prediction program, and corresponding network parameters, the processor calling the program and corresponding network parameters to execute the method described in claim 1.
[0023] Preferably, the system further includes a planning and control module that is communicatively connected to the computing device. The computing device is used to perform format alignment on the road topology vector information and the prediction results of multiple future vehicle instances, and encapsulate them into scene-aware data packets according to a unified data format before sending them to the planning and control module. The planning and control module is used to determine the passable area based on the road topology vector information, and generate at least one of following control information, avoidance control information, and intersection turning control information based on the prediction results of multiple future vehicle instances.
[0024] Preferably, the memory further stores an incremental training program, which the processor executes to form an incremental training sample set based on the synchronous multi-view original images, the corresponding road topology ground truth annotations, and the actual vehicle motion trajectories at each time point, according to... The cumulative loss corresponding to the incremental training sample set is calculated, and the parameters of the image feature extraction module, temporal enhancement module, lane topology decoding branch, and vehicle instance prediction branch are simultaneously fine-tuned based on the cumulative loss. The network weights with fine-tuned parameters are stored in the memory. The cumulative loss corresponding to the incremental training sample set, For a single set of incremental training samples, The incremental training sample set. This is the multi-task integrated loss calculated for a single set of incremental training samples.
[0025] The above technical solution has the following advantages: First, by aligning the features of the fused historical frames to the current vehicle coordinate system based on the vehicle's positioning and attitude information, the spatial misalignment between frames caused by the vehicle's movement can be reduced, allowing subsequent difference features to better reflect changes in the actual scene.
[0026] Second, the bird's-eye view features of adjacent frames after alignment are enhanced by difference and non-cyclic gated spatiotemporal coding, which can reduce the interference of static scene content on dynamic target modeling and highlight the temporal changes caused by vehicle movement.
[0027] Third, the shared image feature extraction results, bird's-eye view spatial mapping results, and time-series enhancement bird's-eye view features of lane topology decoding branches and vehicle instance prediction branches can reduce the redundant calculations caused by performing front-end feature processing separately for the two tasks.
[0028] Fourth, the lane polyline prototype provides a geometric initialization basis for lane queries. Hyperbolic tangent compression and coordinate scaling coefficients can limit the magnitude of single-layer coordinate correction. Query-by-query topology gating can adjust the intensity of topological information propagation of different road elements, thereby improving the stability of lane geometry and connectivity prediction.
[0029] Fifth, the semantic segmentation graph is used to determine the area occupied by the vehicle, and the reverse motion flow graph is used to establish the instance correspondence between adjacent prediction times. The two work together to generate multi-frame future vehicle instance prediction results with cross-frame instance identifiers.
[0030] Sixth, the two-stage training separates the basic bird's-eye view spatial mapping from the temporal, topological, and future instance predictions, which helps reduce mutual interference caused by multiple tasks simultaneously changing shared front-end parameters.
[0031] Seventh, road topology vector information and multi-frame future vehicle instance prediction results can be provided to the planning and control module in a unified format, enabling the planning and control module to simultaneously utilize static road structure and dynamic vehicle changes for control processing. Attached Figure Description
[0032] The present invention will now be described in detail with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a flowchart illustrating the bird's-eye view vehicle instance and lane topology prediction method provided in an embodiment of the present invention. Detailed Implementation
[0033] The following is combined Figure 1 The technical solution of the present invention will be further described below. The embodiments described are used to illustrate the technical concept and implementation process of the present invention and are not intended to unduly limit the present invention. Where there is no technical conflict, the technical features of the various embodiments can be combined with each other.
[0034] The execution entity of this embodiment can be a computing device installed on an autonomous vehicle. The computing device includes a processor and a memory, the memory storing programs and network parameters for executing each processing step. The processor calls the program to process images acquired by the onboard camera and vehicle positioning and attitude information. The computing device can interact with multiple camera devices, the vehicle positioning and attitude information output terminal, and the vehicle planning and control module.
[0035] Example 1 like Figure 1 As shown, this embodiment provides a method for predicting vehicle instances and lane topology from a bird's-eye view. This method sequentially performs multi-view image feature extraction and alignment, bird's-eye view spatial mapping, temporal feature enhancement, geometrically constrained lane topology decoding, and future vehicle instance prediction. The lane topology decoding branch and the vehicle instance prediction branch share the image feature extraction results, bird's-eye view spatial mapping results, and temporal enhancement shared bird's-eye view features.
[0036] The autonomous vehicle is equipped with multiple cameras positioned around it. Each camera continuously captures sequential images of the vehicle's surroundings, and the images from all cameras at the same time are combined into a multi-view image set. The images may include scene elements such as vehicles, lane markings, road boundaries, intersections, and traffic facilities.
[0037] During image acquisition, the vehicle's positioning and attitude information at each acquisition moment is obtained synchronously. This positioning and attitude information describes the vehicle's position and attitude changes between consecutive moments. Multi-view images at the same moment correspond to the vehicle's positioning and attitude information at that moment.
[0038] The computing device processes images from different viewpoints separately through an image feature extraction module. This module can employ a visual transformation network capable of outputting features at multiple spatial scales. Different feature levels represent image content at different spatial scales.
[0039] The computing device determines the corresponding scale alignment coefficient or scale alignment processing based on the downsampling ratio of each feature level, and upsamples the corresponding feature layers to give different feature layers a fusionable spatial size. After scale alignment is completed, each feature layer is fused sequentially through convolutional units, normalization units, and activation units.
[0040] For the The multi-scale fusion relationship of a frame image is represented as follows:
[0041] in, Indicates the first Multi-scale fusion image features of frames. Indicates the first The first frame Each feature layer Indicates the first The scale alignment coefficients or scale alignment processing corresponding to each feature layer This indicates the total number of feature levels.
[0042] Because the vehicle moves between consecutive acquisition moments, the image features from historical moments and the image features from the current moment are in different local reference coordinate systems of the vehicle. The computing device determines the coordinate transformation matrix from the historical frame to the vehicle coordinate system at the current moment based on the vehicle's positioning and attitude information at each moment.
[0043] For the The spatial alignment relationship of the frame history image features is expressed as follows:
[0044] in, Indicates the first Spatially aligned fused image features Indicates the first Multi-scale fused image features before frame space alignment Indicates the first The coordinate transformation matrix from the frame to the vehicle coordinate system at the current moment.
[0045] The computing device processes the fused image features of each historical frame using the corresponding coordinate transformation matrix, so that the fused image features of historical frames are uniformly mapped to the vehicle coordinate system at the current moment. After alignment, the image features from different times correspond to the same or adjacent actual road areas at the same spatial location.
[0046] After feature alignment, the computing device uses an attention-based view transformation mechanism to map multi-view image features to a bird's-eye view space. The computing device sets a set of learnable bird's-eye view query vectors. Each bird's-eye view query vector is used to carry scene information for the corresponding location in the bird's-eye view space.
[0047] The computing device first processes the bird's-eye view query vector. Spatial self-attention operations are performed to establish spatial associations between different bird's-eye view query locations. Subsequently, the bird's-eye view query vectors processed by spatial self-attention are used to perform cross-attention information interaction with the aligned multi-scale fused image features.
[0048] The aggregation relationship of the original bird's-eye view features is represented as follows:
[0049] in, Indicates the original bird's-eye view features. This represents a learnable bird's-eye view query vector. This represents the features of the aligned multi-scale fused image. This indicates a cross-attention information aggregation operation.
[0050] The computing device inputs the raw bird's-eye view features into a sparse coding network. The sparse coding network extracts the road structure over a large spatial range through stepwise downsampling, restores the bird's-eye view spatial resolution through stepwise upsampling, and transmits local spatial details through skip connections between corresponding levels. After processing, refined bird's-eye view features corresponding to each acquisition time are obtained, and the refined bird's-eye view features from multiple consecutive time points are arranged in chronological order to form a bird's-eye view feature sequence.
[0051] Since the features of the fused historical frames have been corrected to the same vehicle coordinate system, the refined bird's-eye view features at adjacent time points lie within a consistent bird's-eye view coordinate frame. The computing device performs a weighted pixel-by-pixel difference calculation on the refined bird's-eye view features of adjacent frames.
[0052] The difference enhancement feature between the current time and the previous time is represented as:
[0053] in, This indicates a difference enhancement feature. This indicates the amplification weighting coefficient for the region of change. This indicates a detailed bird's-eye view of the current moment. This indicates a detailed aerial view of the features from the previous moment.
[0054] When the input includes three or more consecutive frames, the above difference operation can be performed on adjacent frame pairs in chronological order. In areas corresponding to roads, buildings, and fixed transportation facilities, the changes in features at adjacent times are relatively small; in areas with vehicle movement, a more obvious difference response is formed between features at adjacent times.
[0055] The computing device inputs the difference enhancement features into the acyclic gated transform encoder. This encoder includes a temporal attention unit, a spatial attention unit, and a gated linear unit. The acyclic gated transform encoder first performs temporal attention operations to establish feature associations between different acquisition times; then it performs spatial attention operations to establish associations between different locations in the bird's-eye view; the gated linear unit filters information based on the current feature content.
[0056] Temporally enhanced shared bird's-eye view features are represented as follows:
[0057] in, This indicates temporal enhancement shared bird's-eye view features. This represents a loop-free gated spatiotemporal coding operation consisting of temporal attention operations, spatial attention operations, and gated linear processing.
[0058] The computing device inputs the temporally enhanced shared bird's-eye view features into the lane topology decoding branch and the vehicle instance prediction branch respectively, so that the two branches do not need to repeatedly perform multi-view image feature extraction and bird's-eye view spatial mapping.
[0059] In the lane topology decoding branch, the computing device first generates a lane polyline prototype for initializing lane queries. The computing device uniformly divides the bird's-eye view space into grid partitions with a preset number of rows and columns, and counts the number of lane samples in each grid partition based on the lane samples used to train or build the lane prototype.
[0060] No. The density allocation weights for each grid partition are represented as follows:
[0061] in, Indicates the first Density weights are assigned to each grid partition. Indicates the first The number of lane samples within each grid partition.
[0062] The computing device allocates a corresponding number of lane prototypes based on the density weight of each grid partition, and clusters the lane samples within each grid partition to obtain lane polyline prototypes that can represent the typical lane morphology within the corresponding grid partition. The lane polyline prototypes obtained in each grid partition together form the lane query initialization reference set.
[0063] The lane topology decoding branch initializes lane queries using the lane polyline prototype. During decoding, the lane topology decoding branch predicts lane polyline coordinate corrections layer by layer. The normalized coordinates of the lane polyline prototype are used as the initial lane polyline coordinates. , No. The lane polygon coordinates of the layer are updated according to the following formula:
[0064] in, Indicates the first Normalized coordinates of lane polylines after layer iteration Indicates the first Normalized coordinates of lane polylines after layer iteration Indicates the coordinate scaling factor. Indicates the first The original coordinate correction amount for layer prediction.
[0065] Hyperbolic tangent mapping compresses the original coordinate correction to a finite range, and the coordinate scaling factor further controls the coordinate correction magnitude of a single-layer iteration.
[0066] After completing the lane geometry feature update, the lane topology decoding branch generates a soft connectivity adjacency matrix between lanes based on the lane query features. Furthermore, the graph convolutional connectivity aggregation operation is used to extract connectivity features corresponding to each lane query.
[0067] For the For lane group queries, the topology gating value is represented as follows:
[0068] in, Indicates the first The topology gating value corresponding to the lane group query. Represents the gating weight vector. Indicates the first Lane query features corresponding to group lane queries Indicates the gating bias. This represents a nonlinear gating function that maps linear computation results to topological gating values.
[0069] The feature fusion relationship under query-by-query topology gating is represented as follows:
[0070] in, This represents the lane query features after topological inference. This indicates the lane query features input. This represents the multilayer perceptron transformation operation. This represents the aggregation operation on graph convolution connectivity. This represents the soft connectivity adjacency matrix between lanes.
[0071] For lanes requiring connectivity modeling, the topology gate value enables the corresponding graph convolutional connectivity features to participate in lane query updates. For relatively isolated road elements such as zebra crossings, the impact of the corresponding graph convolutional connectivity features on query updates weakens when their topology gate value approaches 0.
[0072] After completing multi-layer decoding, the lane topology decoding branch outputs lane geometry and topology prediction results. These results may include lane centerlines, road boundary types, and lane connectivity within the intersection area.
[0073] In vehicle instance prediction branches, the computing device inputs temporally enhanced shared bird's-eye view features into a multi-layer coding structure. The multi-layer coding structure generates feature representations at different spatial scales.
[0074] The vehicle instance prediction branch uses two parallel output channels: one channel generates the semantic segmentation map for future frames, and the other channel generates the reverse motion flow map for future frames. The output relationship at each future moment is represented as follows:
[0075] in, Indicates the first A semantic segmentation map of future moments. Indicates the first A reverse motion flow graph at a future moment, This represents multi-scale instance prediction computation.
[0076] The semantic segmentation map is used to distinguish between vehicle-occupied areas and background areas within the bird's-eye view space. The reverse motion flow graph records the spatial offset of the corresponding vehicle instance relative to the previous time step. For the vehicle-occupied position in the semantic segmentation map at a future time step, the computing device determines its corresponding position at the previous time step based on the corresponding reverse motion flow information and passes the vehicle instance identifier from the previous time step to the current prediction position. Through the reverse association between consecutive future frames, a multi-frame prediction result of future vehicle instances with instance identifiers is formed.
[0077] In the actual inference process, the computing device aligns the road topology vector information output by the lane topology decoding branch and the multi-frame future vehicle instance prediction results output by the vehicle instance prediction branch with the same format, and encapsulates them into scene-aware data packets according to a unified data format.
[0078] The encapsulation relationship is represented as follows:
[0079] in, This represents a scene-aware data packet. Represents road topology vector information, This represents the prediction results of future vehicle instances across multiple frames. This indicates that the data format is uniformly encapsulated for computation.
[0080] The computing device sends scene-aware data packets to the vehicle planning and control module. The vehicle planning and control module determines the passable area based on the road topology vector information and generates the control information required for following, avoiding, or turning at intersections by combining the prediction results of multiple frames of future vehicle instances.
[0081] Example 2 This embodiment further illustrates the training process of the prediction network.
[0082] The training samples include synchronous multi-view images in continuous time sequence, vehicle positioning and pose information corresponding to each acquisition time, and labeled data corresponding to each prediction task. The labeled data is used to supervise the current frame bird's-eye view segmentation, lane geometry and topology prediction, future vehicle semantic segmentation, reverse motion flow, pixel offset, and target centrality prediction.
[0083] The first stage only iterates parameters for the image feature extraction module and the bird's-eye view spatial mapping module. The temporal enhancement module, lane topology decoding branch, and vehicle instance prediction branch do not participate in the parameter update in the first stage. The computing device calculates the bird's-eye view segmentation task loss for the current frame based on the current frame's bird's-eye view segmentation annotations, and updates the parameters of the image feature extraction module and the bird's-eye view spatial mapping module based on this loss.
[0084] After the first stage is completed, save the network parameters of the image feature extraction module and the bird's-eye view spatial mapping module.
[0085] The second stage fixes the parameters of the image feature extraction module and the bird's-eye view spatial mapping module, which were trained in the first stage. Fixing the parameters means that during the inverse parameter update process in the second stage, the parameters already obtained by the above modules are not changed according to the loss of the second stage.
[0086] The computing device initiates the timing enhancement module, the lane topology decoding branch, and the vehicle instance prediction branch. The lane topology decoding branch outputs lane geometry and topology prediction results. The computing device compares these prediction results with the corresponding annotations to obtain the lane topology regression loss.
[0087] The vehicle instance prediction branch outputs semantic segmentation maps for future frames, inverse motion flow graphs, pixel offset prediction results, and object centrality prediction results. The dynamic weighted loss of the vehicle instance prediction branch is expressed as:
[0088] in, This represents the dynamic weighted loss of the predicted branch for vehicle instances. This represents the total number of future frames to be predicted. Indicates the future frame number. Represents semantic segmentation loss. Indicates the loss of the reverse flow. Indicates pixel offset loss, This represents the target centrality loss.
[0089] , , and These represent the dynamic weights for the corresponding losses. These dynamic weights can be adjusted during training to coordinate the impact of different instance prediction subtasks on parameter updates.
[0090] In the second stage, the lane topology regression loss and the dynamically weighted loss of the vehicle instance prediction branch are used together to update the parameters of the temporal enhancement module, the lane topology decoding branch, and the vehicle instance prediction branch. After the parameter update is completed, the network parameters obtained from the second stage training are saved.
[0091] Example 3 This embodiment further illustrates the process of outputting and using the prediction results.
[0092] The lane topology decoding branch outputs road topology vector information. This information includes the geometric representation of lane centerlines and the connectivity between lanes. The vehicle instance prediction branch outputs vehicle instance prediction results for multiple future timeframes. Each future timeframe's prediction includes the vehicle-occupied area and the corresponding vehicle instance identifier.
[0093] The computing device performs format alignment on the road topology vector information and the prediction results of future vehicle instances from multiple frames, and encapsulates them into a scene-aware data packet according to the following formula:
[0094] The scene-aware data package includes at least the road connectivity topology in the current scene and vehicle instance prediction information for multiple future time points.
[0095] The vehicle planning and control module determines the passable area in the current scenario based on the lane centerline and lane connectivity. For intersection scenarios, the planning and control module can determine the subsequent lanes or road branches that the current vehicle can enter based on lane connectivity.
[0096] The planning and control module can also obtain the future movement trend of each vehicle instance based on its predicted position at different future times. When there is a vehicle traveling in the same direction in front of the vehicle, the planning and control module can generate following control information; when a predicted vehicle instance enters the vehicle's planned passage area, it can generate avoidance control information; when the vehicle is at an intersection and needs to turn, it can generate intersection turning control information by combining the road connectivity topology and the prediction results of future vehicle instances.
[0097] Example 4 This embodiment further illustrates the incremental training method of the network during actual operation.
[0098] After the prediction network completes its initial training and is put into operation, the data acquisition equipment continuously collects real-world road scene data. Each set of real-time acquired scene samples includes synchronized multi-view original images, corresponding road topology ground truth annotations, and the actual vehicle trajectories at each time point. A single set of real-time acquired scene samples is denoted as... Multiple sets of real-time acquired scene samples form an incremental training sample set. .
[0099] The computing device reads incremental training samples according to the data processing flow corresponding to the initial training. For each set of incremental training samples, the computing device performs multi-view feature extraction, historical feature spatial alignment, bird's-eye view spatial mapping, temporal difference enhancement, lane topology decoding, and vehicle instance prediction, and calculates the multi-task comprehensive loss based on the corresponding ground truth data.
[0100] The cumulative loss corresponding to the incremental training sample set is expressed as:
[0101] in, This represents the cumulative loss corresponding to the incremental training sample set. This represents a single set of incremental training samples. This represents the incremental training sample set. This represents the multi-task integrated loss calculated for a single set of incremental training samples.
[0102] The computing device is based on the cumulative loss The parameters of the image feature extraction module, temporal enhancement module, lane topology decoding branch, and vehicle instance prediction branch are fine-tuned. After incremental training is completed, the computing device saves the network weights with fine-tuned parameters and calls the updated network weights in subsequent runs.
[0103] Example 5 This embodiment provides a bird's-eye view vehicle instance and lane topology prediction system.
[0104] The system includes a multi-view image acquisition unit, a vehicle positioning and attitude information acquisition unit, and a computing device; it may also include a planning and control module. The various units interact with each other via the vehicle's data transmission channel.
[0105] The multi-view image acquisition section includes multiple camera devices arranged around the vehicle. These cameras continuously acquire time-series images of the vehicle from different perspectives and send the acquired images to the computing device. The vehicle positioning and attitude information acquisition section sends the vehicle positioning and attitude information corresponding to each acquisition moment to the computing device.
[0106] The computing device includes a processor and a memory. The memory stores image feature extraction programs, bird's-eye view spatial mapping programs, temporal enhancement programs, lane topology decoding programs, vehicle instance prediction programs, and corresponding network parameters. The processor calls these programs and network parameters to complete the corresponding data processing.
[0107] The image feature extraction program performs multi-scale feature extraction and fusion on multi-view images. The bird's-eye view spatial mapping program completes historical feature spatial alignment based on vehicle positioning and attitude information, and uses bird's-eye view query vectors, spatial self-attention, and cross-attention to map image features to the bird's-eye view space.
[0108] The temporal enhancement program performs differential enhancement on the bird's-eye view features of adjacent frames located in the same bird's-eye view coordinate system, and generates temporally enhanced shared bird's-eye view features through a non-cyclic gated transform encoder. The temporally enhanced shared bird's-eye view features are then transmitted to the lane topology decoding program and the vehicle instance prediction program, respectively.
[0109] The lane topology decoding program initializes the lane query based on the lane polyline prototype, updates the lane polyline coordinates layer by layer based on the coordinate correction amount compressed by the hyperbolic tangent function and scaled by the coordinate scaling factor, and generates lane geometry and topology prediction results by using the convolutional connectivity features of the query-by-query topology gating adjustment graph.
[0110] The vehicle instance prediction program generates semantic segmentation maps and inverse motion flow maps for each future frame based on temporal enhanced shared bird's-eye view features, and generates multi-frame future vehicle instance prediction results based on the inverse motion flow map.
[0111] The computing device encapsulates road topology vector information and multi-frame predictions of future vehicle instances into a scene-aware data packet, and sends the scene-aware data packet to the planning and control module. The planning and control module generates vehicle control information based on the road connectivity topology and the locations of future vehicle instances.
[0112] The memory can also store incremental training programs. The processor executes the incremental training program, calculates the cumulative loss based on the incremental training sample set, fine-tunes the parameters of the image feature extraction module, the temporal enhancement module, the lane topology decoding branch, and the vehicle instance prediction branch, and saves the updated network weights to the memory.
[0113] In this embodiment, each program can be implemented by a processor executing pre-stored computer instructions. Feature data and prediction results are passed between programs according to the data processing order described above. Image acquisition, positioning and attitude information acquisition, network inference, and planning and control can be executed continuously during vehicle operation.
[0114] In an implementation that uses multiple camera devices as perception inputs, the prediction network can be processed directly based on multi-view images without requiring point clouds as a necessary input to the prediction network.
[0115] The above embodiments illustrate the technical solution, network structure, data processing flow, training method, prediction result usage method, incremental update method, and system composition of the present invention. The technical features in each embodiment can be combined and used together where there is no conflict.
[0116] Those skilled in the art can make conventional adjustments to the image acquisition, feature processing, network execution, and data transmission processes based on the vehicle's perception hardware configuration, computing resources, and application scenarios. These adjustments should not alter the data processing relationships between historical feature spatial alignment, bird's-eye view feature difference enhancement, shared bird's-eye view temporal representation, geometrically constrained lane topology decoding, and future vehicle instance prediction.
Claims
1. A method for predicting vehicle instances and lane topology from a bird's-eye view, characterized in that, include: The system acquires continuous temporal images of the vehicle from multiple perspectives and the vehicle's positioning attitude information. It then performs multi-scale feature extraction and fusion on the images from each perspective. Based on the vehicle's positioning attitude information, it determines the coordinate transformation matrix of the vehicle coordinate system from the historical frame to the current time and uses the coordinate transformation matrix to align the features of the fused historical frame images. Spatial self-attention processing is performed on the learnable bird's-eye view query vector, and the processed bird's-eye view query vector is made to interact with the aligned fused image features across attention information to generate original bird's-eye view features. The original bird's-eye view features are then refined through a sparse coding network to obtain continuous frame bird's-eye view features. Weighted pixel-by-pixel difference calculation is performed on the bird's-eye view features of adjacent frames located in the same bird's-eye view coordinate system. Temporal attention operation, spatial attention operation and gated linear processing are performed sequentially through a non-cyclic gated transform encoder to obtain temporally enhanced shared bird's-eye view features. The bird's-eye view space is divided into grid partitions. The number of lane prototypes is allocated according to the lane sample density in each grid partition. Lane polyline prototypes are generated by clustering in each grid partition and lane queries are initialized with the lane polyline prototypes. The temporal enhanced shared bird's-eye view features are input into the lane topology decoding branch and the vehicle instance prediction branch, respectively. The lane topology decoding branch updates the lane polyline coordinates layer by layer based on the coordinate correction amount compressed by the hyperbolic tangent function and scaled by the coordinate scaling factor. It generates a topology gating value based on the query features of each lane, and uses the topology gating value to weight the graph convolutional connectivity features obtained based on the soft connectivity adjacency matrix between lanes to obtain the lane geometry and topology prediction results. The vehicle instance prediction branch generates semantic segmentation maps and reverse motion flow graphs for each future frame, and transmits vehicle instance identifiers based on the reverse motion flow graph to obtain multi-frame future vehicle instance prediction results.
2. The method for predicting vehicle instances and lane topology from a bird's-eye view according to claim 1, characterized in that, The multi-scale feature extraction and fusion includes: determining the corresponding scale alignment coefficient based on the downsampling ratio of each feature level; upsampling the corresponding feature layers using the scale alignment coefficients; and fusing the upsampled feature layers sequentially through a convolution unit, a normalization unit, and an activation unit to obtain multi-scale fused image features; The spatial alignment relationship of frame history image features satisfies ,in, For the first Spatially aligned fused image features For the first Multi-scale fused image features before frame space alignment For the first The coordinate transformation matrix from the frame to the vehicle coordinate system at the current moment.
3. The bird's-eye view vehicle instance and lane topology prediction method according to claim 2, characterized in that, The bird's-eye view query vector after spatial self-attention processing is interacted with the aligned multi-scale fused image features across attentional information to obtain the original bird's-eye view features, which satisfy the following conditions: ,in, The original bird's-eye view features, For the learnable bird's-eye view query vector, For the aligned multi-scale fused image features, For cross-attention information aggregation operations, the sparse coding network performs multi-scale spatial refinement of the original bird's-eye view features through step-by-step downsampling, step-by-step upsampling, and skip connections between corresponding levels.
4. The method for predicting vehicle instances and lane topology from a bird's-eye view according to claim 3, characterized in that, The difference enhancement feature obtained by the weighted pixel-by-pixel difference operation satisfies ,in, For the difference enhancement feature, Increase the weighting coefficient for the region of change. A detailed aerial view of the current moment. The bird's-eye view features are refined from the previous time step; the temporally enhanced shared bird's-eye view features satisfy... ,in, For the time-series enhanced shared bird's-eye view features, It is a loop-free gated spatiotemporal coding operation consisting of a temporal attention unit, a spatial attention unit, and a gated linear unit, wherein the temporal attention unit is executed before the spatial attention unit.
5. The method for predicting vehicle instances and lane topology from a bird's-eye view according to claim 1, characterized in that, The bird's-eye view space is uniformly divided into grid partitions with a preset number of rows and columns, and the density allocation weight is determined according to the number of lane samples in each grid partition. The density allocation weights of each grid partition satisfy... ,in, For the first Density weights are assigned to each grid partition. For the first The number of lane samples within each grid partition; the number of lane prototypes in each grid partition is allocated according to the density allocation weight, and the lane samples are clustered within each grid partition to obtain the lane polyline prototype.
6. The method for predicting vehicle instances and lane topology from a bird's-eye view according to claim 5, characterized in that, The normalized coordinates of the lane polyline prototype are used as the initial lane polyline coordinates. The lane polyline coordinates are based on Iterative updates are performed layer by layer, where... For the lane topology decoding branch number Normalized coordinates of lane polylines after layer iteration For the first Normalized coordinates of lane polylines after layer iteration This is the coordinate scaling factor. For the first The original coordinate correction amount for layer prediction.
7. The method for predicting vehicle instances and lane topology from a bird's-eye view according to claim 6, characterized in that, No. The topology gating value corresponding to the lane group query satisfies ,in, For the first The topology gating value corresponding to the lane group query. For the gated weight vector, For the first Lane query features corresponding to group lane queries This is the gating bias. A nonlinear gating function is used to map the linear calculation results to topological gating values; based on the lane query features transformed by the multilayer perceptron and the soft connectivity adjacency matrix between lanes, graph convolutional connectivity features are obtained, and the first... The topological gating value corresponding to the lane query group is used as the weighting coefficient of the graph convolutional connectivity feature corresponding to the lane query group, and the weighted graph convolutional connectivity feature is fused with the corresponding original lane query feature and multilayer perceptron transform feature.
8. The method for predicting vehicle instances and lane topology from a bird's-eye view according to claim 1, characterized in that, The vehicle instance prediction branch includes a multi-layer coding structure and two parallel output channels. The multi-layer coding structure generates feature representations corresponding to different spatial scales based on the temporally enhanced shared bird's-eye view features. The two parallel output channels respectively generate the first... Semantic segmentation graphs and reverse motion flow graphs for each future time step, satisfying... ,in, For the first A semantic segmentation map for a future moment, wherein the semantic segmentation map represents the area occupied by vehicles in a bird's-eye view space. For the first A reverse motion flow graph for each future time step, wherein the reverse motion flow graph records the spatial offset of the corresponding vehicle instance at the previous time step. This is for multi-scale instance prediction computation.
9. The method for predicting vehicle instances and lane topology from bird's-eye view images according to claim 1, characterized in that, Before performing the prediction, the network used to perform the bird's-eye view vehicle instance and lane topology prediction method is trained in two stages: In the first stage, only the image feature extraction module and the bird's-eye view spatial mapping module are iterated, and the temporal enhancement module, lane topology decoding branch and vehicle instance prediction branch are not started, and the current frame bird's-eye view segmentation task is used as the training target; In the second stage, the parameters of the image feature extraction module and the bird's-eye view spatial mapping module trained in the first stage are fixed, and the parameters of the temporal enhancement module, lane topology decoding branch and vehicle instance prediction branch are updated.
10. The method for predicting vehicle instances and lane topology from a bird's-eye view according to claim 9, characterized in that, The second stage employs a multi-task loss mechanism, including lane topology regression loss and vehicle instance prediction branch dynamic weighting loss, wherein the vehicle instance prediction branch dynamic weighting loss satisfies ,in, Predict the dynamic weighted loss of the branch for the vehicle instance. The total number of future frames to be predicted. For future frame numbers, For semantic segmentation loss, For reverse flow loss, For pixel offset loss, For target centrality loss, , , and These are the dynamic weights corresponding to the loss, and the dynamic weights are adaptively adjusted with each training round.
11. The method for predicting vehicle instances and lane topology from bird's-eye view images according to claim 1, characterized in that, It also includes: aligning the road topology vector information in the lane geometry and topology prediction results with the multi-frame future vehicle instance prediction results in terms of format, and encapsulating them into a scene-aware data packet according to a unified data format; sending the scene-aware data packet to the vehicle's planning and control module, wherein the planning and control module determines the passable area based on the road topology vector information, and generates at least one of following control information, avoidance control information, and intersection turning control information based on the multi-frame future vehicle instance prediction results.
12. The method for predicting vehicle instances and lane topology from a bird's-eye view according to claim 9, characterized in that, Also includes: During the model execution phase, synchronous multi-view original images and the actual vehicle motion trajectories at each time point are continuously acquired, and road topology ground truth annotations corresponding to the synchronous multi-view original images are obtained to form an incremental training sample set; according to Calculate the cumulative loss corresponding to the incremental training sample set, where, The cumulative loss corresponding to the incremental training sample set, For a single set of incremental training samples, The incremental training sample set, The multi-task comprehensive loss is calculated for a single set of incremental training samples. Based on the accumulated loss, the parameters of the image feature extraction module, temporal enhancement module, lane topology decoding branch and vehicle instance prediction branch are simultaneously fine-tuned, and the network weights after parameter fine-tuning are saved.
13. A bird's-eye view vehicle instance and lane topology prediction system, characterized in that, The system includes a multi-view image acquisition section, a vehicle positioning and attitude information acquisition section, and a computing device. The multi-view image acquisition section includes multiple cameras arranged around the vehicle to acquire continuous temporal images of the vehicle from multiple perspectives. The vehicle positioning and attitude information acquisition section is used to acquire vehicle positioning and attitude information corresponding to each acquisition time. The computing device includes a processor and a memory. The memory stores an image feature extraction program, a bird's-eye view spatial mapping program, a temporal enhancement program, a lane topology decoding program, a vehicle instance prediction program, and corresponding network parameters. The processor calls the program and the corresponding network parameters to execute the method described in claim 1.
14. The bird's-eye view vehicle instance and lane topology prediction system according to claim 13, characterized in that, It also includes a planning and control module that is communicatively connected to the computing device. The computing device is used to perform format alignment on the road topology vector information and the prediction results of multiple future vehicle instances, and encapsulate them into scene-aware data packets according to a unified data format before sending them to the planning and control module. The planning and control module is used to determine the passable area based on the road topology vector information, and generate at least one of following control information, avoidance control information, and intersection turning control information based on the prediction results of multiple future vehicle instances.
15. The bird's-eye view vehicle instance and lane topology prediction system according to claim 13, characterized in that, The memory also stores an incremental training program, which the processor executes to form an incremental training sample set based on the synchronous multi-view original images, corresponding road topology ground truth annotations, and the actual vehicle motion trajectories at each time point. The cumulative loss corresponding to the incremental training sample set is calculated, and the parameters of the image feature extraction module, temporal enhancement module, lane topology decoding branch, and vehicle instance prediction branch are simultaneously fine-tuned based on the cumulative loss. The network weights with fine-tuned parameters are stored in the memory. The cumulative loss corresponding to the incremental training sample set, For a single set of incremental training samples, The incremental training sample set, This is the multi-task integrated loss calculated for a single set of incremental training samples.
Citation Information
Patent Citations
Road topology prediction method of high-precision map based on Transform
CN117037095A
Scene topology understanding method and device, storage medium and program product
CN120526426A