Multi-modal fusion road topology reasoning method and system
Through a multimodal fusion road topology reasoning method, crowdsourced data and multimodal models are utilized to reduce reliance on manual annotation, lower data costs, and improve the accuracy and consistency of lane topology reasoning. This solves the problems of high cost and insufficient annotation in traditional methods and achieves high-precision road topology mapping.
Patent Information
- Application Number
- CN202510895881.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-06-30
AI Technical Summary
In existing technologies for autonomous driving, traditional road topology mapping methods rely on manual high-precision maps and rule templates, resulting in high costs and difficulty adapting to dynamic environments. In addition, data-driven methods rely on large amounts of labeled data, resulting in high labeling costs and insufficient supervision performance.
A multimodal fusion road topology inference method is adopted. By collecting historical images and point clouds from crowdsourcing datasets, a multimodal model is used to extract lane masks and convert them into semantic seed points. Pseudo-labels are set, and a lane detection model is trained. The lane masks are mapped to the bird's-eye view feature map, the topological relationships are calculated, and the road topology inference results are obtained using a graph neural network.
It reduces data preparation costs, improves the accuracy and topological consistency of lane topology reasoning, can effectively capture the spatial structure between lane lines, and provide high-precision lane topology reasoning results with good topological consistency.
Smart Images

Figure CN120725147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving perception technology, and in particular to a multimodal fusion road topology reasoning method and system. Background Art
[0002] In recent years, with the rapid development of intelligent connected vehicles and autonomous driving technology, the requirements for road semantic understanding and structured mapping in complex traffic environments have increased significantly. As a key component of the perception layer of autonomous driving systems, accurate understanding of road topology not only directly impacts the performance of path planning, decision-making, and control, but also the safety and interpretability of the entire system.
[0003] However, traditional road topology mapping methods often rely on artificial high-precision maps (HD maps) and the prior structure of rule templates, combined with on-board camera or radar data for offline matching and mapping. While these methods offer high accuracy, they suffer from two major drawbacks: First, their reliance on HD maps makes system construction and maintenance expensive and difficult to adapt to dynamic road environments; second, their over-reliance on rule-based prior models lacks flexibility and scenario generalization, making topology misjudgment particularly prone to complex lane configurations (such as merges, forks, and U-turns).
[0004] With the rise of crowdsourced map data and the development of data-driven methods, large amounts of image and point cloud data collected by vehicles are being used to train lane detection models, making it possible to automatically infer accurate lane positions from perception results. However, current data-driven methods generally rely on large amounts of annotated data, which in real-world applications faces challenges such as high annotation costs and insufficient supervision. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a multimodal fusion road topology reasoning method and system to obtain lane topology reasoning results with high precision and good topological consistency, while reducing data preparation costs.
[0006] In a first aspect, the present invention provides a multimodal fusion road topology reasoning method, the method comprising: Collect historical images and historical point clouds from a crowdsourced dataset, extract lane masks from the historical images using a multimodal model, convert the lane masks into semantic seed points in the historical point clouds, and perform pseudo-labeling based on the semantic seed points. The lane detection model is obtained by model training based on historical images and historical point clouds with pseudo labels; Obtaining lane line information from real-time images and real-time point clouds using the lane detection model, and mapping the lane line information into a bird's-eye view feature map; The topological relationship between lane lines is calculated based on the lane line information in the bird's-eye view feature map to obtain a topological reasoning matrix, and the corresponding road topology reasoning result is obtained based on the graph neural network and the topological reasoning matrix.
[0007] In an optional embodiment, the step of extracting the lane mask in the historical image using a multimodal model includes: Segment the historical image using a multimodal model to generate a background semantic mask. and obtaining a description text of each background semantic mask; The similarity between the description text and predefined category text is calculated, and the background semantic mask is filtered according to the obtained similarity value to determine the lane mask.
[0008] In an optional embodiment, the step of converting the lane mask into semantic seed points in the historical point cloud includes: Obtaining the maximum and minimum values of the lane mask in the pixel coordinate system; performing a boundary shrinkage process on the lane mask based on the maximum value and the minimum value; The lane mask after the boundary shrinkage processing is converted from the historical image to the historical point cloud, and the area corresponding to the lane mask is determined as a semantic seed point.
[0009] In an optional embodiment, the step of setting a pseudo label based on the semantic seed point includes: performing a clustering operation on the semantic seed points in the historical point cloud to determine a plurality of clusters; A bounding box is fitted for each of the clusters, and a pseudo label is obtained based on the bounding box.
[0010] In an optional embodiment, the step of performing model training based on historical images and historical point clouds with pseudo labels to obtain a lane detection model includes: For each pseudo label in the historical image, respectively calculating a distribution constraint score and a meta-shape constraint score of the pseudo label; After normalizing the distribution constraint score and the meta-shape constraint score, they are accumulated according to the set weight factor to obtain a distribution shape score; The pseudo labels are screened based on the distribution shape scores, and a lane detection model is obtained by performing model training based on historical images and historical point clouds corresponding to the screened pseudo labels.
[0011] In an optional embodiment, the real-time image includes images from multiple perspectives; The step of mapping the lane line information into the bird's-eye view feature map includes: Extracting image features of images at each viewing angle in the real-time image, and converting each pixel point in the image at each viewing angle into a three-dimensional point in a ground coordinate system; Project the image features of each perspective into the bird's-eye view space according to the three-dimensional coordinates of the three-dimensional points of the image of each perspective in the ground coordinate system to obtain an initial bird's-eye view feature map; The two-dimensional coordinates and lane line category vector of each lane line in the lane line information in the bird's-eye view space are obtained, and the two-dimensional coordinates and lane line category vector of each lane line are written into the initial bird's-eye view feature map to obtain a final bird's-eye view feature map.
[0012] In an optional embodiment, the step of calculating the topological relationship between lane lines based on the lane line information in the bird's-eye view feature image to obtain a topological reasoning matrix includes: Obtaining a lane query based on lane line information in the final bird's-eye view feature map through a lane deformable decoder, and generating a plurality of directional lane lines based on the lane query; Calculating a lane geometry topology matrix based on the plurality of directional lane lines, and calculating a lane similarity topology matrix based on the lane query; The lane geometry topology matrix and the lane similarity topology matrix are combined to obtain a topology reasoning matrix.
[0013] In an optional embodiment, the step of calculating a lane geometric topology matrix based on the plurality of directional lane lines and calculating a lane similarity topology matrix based on the lane query includes: Calculating a lane geometric distance between any two directional lane lines among the plurality of directional lane lines to obtain a lane geometric distance matrix; Obtaining a lane geometric topology matrix based on the lane geometric distance matrix and the learned mapping function; Using two multi-layer perceptrons to encode the lane query respectively to obtain encoding results, and performing inner product calculation on the two encoding results to obtain lane similarity; The lane similarity is mapped to the lane topology to obtain a lane similarity topology matrix.
[0014] In an optional embodiment, the step of obtaining a corresponding road topology reasoning result based on the graph neural network and the topology reasoning matrix includes: Inputting the topological reasoning matrix and the obtained lane query into a graph neural network to aggregate information of adjacent lane lines corresponding to the lane query to obtain an enhanced lane query; The enhanced lane query is processed using the lane header function to obtain the corresponding road topology inference result.
[0015] In a second aspect, the present invention provides a multimodal fusion road topology reasoning system, the system comprising: A setting module is used to collect historical images and historical point clouds from a crowdsourced dataset, extract lane masks from the historical images using a multimodal model, convert the lane masks into semantic seed points in the historical point clouds, and perform pseudo-label setting based on the semantic seed points; A training module is used to train a lane detection model based on historical images and historical point clouds with pseudo labels; a mapping module, configured to obtain lane line information from the real-time image and the real-time point cloud using the lane detection model, and map the lane line information into a bird's-eye view feature map; An inference module is used to calculate the topological relationship between lane lines based on the lane line information in the bird's-eye view feature map, obtain a topological inference matrix, and obtain corresponding road topology inference results based on the graph neural network and the topological inference matrix.
[0016] The present invention provides a multimodal fusion road topology reasoning method and system. By collecting historical images and historical point clouds from a crowdsourced dataset, a multimodal model is used to extract lane masks from the historical images, the lane masks are converted into semantic seed points in the historical point clouds, and pseudo-labels are set based on the semantic seed points. A lane detection model is obtained by model training based on historical images and historical point clouds with pseudo-labels. The lane detection model is used to obtain lane line information from real-time images and real-time point clouds, and the lane line information is mapped into a bird's-eye view feature map. The topological relationship between lane lines is calculated based on the lane line information in the bird's-eye view feature map to obtain a topological reasoning matrix. Then, based on a graph neural network and the topological reasoning matrix, the corresponding road topology reasoning result is obtained.
[0017] In this scheme, the semantic mask extraction and pseudo-label generation mechanism guided by a multimodal model reduces reliance on manual annotation, significantly lowering data preparation costs while ensuring detection quality. Furthermore, the introduction of a lane line representation method based on a unified bird's-eye view spatial mapping and the effective fusion of lane topology reasoning can effectively capture the spatial structure between lane lines, resulting in highly accurate lane topology reasoning results with good topological consistency. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1A flowchart of a multimodal fusion road topology reasoning method provided by an embodiment of the present invention; Figure 2 for Figure 1 Flowchart of the sub-steps included in S11; Figure 3 for Figure 1 Flowchart of the sub-steps included in S12; Figure 4 for Figure 1 Flowchart of the sub-steps contained in S13; Figure 5 for Figure 1 Flowchart of the sub-steps included in S14; Figure 6 for Figure 1 Flowchart of the sub-steps included in S15; Figure 7 for Figure 1 A flowchart of the sub-steps included in S16; Figure 8 for Figure 1 Flowchart of the sub-steps included in S17; Figure 9 Provides a functional module block diagram of a multimodal fusion road topology reasoning system for an embodiment of the present invention; Figure 10 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.
[0021] See also Figure 1 The following is a flowchart of a multimodal fusion road topology inference method provided in an embodiment of the present invention. This multimodal fusion road topology inference method can be performed by a multimodal fusion road topology inference system. This multimodal fusion road topology inference system can be implemented using software and / or hardware and configured in an electronic device, such as a computer or server, for example, a server in a backend control platform. The detailed steps of this multimodal fusion road topology inference method are described below.
[0022] S11, collecting historical images and historical point clouds from a crowdsourcing dataset, and extracting lane masks from the historical images using a multimodal model.
[0023] S12: Convert the lane mask into a semantic seed point in the historical point cloud.
[0024] S13: Setting pseudo labels based on the semantic seed points.
[0025] S14, performing model training based on historical images and historical point clouds with pseudo labels to obtain a lane detection model.
[0026] S15, using the lane detection model to obtain lane line information from the real-time image and the real-time point cloud, and mapping the lane line information to the bird's-eye view feature map.
[0027] S16, calculating the topological relationship between lane lines based on the lane line information in the bird's-eye view feature image to obtain a topological inference matrix.
[0028] S17, obtaining a corresponding road topology reasoning result based on the graph neural network and the topology reasoning matrix.
[0029] The multimodal fusion road topology inference method provided in this embodiment uses a semantic mask extraction and pseudo-label generation mechanism guided by a multimodal model to reduce reliance on manual annotation, significantly reducing data preparation costs while ensuring detection quality. Furthermore, the introduction of a lane line representation method based on a unified bird's-eye view spatial mapping and the effective integration of lane topology inference effectively capture the spatial structure between lane lines, resulting in highly accurate lane topology inference results with good topological consistency.
[0030] The specific implementation of each of the above steps is described in detail below.
[0031] In this embodiment, the historical images and historical point clouds in the crowdsourced dataset can be uploaded by a large number of users and contain rich road-related information. Each historical image is captured by the on-board surround-view camera, and the historical point cloud is a point cloud image detected by the on-board radar equipment.
[0032] See also Figure 2 In this embodiment, the step of extracting lane masks from historical images using a multimodal model can be implemented in the following manner: S111 , segmenting the historical image using a multimodal model, generating background semantic masks, and obtaining description texts of each background semantic mask.
[0033] S112 , calculating the similarity between the description text and predefined category text, and filtering the background semantic mask according to the obtained similarity value to determine a lane mask.
[0034] In this implementation, one of the historical images and its historical point cloud For example, and Indicates the height and width of the image. The multimodal model may include an image segmentation model, such as the FastSAM model, which is used to segment historical images. Perform segmentation and generate background semantic masks that are independent of the category ,in represents the spatial dimension of each background semantic mask, Indicates the number of background semantic masks.
[0035] Next, the historical image and background semantic mask Input the image segmentation model, such as the SemanticSAM model, to obtain the description text of each background semantic mask , Represents the semantic description of each background semantic mask. The semantic description can be understood as a natural language phrase generated by the SemanticSAM model for the background semantic mask in the historical image, such as lane, pedestrian, etc.
[0036] Then calculate the description text and predefined category text The similarity between them can be measured by cosine similarity. , It is the descriptor of the category, mainly including the categories that need to be identified, such as various lanes. Next, based on the obtained similarity value, the background semantic mask of no interest is filtered out to obtain the lane mask Specifically, by calculating The similarity value between them is compared with a preset threshold, and the background semantic mask corresponding to the similarity value greater than the preset threshold is determined. The determined background semantic mask is the lane mask.
[0037] Based on the lane mask, the lane mask is converted into semantic seed points in the historical point cloud. For details, see Figure 3 , this step can be achieved by: S121: Obtain the maximum and minimum values of the lane mask in the pixel coordinate system.
[0038] S122: Perform boundary shrinkage processing on the lane mask based on the maximum value and the minimum value.
[0039] S123 , converting the lane mask after the boundary shrinkage processing from the historical image to the historical point cloud, and determining the area corresponding to the lane mask as a semantic seed point.
[0040] In this embodiment, for each lane mask, its pixel coordinate system is calculated The maximum and minimum values in are recorded as Then based on the maximum and minimum values, the boundaries of the lane mask are constrained by shrinking its bounds according to the following formula: ] ] Where γ is the shrinkage factor. And the lane mask after the boundary shrinkage process is recorded as .
[0041] In addition, the camera's intrinsic and extrinsic matrix is used to convert the shrunken lane mask into From historical images in 2D image format Transfer to historical point cloud in the form of 3D radar point cloud In the radar point cloud Accurate cross-modal semantic hints. Lane mask after shrinking the boundary and converting The determined area is defined as a semantic seed point to obtain a historical point cloud containing the semantic seed point.
[0042] Among them, the intrinsic parameter matrix contains information such as the focal length and optical center of the camera, while the extrinsic parameter matrix describes the relationship between the camera and the world coordinate system.
[0043] After obtaining the historical point cloud containing semantic seed points, pseudo labels are set based on the semantic seed points. For details, please refer to Figure 4 , this step can be achieved by: S131 , performing a clustering operation on the semantic seed points in the historical point cloud to determine a plurality of clusters.
[0044] S132: Fit a bounding box for each of the clusters, and obtain a pseudo label based on the bounding box.
[0045] In this embodiment, the historical point cloud of the laser radar is represented as , defined by The covered semantic seed points are ,in . use Represents the kth instance in the point cloud framework. Design a dynamic clustering radius update function to dynamically update the clustering radius r:
[0046] in, It is a hyperparameter set based on experience, and δ is an adjustment factor used to prevent r from being too small. is the number of seed points in the current instance, and t represents the tth semantic seed point in the kth instance. By applying this formula, the cluster radius r is dynamically updated during the clustering process. Next, the DBSCAN clustering method is used to cluster the historical point cloud. In the cluster, the radius of the cluster is dynamically updated with the current semantic seed point as the center. Perform density clustering to obtain multiple clusters. Next, fit the bounding box for each cluster , and the bounding box Add to the result set. Use this method to traverse all instances in the historical point cloud and obtain the historical image The corresponding pseudo-label set .
[0047] In this embodiment, the lane detection model can be obtained by model training using historical images and historical point clouds with pseudo labels. For details, please refer to Figure 5 , this step can be achieved by: S141 , for each pseudo label in the historical image, respectively calculating a distribution constraint score and a meta-shape constraint score of the pseudo label.
[0048] S142 , normalizing the distribution constraint score and the meta-shape constraint score, and then accumulating them according to a set weight factor to obtain a distribution shape score.
[0049] S143: Filter the pseudo labels based on the distribution shape scores, and perform model training based on historical images and historical point clouds corresponding to the filtered pseudo labels to obtain a lane detection model.
[0050] To quantify the credibility of different pseudo-labels, this embodiment sets a distribution shape score, which is obtained by adding the distribution constraint score and the meta-shape constraint score according to the weights. The calculation methods of the distribution constraint score and the meta-shape constraint score are described below.
[0051] For distribution constraint score calculation, first, a pseudo label is calculated A random variable , Represents the semantic seed point within the bounding box The distance to the boundary of the bounding box. At the same time, due to the high-quality pseudo label Its internal semantic seed point The distance to the boundary follows a roughly Gaussian distribution . Thus, by calculating each pseudo label The corresponding random variable With standard Gaussian distribution The similarity between them can be used to obtain the distribution constraint score. The specific formula is:
[0052] in, represents the logarithmic function, It is a pseudo label The semantic seed point set within Indicates the number of semantic seed points.
[0053] For the element shape constraint score, define Representation category The shape of the meta-instance, where are the normalized length, width, and height respectively. Combined with the above, the element shape constraint score is obtained through processing. The specific formula is:
[0054] in, represents the normalized KL divergence function, Represents pseudo labels standardized shape.
[0055] Finally, the final distribution shape score is obtained by combining the normalized distribution constraint score and the meta-shape constraint score. :
[0056] in, and is the weight factor, and After normalization 、 .
[0057] On this basis, the distribution shape score of each pseudo label will be obtained Instead of the confidence score of each candidate pseudo-label in the traditional non-maximum suppression method, the obtained pseudo-label set Through the traditional non-maximum suppression method, high-quality pseudo labels are screened and low-quality pseudo labels are suppressed to obtain a high-quality pseudo label set. .
[0058] Finally, the pseudo-label set As a historical image and historical point clouds The historical images and point clouds in the training set selected from the crowdsourced dataset are annotated in the same way as above. The training set is then used for deep learning training to obtain a trained lane detection model.
[0059] In the actual application stage, the pre-trained lane detection model is used to process the real-time images and real-time point clouds to obtain lane line information. The real-time images include images from multiple perspectives and are real-time surround images obtained by on-board sensors. The lane line information is then mapped to the bird's-eye view feature map. For details, please refer to Figure 6 , this step is achieved by: S151 , extracting image features of images of each viewing angle in the real-time image, and converting each pixel point in the image of each viewing angle into a three-dimensional point in a ground coordinate system.
[0060] S152 , projecting the image features of each perspective into the bird's-eye view space according to the three-dimensional coordinates of the three-dimensional points of the image of each perspective in the ground coordinate system, to obtain an initial bird's-eye view feature map.
[0061] S153, obtaining the two-dimensional coordinates and lane line category vector of each lane line in the lane line information in the bird's-eye view space, and writing the two-dimensional coordinates and lane line category vector of each lane line into the initial bird's-eye view feature map to obtain a final bird's-eye view feature map.
[0062] In this embodiment, real-time surround view images are obtained based on vehicle-mounted sensors. and real-time point cloud ,in Indicates the number of different perspectives of the surround image, which is generally six. The CNN model can be used to extract multi-scale image features from each perspective:
[0063] Among them, Backbone represents the main network structure in the CNN model.
[0064] Then, using the camera's extrinsic and intrinsic matrix, the pixel points corresponding to each pixel in the image of each perspective are projected to the ground coordinate system:
[0065] in, is the pixel coordinate on the image plane, is the three-dimensional coordinate corresponding to the pixel point, is the intrinsic parameter matrix of the camera, is the extrinsic parameter matrix of the camera.
[0066] Then, the image features of all perspectives are combined with the three-dimensional coordinates through the view transformation method and reprojected into a unified bird's-eye view space (BEV space) to obtain the initial bird's-eye view feature map .
[0067] Using lane detection models to process real-time images and real-time point cloud , get the three-dimensional position of each lane line .in, represents the three-dimensional position of the point, Represents the category label of the lane line, Indicates the The number of 3D points contained in the lane line, Indicates the number of recognized lane lines.
[0068] Then, perform two-dimensional sampling on the lane line to obtain the two-dimensional coordinates , and each point of the lane line Mapped to the BEV grid, the specific mapping formula is as follows:
[0069]
[0070] in, , Indicates the minimum boundary value of the BEV area; , Indicates the actual size of each grid in the BEV image.
[0071] The detected lane line category , lane line two-dimensional coordinates and other features are encoded into vectors and written into the initial bird's-eye view feature map to obtain the enhanced BEV feature map containing the fusion image and point cloud information and lane line position information. , where C represents the channel dimension of the BEV feature, which serves as the final bird's-eye view feature map.
[0072] On this basis, the topological relationship between lane lines is calculated based on the lane line information in the bird's-eye view feature map to obtain the topological reasoning matrix. For details, please refer to Figure 7 , this step can be achieved by: S161 , obtaining a lane query based on lane line information in the final bird's-eye view feature map through a lane deformable encoder, and generating a plurality of directional lane lines based on the lane query.
[0073] S162: Calculate a lane geometry topology matrix based on the plurality of directional lane lines, and calculate a lane similarity topology matrix based on the lane query.
[0074] S163 , combining the lane geometric topology matrix and the lane similarity topology matrix to obtain a topology reasoning matrix.
[0075] In this embodiment, a set of learnable initial query vectors is defined ,in Is the upper limit of the number of lanes, C is the channel dimension of the query. and the initial query vector Input lane deformable decoder, through deformable attention mechanism, by Extract the lane semantics of important locations and output lane queries .
[0076] The initial query vector can be understood as an attention pointer, which is used to guide the lane deformable decoder to obtain the bird's-eye view feature map. Extract the corresponding lane information, which is the initial parameter. Lane query includes the information from the bird's eye view feature map The lane information of different lanes extracted can be used for subsequent lane line generation.
[0077] On this basis, the lane head function is used to query the lane And the lane line information in the bird's-eye view feature map generates multiple directional lane lines .
[0078] Then, a lane geometry topology matrix is calculated based on multiple directional lane lines, and a lane similarity topology matrix is calculated based on the lane query in the following manner: The lane geometric distance between any two directional lane lines among the multiple directional lane lines is calculated to obtain a lane geometric distance matrix; a lane geometric topology matrix is obtained based on the lane geometric distance matrix and the learned mapping function; the lane queries are respectively encoded using two multi-layer perceptrons to obtain encoding results, and an inner product is calculated on the two encoding results to obtain lane similarity; the lane similarity is mapped to the lane topology to obtain a lane similarity topology matrix.
[0079] In this embodiment, the connectivity between two different directional lanes is evaluated by calculating the distance between the start and end points of these directional lanes. The specific formula is as follows:
[0080]
[0081]
[0082] Where D is the lane geometry distance matrix, and Respectively represent the previous and next directional lane lines, Indicates a directional lane line and The geometric distance between Indicates a directional lane line The last point, Indicates a directional lane line The first point.
[0083] Through learnable mapping functions , mapping lane geometry distance to lane topology (lane topology refers to the connection and adjacent relationship between lanes for subsequent lane line generation). The specific formula is:
[0084] in, , is the lane geometry distance matrix The standard deviation of and It is a learnable parameter, that is, an adjustable parameter, and machine learning can be used to obtain better parameters.
[0085] Finally, through the function and lane geometry distance matrix , we can get the lane geometry topology matrix:
[0086] For the lane similarity topology method, two independent multilayer perceptrons are first used to perform lane query The lane similarity is obtained by encoding and then performing inner product processing on the encoding results of the two multi-layer perceptrons. The lane similarity is then mapped to the lane topology using the sigmoid function to obtain the lane similarity topology matrix. The specific formula is as follows:
[0087]
[0088]
[0089] in, ∈ , express The similarity, , represents the number of lane queries, represents the inner product operation of the matrix, Represents the transpose operation of a matrix.
[0090] After obtaining the lane geometry topology matrix and the lane similarity topology matrix, the two lane topology inference results are combined and weighted using learnable coefficients to form the final topology inference matrix. The formula is as follows:
[0091] in, and is a learnable parameter.
[0092] Based on the above, the corresponding road topology reasoning results are obtained based on the graph neural network and the topology reasoning matrix. For details, please refer to Figure 8 , this step can be achieved by: S171: Input the topological reasoning matrix and the obtained lane query into a graph neural network to aggregate information of adjacent lane lines corresponding to the lane query to obtain an enhanced lane query.
[0093] S172: Use the lane header function to process the enhanced lane query to obtain the corresponding road topology inference result.
[0094] In this embodiment, the lane query and the final topological inference matrix The input is fed into a graph neural network for multi-layer graph convolution processing, which aggregates the information of adjacent lane lines and ultimately outputs an enhanced lane query that improves the accuracy of lane topology reasoning. Finally, the enhanced lane query is performed through the lane header function Processing is performed to obtain the final road topology reasoning result.
[0095] Enhanced lane queries incorporate the topological inference matrix. This enhanced lane query includes not only individual lane information but also topological information about lanes, such as connected lanes. The resulting road topology inference results in the final lane map, which includes information such as the actual spatial shape of lane lines, lane connectivity, and lane types.
[0096] The multimodal fusion road topology inference method provided in this embodiment, on the one hand, reduces reliance on manual annotation through a semantic mask extraction and pseudo-label generation mechanism guided by a large-scale multimodal model, significantly reducing data preparation costs while ensuring detection quality. On the other hand, by introducing a lane line representation method based on a unified BEV spatial mapping and effectively integrating multiple lane topology inference methods, it effectively captures the spatial structural relationship between lane segments and outputs lane topology inference results with high accuracy, good topological consistency, and semantic interpretability.
[0097] In summary, this solution effectively leverages crowdsourced data from multimodal on-board sensors (such as cameras and radars), integrating image and point cloud information. Based on unsupervised pseudo-label generation and multimodal semantic alignment, it constructs a highly robust lane detection model. Furthermore, it effectively utilizes bird's-eye-view (BEV) features, integrates geometric relationships and semantic similarity into lane topology inference, and utilizes graph neural networks to implement lane topology inference. This approach offers strong interpretability and engineering deployability, providing new insights for efficient, safe, and cost-effective road mapping and perception systems.
[0098] Based on the same inventive concept, please refer to Figure 9, an embodiment of the present invention also provides a functional module diagram of a multimodal fusion road topology reasoning system. This embodiment can divide the functional modules of the multimodal fusion road topology reasoning system according to the above method embodiment. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is schematic and is only a logical functional division. There may be other division methods in actual implementation.
[0099] For example, when each functional module is divided into corresponding functional modules, Figure 9 The multimodal fusion road topology inference system shown is only a schematic diagram. The multimodal fusion road topology inference system can include a setup module, a training module, a mapping module, and an inference module. The functions of each functional module of the multimodal fusion road topology inference system are described in detail below.
[0100] A setting module is used to collect historical images and historical point clouds from a crowdsourced dataset, extract lane masks from the historical images using a multimodal model, convert the lane masks into semantic seed points in the historical point clouds, and perform pseudo-label setting based on the semantic seed points; A training module is used to train a lane detection model based on historical images and historical point clouds with pseudo labels; a mapping module, configured to obtain lane line information from the real-time image and the real-time point cloud using the lane detection model, and map the lane line information into a bird's-eye view feature map; An inference module is used to calculate the topological relationship between lane lines based on the lane line information in the bird's-eye view feature map, obtain a topological inference matrix, and obtain corresponding road topology inference results based on the graph neural network and the topological inference matrix.
[0101] The multimodal fusion road topology reasoning system provided in this embodiment can be used to execute the multimodal fusion road topology reasoning method under any implementation method of the above embodiments. For any details not provided in this embodiment, please refer to the corresponding description of the above embodiments, and this embodiment will not be repeated here.
[0102] See also Figure 10, is a block diagram of the structure of an electronic device provided in an embodiment of the present invention. This electronic device may be a computer device, server, or the like in an autonomous driving control platform. The electronic device includes a memory, a processor, and a communication module. The memory, processor, and communication module components are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines.
[0103] Memory is used to store computer programs or data. Memory can include, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM).
[0104] The processor is used to read / write data or programs stored in the memory and execute the multimodal fusion road topology reasoning method provided by any embodiment of the present invention.
[0105] The communication module is used to establish a communication connection between the electronic device and other communication terminals through the network, and is used to send and receive data through the network.
[0106] It should be understood that Figure 10 The structure shown is only a schematic diagram of the structure of the electronic device. The electronic device may also include Figure 10 More or fewer components than shown, or with Figure 10 Different configurations shown.
[0107] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are executed, the multimodal fusion road topology reasoning method provided in the above embodiment is implemented.
[0108] Specifically, the computer-readable storage medium can be a general-purpose storage medium, such as a removable disk or hard disk. When the computer program on the computer-readable storage medium is executed, the multimodal fusion road topology inference method described above can be implemented. The processes involved in executing the computer-readable storage medium and its executable instructions can be found in the description of the aforementioned method embodiments and will not be further elaborated here.
[0109] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, the indirect coupling or communication connection of the device or unit may be electrical, mechanical or other forms.
[0110] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0111] Furthermore, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0112] It should be noted that if a function is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0113] The foregoing description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A multimodal fusion road topology reasoning method, characterized by: The method comprises: Collect historical images and historical point clouds from a crowdsourced dataset, extract lane masks from the historical images using a multimodal model, convert the lane masks into semantic seed points in the historical point clouds, and perform pseudo-labeling based on the semantic seed points. The lane detection model is obtained by model training based on historical images and historical point clouds with pseudo labels; Obtaining lane line information from real-time images and real-time point clouds using the lane detection model, and mapping the lane line information into a bird's-eye view feature map; The topological relationship between lane lines is calculated based on the lane line information in the bird's-eye view feature map to obtain a topological reasoning matrix, and the corresponding road topology reasoning result is obtained based on the graph neural network and the topological reasoning matrix.
2. The multimodal fusion road topology reasoning method according to claim 1 is characterized in that: The step of extracting the lane mask in the historical image using the multimodal model includes: Segmenting the historical image using a multimodal model to generate background semantic masks, and obtaining description text for each background semantic mask; The similarity between the description text and predefined category text is calculated, and the background semantic mask is filtered according to the obtained similarity value to determine the lane mask.
3. The multimodal fusion road topology reasoning method according to claim 1 is characterized in that: The step of converting the lane mask into semantic seed points in the historical point cloud comprises: Obtaining the maximum and minimum values of the lane mask in the pixel coordinate system; performing a boundary shrinkage process on the lane mask based on the maximum value and the minimum value; The lane mask after the boundary shrinkage processing is converted from the historical image to the historical point cloud, and the area corresponding to the lane mask is determined as a semantic seed point.
4. The multimodal fusion road topology reasoning method according to claim 1, characterized in that: The step of setting a pseudo label based on the semantic seed point includes: performing a clustering operation on the semantic seed points in the historical point cloud to determine a plurality of clusters; A bounding box is fitted for each of the clusters, and a pseudo label is obtained based on the bounding box.
5. The multimodal fusion road topology reasoning method according to claim 1, characterized in that: The step of performing model training based on historical images and historical point clouds with pseudo labels to obtain a lane detection model includes: For each pseudo label in the historical image, respectively calculating a distribution constraint score and a meta-shape constraint score of the pseudo label; After normalizing the distribution constraint score and the meta-shape constraint score, they are accumulated according to the set weight factor to obtain a distribution shape score; The pseudo labels are screened based on the distribution shape scores, and a lane detection model is obtained by performing model training based on historical images and historical point clouds corresponding to the screened pseudo labels.
6. The multimodal fusion road topology reasoning method according to claim 1, characterized in that: The real-time image includes images from multiple perspectives; The step of mapping the lane line information into the bird's-eye view feature map includes: Extracting image features of images at each viewing angle in the real-time image, and converting each pixel point in the image at each viewing angle into a three-dimensional point in a ground coordinate system; Project the image features of each perspective into the bird's-eye view space according to the three-dimensional coordinates of the three-dimensional points of the image of each perspective in the ground coordinate system to obtain an initial bird's-eye view feature map; The two-dimensional coordinates and lane line category vector of each lane line in the bird's-eye view space are extracted from the lane line information, and the two-dimensional coordinates and lane line category vector of each lane line are written into the initial bird's-eye view feature map to obtain a final bird's-eye view feature map.
7. The multimodal fusion road topology reasoning method according to claim 6, characterized in that: The step of calculating the topological relationship between lane lines based on the lane line information in the bird's-eye view feature map to obtain a topological reasoning matrix includes: Obtaining a lane query based on lane line information in the final bird's-eye view feature map through a lane deformable decoder, and generating a plurality of directional lane lines based on the lane query; Calculating a lane geometry topology matrix based on the plurality of directional lane lines, and calculating a lane similarity topology matrix based on the lane query; The lane geometry topology matrix and the lane similarity topology matrix are combined to obtain a topology reasoning matrix.
8. The multimodal fusion road topology reasoning method according to claim 7 is characterized in that: The step of calculating a lane geometric topology matrix based on the plurality of directional lane lines and calculating a lane similarity topology matrix based on the lane query comprises: Calculating a lane geometric distance between any two directional lane lines among the plurality of directional lane lines to obtain a lane geometric distance matrix; Obtaining a lane geometric topology matrix based on the lane geometric distance matrix and the learned mapping function; Using two multi-layer perceptrons to encode the lane query respectively to obtain encoding results, and performing inner product calculation on the two encoding results to obtain lane similarity; The lane similarity is mapped to the lane topology to obtain a lane similarity topology matrix.
9. The multimodal fusion road topology reasoning method according to claim 1, characterized in that: The step of obtaining the corresponding road topology reasoning result based on the graph neural network and the topology reasoning matrix includes: Inputting the topological reasoning matrix and the obtained lane query into a graph neural network to aggregate information of adjacent lane lines corresponding to the lane query to obtain an enhanced lane query; The enhanced lane query is processed using the lane header function to obtain the corresponding road topology inference result.
10. A multimodal fusion road topology reasoning system, characterized by: The system comprises: A setting module is used to collect historical images and historical point clouds from a crowdsourced dataset, extract lane masks from the historical images using a multimodal model, convert the lane masks into semantic seed points in the historical point clouds, and perform pseudo-label setting based on the semantic seed points; A training module is used to train a lane detection model based on historical images and historical point clouds with pseudo labels; a mapping module, configured to obtain lane line information from the real-time image and the real-time point cloud using the lane detection model, and map the lane line information into a bird's-eye view feature map; An inference module is used to calculate the topological relationship between lane lines based on the lane line information in the bird's-eye view feature map, obtain a topological inference matrix, and obtain corresponding road topology inference results based on the graph neural network and the topological inference matrix.
Citation Information
Patent Citations
Model pre-training method, model training method, object processing method and device
CN116740498A
Method and system for establishing lane line map and extracting topological structure
CN117671494A
Three-dimensional semantic segmentation method based on camera and laser fusion
CN117934831A
Self-adaptive perception positioning method based on unsupervised fusion BEV
CN119295877A
Road element detection result optimization method, system and device and storage medium
CN119693287A
Cited By
Unmanned aerial vehicle aerial image road detection method based on deep learning
CN121170654A
A deep learning-based unmanned aerial vehicle aerial image road detection method
CN121170654B