A method for reasoning about the topological relationship of a map integrating an SD map

By integrating SD map position coding and 3D lane coordinate generation position embedding, the problem of insufficient lane detection accuracy and real-time in topological map construction is solved, and more accurate topological relationship inference and planning are achieved.

CN119808963BActive Publication Date: 2025-07-18HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510286364.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-18
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing topological map construction methods have shortcomings in lane detection accuracy and real-time performance, and are particularly susceptible to the inherent endpoint shift in lane detection and ignore the geometric characteristics of the lane.

Method used

By integrating SD map position coding and 3D lane coordinates, lane position embedding is generated, and SD map assists in initialization of topological relationships, SD map features and 3D lane features are fused for topological relationship inference, to weaken the impact of detection deviations on topological relationships.

Benefits of technology

It significantly improves the accuracy of topological map construction, especially in occlusion environments, which can provide necessary prior information to ensure more accurate planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808963B_ABST
    Figure CN119808963B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for reasoning about the topological relationship of a map integrating an SD map, belonging to the technical field of 3D reconstruction. The method includes the following steps: obtaining an SD map and performing vectorization processing to obtain an SD map polyline sequence; encoding the SD map polyline sequence to obtain SD map features; using SD map position encoding to adjust the 3D lane coordinates to obtain lane position embeddings; fusing the SD map features and the 3D lane features to perform topological relationship reasoning between lanes. By combining the 3D lane coordinates with the SD map position encoding, the present invention generates lane position embeddings. This process uses the SD map to assist in the initialization of the topological relationship and effectively incorporates the lane geometric information provided by the SD map in the topological reasoning stage. The method not only considers the inherent geometric features of the lanes but also significantly reduces the influence of the inherent endpoint shift in lane detection on the topological relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of map construction, and particularly relates to a method for inferring the topological relationship of a map integrating an SD map. Background Art

[0002] High-definition (HD) maps are the basis for autonomous vehicles (AVs) to perceive the surrounding static environment, providing key information for downstream behavior prediction and motion planning modules. Currently, the method of online high-definition vector maps can directly extract data from in-vehicle sensors to predict high-definition map information, thus realizing online integration with downstream tasks. On this basis, the lanes and the relationships between traffic elements in the driving scene are captured to construct a topological map, showing the drivable routes in autonomous driving and providing clear navigation signals for downstream tasks such as motion prediction and planning.

[0003] The construction of a topological map includes two stages: map element detection and topological relationship inference. The inference performance is limited by the detection accuracy. Therefore, many current panoramic topological map construction methods focus on improving the performance of the detection stage. The map element detection stage realizes end-to-end accurate high-definition map construction based on the DETR architecture for object detection. However, map elements usually have irregular and elongated shapes, and traditional object detection architectures cannot be directly applied to map construction problems. Currently, aiming at the characteristics of map elements, the improvement of the detection architecture is divided into two types of ideas: integrating supplementary information to enhance map perception and improving the DETR object detection architecture for map element features.

[0004] Integrating supplementary information to enhance map perception: The supplementary information includes key prior information of road topology, which can supplement the in-vehicle camera for lane topology inference. For example, using the SD map for query initialization, the query can focus on specific regions of interest to effectively search for map elements and improve the model convergence speed; integrating historical map information into the detection framework to improve the temporal consistency and quality of the vectorized local high-definition map, which performs well in challenging environments such as occlusion. However, the increase in memory and latency costs makes it difficult to integrate in real time into downstream prediction and planning tasks.

[0005] Improvement of the DETR object detection architecture for map element features: Using hierarchical queries to achieve structured and linearized modeling of map elements, that is, performing instance-level matching and point-level matching in sequence. However, in the absence of structural guidance, a large number of vertices need to be predicted to ensure the shape of the elements, posing challenges to the convergence speed and performance.

[0006] In summary, the existing technologies have the following disadvantages:

[0007] The current research on topological map construction mainly focuses on the accuracy and real-time performance of lane detection. It believes that perception is better than reasoning, improves reasoning performance by enhancing lane perception, and directly uses MLP to learn lane topology from lane queries. However, it ignores the inherent geometric features of lanes themselves and is vulnerable to the inherent endpoint shift in lane detection. Summary of the Invention

[0008] An object of an embodiment of the present invention is to provide a method, an electronic device, and a storage medium for inferring map topological relationships by integrating an SD map. For the situation of inherent endpoint shift in lane detection, it uses the position encoding of the SD map to assist in predicting the three-dimensional lane coordinates to generate position embeddings, fine-tune the lane connectivity, and weaken the influence of detection deviation on topological relationships. By using the SD map position prior, it weakens the dependence of topological reasoning on detection, further improves the accuracy of topological map construction, and thus can solve at least one technical problem involved in the background art.

[0009] To solve the above technical problems, the present invention is implemented as follows:

[0010] In a first aspect, an embodiment of the present invention provides a method for inferring map topological relationships by integrating an SD map, including the following steps:

[0011] Step S1: Obtain an SD map and perform vectorization processing to obtain an SD map polyline sequence, specifically including:

[0012] Uniformly sample N fixed numbers from M lanes of the SD map for polyline conversion, denoted as , where both M and N are positive integers, respectively represent the two-dimensional horizontal and vertical axis coordinate points of the SD map lane polyline points;

[0013] Use sine embedding with different frequencies to encode the positions of the polyline points to obtain the SD map position encoding , where , and d is the embedding dimension;

[0014] For the lane types in the SD map, use a one-hot vector of road type labels with dimension K to represent them, and map the lane types to a one-dimensional vector with a length of K. Among them, the element corresponding to the road type is 1, and the rest of the elements are 0;

[0015] Connect the polyline position encoding and the road type one-hot vector to obtain the SD map polyline sequence , where ;

[0016] Step S2: Encode the SD map polyline sequence to obtain the SD map features;

[0017] Step S3: Adjust the 3D lane coordinates using the SD map location encoding to obtain the lane location embedding;

[0018] Step S4: Fuse the SD map features and the 3D lane features to perform lane - to - lane topological relationship reasoning.

[0019] In a second aspect, an embodiment of the present invention provides an electronic device, including:

[0020] At least one processor;

[0021] At least one memory for storing at least one program;

[0022] When the at least one program is executed by the at least one processor, the at least one processor implements the steps of the method described in the first aspect.

[0023] In a third aspect, an embodiment of the present invention provides a readable storage medium, on which a program or an instruction is stored, and when the program or the instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0024] In a fourth aspect, an embodiment of the present invention provides a chip, the chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run a program or an instruction to implement the method described in the first aspect.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0026] 1. By combining the 3D lane coordinates with the SD map location encoding, a lane location embedding is generated. This process uses the SD map to assist in the initialization of the topological relationship and effectively incorporates the lane geometric information provided by the SD map during the topological reasoning stage. This method not only considers the inherent geometric features of the lanes but also significantly reduces the impact of the inherent endpoint shift in lane detection on the topological relationship.

[0027] 2. By fusing the SD map features and the lane features to perform topological relationship reasoning, the SD map contains key information about the road topology and provides valuable supplements for the on - vehicle camera to perform lane topological reasoning. When the merging or exit roads are not visible in the camera image due to occlusion or other reasons, the SD map can provide necessary prior information for downstream planning, thus achieving more accurate planning. Description of the Drawings

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where:

[0029] Figure 1 is a flowchart of a method for reasoning about the map topological relationship integrating the SD map provided by the present invention;

[0030] Figure 2 is one of the schematic hardware structures of the electronic device provided by the present invention;

[0031] Figure 3 is the second schematic hardware structure of the electronic device provided by the present invention. Detailed Embodiments

[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0033] The terms "first", "second", etc. in the description and claims of the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally represents an "or" relationship between the associated objects before and after.

[0034] The following will, with reference to the drawings, describe in detail the method for reasoning about the map topological relationship integrating the SD map provided in the embodiments of the present invention through specific embodiments and their application scenarios.

[0035] Please refer to Figure 1 As shown, the embodiments of the present invention provide a method for reasoning about the map topological relationship integrating the SD map, including the following steps:

[0036] Step S1, obtain the SD map and perform vectorization processing to obtain the SD map polyline sequence, specifically including:

[0037] Uniformly sample N fixed numbers from M lanes of the SD map and perform piecewise linearization, denoted as , both M and N are positive integers, , respectively represent the two-dimensional abscissa and ordinate of N points on the lane after piecewise linearization in the SD map;

[0038] Use sinusoidal embedding encoding with different frequencies to encode the positions of the piecewise points, and obtain the SD map position encoding , where, , d is the embedding dimension;

[0039] For the lane types in the SD map, use the one-hot vector of the road type labels with dimension K to represent them, and map the lane types to a one-dimensional vector with a length of K. Among them, the element corresponding to the road type is 1, and the rest of the elements are 0;

[0040] Connect the piecewise position encoding and the road type one-hot vector to obtain the SD map piecewise sequence , where, ;

[0041] Step S2: Encode the SD map piecewise sequence to obtain the SD map features;

[0042] Step S3: Use the SD map position encoding to adjust the 3D lane coordinates to obtain the lane position embedding;

[0043] Step S4: Fuse the SD map features and the 3D lane features to perform inference on the topological relationship between lanes.

[0044] In step S1, obtain the SD map from OpenStreetMap.

[0045] Step S2 specifically includes:

[0046] Embed the SD map piecewise sequence into a linear layer for encoding, and then use the multi-head self-attention mechanism to extract the SD map features from the SD map , where, , LN represents layer normalization; represents the multi-head self-attention mechanism.

[0047] Step S3 specifically includes:

[0048] For the lane position query obtained in the map element detection stage and the SD map position encoding , use the multi-scale deformable attention mechanism to achieve the purpose of aggregating information using the SD map position at the lane piecewise points, and predict the fused position embedding , where, Represents the multi-scale deformable attention mechanism;

[0049] Step S4 specifically includes:

[0050] Embed the SD map features into the 3D lane query as the fused lane features for the input of the topological reasoning part , and at the same time, in order to fuse and distinguish lane information, embed the predicted fused position and integrate it into the lane query feature to enhance topological modeling, specifically expressed as:

[0051] ;

[0052] ;

[0053] ;

[0054] In the formula, and are the feature and position embeddings of the m-th lane and the n-th lane respectively, where m and n are both positive integers; MLP represents the multi-layer perceptron; represents function.

[0055] As Figure 2 shown, an embodiment of the present invention also provides an electronic device 600. The electronic device 600 includes a processor 601, a memory 602, a program or instruction stored on the memory 602 and executable on the processor 601. When the program or instruction is executed by the processor 601, it implements each process of the above-mentioned method embodiment for inferring the map topological relationship of the fused SD map, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0056] It should be noted that the first electronic device in the embodiment of the present invention includes the above-mentioned mobile electronic device and non-mobile electronic device.

[0057] Figure 3 Is a schematic diagram of the hardware structure of an electronic device for implementing an embodiment of the present invention.

[0058] The electronic device 700 includes but is not limited to: a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, and a processor 710 and other components.

[0059] Those skilled in the art can understand that the electronic device 700 may further include a power source (such as a battery) for powering each component. The power source can be logically connected to the processor 710 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 3 The structure of the electronic device shown in Figure 3 does not limit the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0060] It should be understood that in the embodiments of the present invention, the input unit 704 may include a Graphics Processing Unit (GPU) 7041 and a microphone 7042. The graphics processor 7041 processes the image data of the static image or video obtained by the image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 706 may include a display panel 7061, and the display panel 7061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 707 includes a touch panel 7071 and other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 may include two parts: a touch detection device and a touch controller. The other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, which will not be elaborated here. The memory 709 can be used to store software programs and various data, including but not limited to application programs and operating systems. The processor 710 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 710.

[0061] The embodiments of the present invention also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the map topology relationship reasoning method integrating the SD map, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0062] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc.

[0063] Another embodiment of the present invention provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run programs or instructions to implement each process of the above-described embodiment of the method for inferring the map topological relationship of the fused SD map, and can achieve the same technical effects. To avoid repetition, details are not described herein again.

[0064] It should be understood that the chip mentioned in the embodiments of the present invention may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip.

[0065] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0066] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0067] Another embodiment of the present invention provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run programs or instructions to implement each process of the above-described embodiment of the method for inferring the map topological relationship of the fused SD map, and can achieve the same technical effects. To avoid repetition, details are not described herein again.

[0068] It should be understood that the chip mentioned in the embodiments of the present invention may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0069] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims, and all of them fall within the protection scope of the present invention.

Claims

1. A method for reasoning about the topological relationship of a map integrating an SD map, characterized in that, Including the following steps: Step S1: Obtain an SD map and perform vectorization processing to obtain an SD map polyline sequence, specifically including: For the SD map Uniform sampling of lanes Polyline approximation is performed on a fixed number of points, denoted as , and Both are positive integers, representing the two-dimensional abscissa and ordinate of the points on the polyline of the lane in the SD map respectively; The positions of the broken line points are encoded by sine embeddings with different frequencies to obtain the SD map position encoding , where , is the embedding dimension; For the lane types in the SD map, use the one-hot vector of the road type label with dimension to represent, and map the lane type into a one-dimensional vector with a length of where the element corresponding to the road type is 1 and the rest of the elements are 0. Connect the polyline position encoding with the road type one-hot vector to obtain the SD map polyline sequence , where ; Step S2: Encode the SD map polyline sequence to obtain SD map features; Step S3: Use SD map position encoding to adjust the 3D lane coordinates to obtain lane position embeddings; Step S4: Fuse the SD map features and 3D lane features to perform inter-lane topological relationship reasoning, specifically including: Embed SD map features into the 3D lane query as the fused lane features for the input of the topological reasoning part , and at the same time, in order to fuse the lane geometric information of the SD map, embed the fusion position and integrate it into the lane query features to enhance topological modeling, specifically expressed as: Wherein, , and , are respectively the feature and position embeddings of the -th lane and the -th lane, and are both positive integers; represents a multi-layer perceptron; represents function.

2. The method according to claim 1, wherein In step S1, obtain the SD map from OpenStreetMap.

3. The method according to claim 1, characterized in that, Step S2 specifically includes: Embed the SD map polyline sequence into a linear layer for encoding, and then use the multi-head self-attention mechanism to extract SD map features from the SD map , where , LN represents layer normalization; represents the multi-head self-attention mechanism.

4. The method according to claim 3, characterized in that, Step S3 specifically includes: To utilize the prior information of the lane position in the SD map, encode the SD map position Embed the lane position query in the map element detection stage through the multi-scale deformable attention mechanism , and obtain the fused position embedding , where represents the multi-scale deformable attention mechanism