Driving scene topology reasoning method and electronic device based on machine learning model

Through the cascading loop module structure based on machine learning model, the accuracy of topological reasoning in driving scenarios is enhanced, the problem of inaccurate topological reasoning in the existing technology is solved, and the accuracy and safety of navigation and decision-making of autonomous driving is improved.

CN120258153BActive Publication Date: 2025-08-26NULLMAX INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510732754.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-26
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The topological reasoning results of driving scenarios in the prior art are not accurate enough, which affects the accuracy and safety of autonomous driving navigation and decision-making.

Method used

The cascading loop module structure based on machine learning model is adopted, including the traffic element detection module and the lane centerline detection module. Through the interaction and attention weight enhancement inference of multiple loop modules, the fine-grained point-to-instance relationship and global topological connection in driving scenarios are captured.

Benefits of technology

It improves the accuracy of the topological relationship between the lane center line and between traffic elements and the lane center line during vehicle driving, and improves the accuracy and safety of navigation and decision-making of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258153B_ABST
    Figure CN120258153B_ABST
Patent Text Reader

Abstract

The present application discloses a driving scene topology reasoning method and electronic device based on a machine learning model. The model includes multiple loop modules arranged in a cascade manner. Each loop module includes a traffic element detection module and a lane centerline detection module. The lane centerline detection module includes a lane centerline decoding module and a lane centerline topology reasoning module. In this method, the lane centerline decoding modules at each level obtain updated lane centerline feature information and attention weight information input to the next level lane centerline decoding module and the same level lane centerline topology reasoning module based on the input lane centerline feature information and lane centerline topology reasoning information, thereby realizing cyclic interaction until the topological relationship corresponding to the target driving scene is obtained. The accuracy or precision of the driving scene topological relationship reasoning is improved, thereby improving the accuracy of autonomous driving control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and in particular to a driving scene topology reasoning method and electronic device based on a machine learning model. Background Art

[0002] Driving scene topology reasoning refers to the technology used by autonomous driving systems to infer the scene's topological structure and other topological reasoning results by analyzing the spatial relationships and logical structures of elements such as lanes, intersections, vehicles, and pedestrians within the driving scene. This reasoning can help autonomous driving systems better understand autonomous driving-related information such as road layout and driving paths, thereby improving the accuracy of autonomous driving control aspects such as navigation and decision-making. Therefore, the accuracy or precision of driving scene topology reasoning results will directly or indirectly affect the accuracy of autonomous driving control aspects such as navigation and decision-making, and is crucial to the accuracy and safety of autonomous driving.

[0003] How to improve the accuracy or precision of driving scene topology reasoning results, thereby improving the accuracy of autonomous driving control such as autonomous driving navigation and decision-making, and thus improving the accuracy and safety of autonomous driving, is a direction that is currently being explored in the field. Summary of the Invention

[0004] This application provides a driving scene topology reasoning method based on a machine learning model, which is used to solve the problem in the existing technology of how to improve the accuracy or precision of the driving scene topology reasoning results, thereby improving the accuracy of autonomous driving control such as autonomous driving navigation and decision-making, so as to improve the accuracy and safety of autonomous driving.

[0005] In order to solve the above technical problems, on the first aspect, the implementation method of the present application discloses a driving scene topology reasoning method based on a machine learning model, the machine learning model includes a plurality of loop modules arranged in a cascade manner, each loop module includes a traffic element detection module and a lane centerline detection module, the traffic element detection module includes a traffic element decoding module and a traffic element and lane centerline topology reasoning module, the lane centerline detection module includes a lane centerline decoding module and a lane centerline topology reasoning module, the method includes: the lane centerline decoding modules at each level obtain updated lane centerline feature information and attention weight information based on the input lane centerline feature information and lane centerline topology reasoning information, and input the updated lane centerline feature information to the lane centerline topology reasoning module of the same level, the traffic element and lane centerline topology reasoning module of the same level and the lane centerline decoding module of the next level, and the attention weight information is obtained. The information is input into the lane centerline topology reasoning module of the same level; the lane centerline topology reasoning modules of each level obtain updated lane centerline topology reasoning information based on the input updated lane centerline feature information and attention weight information, and input the updated lane centerline topology reasoning information into the lane centerline decoding module of the next level until the target lane centerline topological relationship is obtained; the traffic element decoding modules of each level obtain updated traffic element feature information based on the input traffic element feature information, and input the updated traffic element feature information into the traffic element decoding module of the next level and the traffic element and lane centerline topology reasoning module of the same level; the traffic element and lane centerline topology reasoning modules of each level obtain updated traffic element and lane centerline topology reasoning information based on the input updated traffic element feature information and the updated lane centerline feature information, until the target traffic element and lane centerline topological relationship is obtained.

[0006] The target lane centerline topological relationship is the topological relationship between lane centerlines corresponding to the target driving scenario, and the target traffic element and lane centerline topological relationship is the topological relationship between lane centerlines and traffic elements corresponding to the target driving scenario.

[0007] The above technical solution employs multiple cascaded loop modules, including a traffic element detection module comprising a traffic element decoding module and a traffic element and lane centerline topology inference module, and a lane centerline detection module comprising a lane centerline decoding module and a lane centerline topology inference module. This allows for the capture of fine-grained point-to-instance relationships and global topological connections based on lane centerline feature information and attention weight information, as well as traffic element feature information and traffic element and lane centerline topology inference information in the driving scene. Furthermore, through loop-enhanced reasoning of the output information from the traffic element decoding module, traffic element and lane centerline topology inference module, and lane centerline decoding module and lane centerline topology inference module, the accuracy or precision of driving scene topological inference results, such as the topological relationships between lane centerlines and the topological relationships between traffic elements and lane centerlines, obtained during vehicle driving, is effectively improved. This improves the accuracy of autonomous driving control aspects such as navigation and decision-making, and thus enhances the accuracy and safety of autonomous driving.

[0008] According to another specific implementation method of the present application, the lane centerline feature information includes point-level lane centerline feature information and instance-level lane centerline feature information corresponding to the target driving scene image, and the updated lane centerline feature information includes updated point-level lane centerline feature information and updated instance-level lane centerline feature information.

[0009] According to another specific implementation of the present application, the attention weight information includes point-to-instance attention weight and instance-to-instance attention weight.

[0010] That is, according to another specific implementation of the present application, the method includes: the lane centerline decoding modules at each level obtain updated point-level lane centerline feature information and updated instance-level lane centerline feature information, as well as point-to-instance attention weights and instance-to-instance attention weights based on the point-level lane centerline feature information, instance-level lane centerline feature information and lane centerline topology reasoning information corresponding to the target driving scene image in the input lane centerline decoding module, input the updated point-level lane centerline feature information and updated instance-level lane centerline feature information to the next-level lane centerline decoding module, input the updated point-level lane centerline feature information and updated instance-level lane centerline feature information, as well as the point-to-instance attention weights and instance-to-instance attention weights to the lane centerline topology reasoning module at the same level, and input the updated point-level lane centerline feature information and updated instance-level lane centerline feature information to the traffic element and lane centerline topology reasoning module at the same level; the lane centerline topology reasoning modules at each level obtain updated point-level lane centerline feature information and updated instance-level lane centerline feature information based on the updated point-level lane centerline feature information and updated instance-level lane centerline feature information in the input lane centerline topology reasoning module. Lane centerline feature information, as well as point-to-instance attention weights and instance-to-instance attention weights, are used to obtain updated lane centerline topology reasoning information, and the updated lane centerline topology reasoning information is input into the next-level lane centerline decoding module until the lane centerline detection result information corresponding to the target driving scene and the topological relationship between the lane centerlines are obtained; traffic element decoding modules at all levels obtain updated traffic element feature information based on the traffic element feature information corresponding to the target driving scene image in the input traffic element decoding module, and input the updated traffic element feature information into the next-level traffic element decoding module, and input the updated traffic element feature information into the same-level traffic element and lane centerline topology reasoning module; traffic element and lane centerline topology reasoning modules at all levels obtain updated traffic element and lane centerline topology reasoning information based on the updated traffic element feature information in the input traffic element and lane centerline topology reasoning module, as well as the updated point-level lane centerline feature information and the updated instance-level lane centerline feature information, until the topological relationship between the lane centerline and traffic elements corresponding to the target driving scene is obtained.

[0011] Using the above technical solution, multiple loop modules are set up in a cascade manner, and a traffic element detection module including a traffic element decoding module and a traffic element and lane centerline topology reasoning module is set in the loop module, and a lane centerline detection module including a lane centerline decoding module and a lane centerline topology reasoning module is set. Among them, the lane centerline decoding modules at each level (i.e., each layer) perform corresponding processing based on the point-level lane centerline feature information, instance-level lane centerline feature information and topology reasoning information corresponding to the target driving scene image input therein, and obtain updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weight and instance-to-instance attention weight, and then input the updated point-level lane centerline feature information and the updated instance-level lane centerline feature information into the next-level lane centerline decoding module for corresponding processing, and input the updated point-level lane centerline feature information and the updated instance-level lane centerline feature information, as well as the point-to-instance attention weight and the instance-to-instance attention weight into the lane centerline topology reasoning module at the same level for corresponding processing, and input the updated point-level lane centerline feature information and the updated instance-level lane centerline feature information into the traffic element and lane centerline topology reasoning module at the same level. Each level (i.e., each layer) of the lane centerline topology inference module processes the updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weights, and instance-to-instance attention weights input to it, obtaining updated lane centerline topology inference information. This updated lane centerline topology inference information is then input into the next-level lane centerline decoding module for corresponding processing until the topological relationship between lane centerlines corresponding to the target driving scene is obtained. Furthermore, each level of the traffic element decoding module obtains updated traffic element feature information based on the traffic element feature information corresponding to the target driving scene image input into the traffic element decoding module. This updated traffic element feature information is then input into the next-level traffic element decoding module and into the same-level traffic element and lane centerline topology inference module. Each level of the traffic element and lane centerline topology inference module obtains updated traffic element and lane centerline topology inference information based on the updated traffic element feature information, updated point-level lane centerline feature information, and updated instance-level lane centerline feature information, until the topological relationship between traffic elements and lane centerlines corresponding to the target driving scene is obtained.

[0012] This system captures fine-grained point-to-instance relationships and global topological connections based on point-level and instance-level lane centerline feature information, point-to-instance and instance-to-instance attention weights, and traffic element feature information and traffic element and lane centerline topology reasoning information corresponding to the target driving scene image. Furthermore, through the recurrent enhanced reasoning of the output information from the traffic element decoding module and the traffic element and lane centerline topology reasoning module, and the lane centerline decoding module and the lane centerline topology reasoning module, the accuracy or precision of the driving scene topology reasoning results, such as the topological relationships between lane centerlines and the topological relationships between traffic elements and lane centerlines, obtained during vehicle driving, is effectively improved. This improves the accuracy of autonomous driving control aspects such as navigation and decision-making, and ultimately the accuracy and safety of autonomous driving.

[0013] According to another specific implementation of the present application, the machine learning model also includes an initial processing module, and the method also includes obtaining point-level lane centerline feature information, instance-level lane centerline feature information, and lane centerline topology reasoning information input into the first-level lane centerline decoding module in the following manner: the initial processing module performs perspective feature conversion processing on the target driving scene image to obtain target perspective feature information, performs query processing based on the target perspective feature information, and obtains point-level lane centerline feature information, instance-level lane centerline feature information, and lane centerline topology reasoning information input into the first-level lane centerline decoding module.

[0014] According to another specific implementation of the present application, the target perspective feature information is bird's-eye view perspective feature information.

[0015] By adopting the above technical solution, the target driving scene image is subjected to image perspective feature conversion processing through the initial processing module to obtain target perspective feature information (for example, converted into bird's-eye view perspective feature information). Query processing is performed based on the target perspective feature information to obtain point-level lane centerline feature information, instance-level lane centerline feature information, and topological reasoning information, which are input into the first-level lane centerline decoding module. This can more accurately reflect the local and global information of the target driving scene image, and effectively increase the accuracy of the traffic element detection result information, lane centerline detection result information, topological relationship between lane centerlines, and topological relationship between lane centerlines and traffic elements obtained during vehicle driving.

[0016] According to another specific implementation of the present application, the method also includes performing second perspective feature conversion processing on the target driving scene image through the initial processing module to obtain second perspective feature information, performing traffic element query processing on the second perspective feature information to obtain traffic element feature information input into the first-level traffic element decoding module.

[0017] The second viewing angle characteristic information may be, for example, perspective view (PV) characteristic information.

[0018] According to another specific implementation of the present application, the method further includes performing feature conversion processing on the perspective view feature information to obtain bird's-eye view feature information.

[0019] According to another specific implementation of the present application, multiple loop modules include a first loop module and a second loop module, the first loop module is the previous level module of the second loop module, the first loop module includes a first traffic element detection module and a first lane centerline detection module, the second loop module includes a second traffic element detection module and a second lane centerline detection module, the first traffic element detection module includes a first traffic element decoding module and a first traffic element and lane centerline topology reasoning module, the first lane centerline detection module includes a first lane centerline decoding module and a first lane centerline topology reasoning module, the second traffic element detection module includes a second traffic element decoding module and a second traffic element and lane centerline topology reasoning module, and the second lane centerline detection module includes a second lane centerline decoding module. The method includes: the first lane centerline decoding module obtains second point-level lane centerline feature information, second instance-level lane centerline feature information, first point-to-instance attention weight and first instance-to-instance attention weight based on the first point-level lane centerline feature information, first instance-level lane centerline feature information and first lane centerline topology reasoning information input into the first lane centerline decoding module, and combines the second point-level lane centerline feature information and the second instance-level lane centerline feature information into the first lane centerline decoding module. The feature information is input into the first lane centerline topology reasoning module, the first traffic element and lane centerline topology reasoning module and the second lane centerline decoding module, and the first point-to-instance attention weight and the first instance-to-instance attention weight are input into the first lane centerline topology reasoning module; the first lane centerline topology reasoning module obtains the second lane centerline topology reasoning information based on the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight and the first instance-to-instance attention weight input into the first lane centerline topology reasoning module, and inputs the second lane centerline topology reasoning information into the second centerline decoding module; the first traffic element decoding module obtains the second traffic element feature information based on the first traffic element feature information input into the first traffic element decoding module, and inputs the second traffic element feature information into the second traffic element decoding module and the first traffic element and lane centerline topology reasoning module; the first traffic element and lane centerline topology reasoning module obtains the second traffic element and lane centerline topology reasoning information based on the second traffic element feature information input into the first traffic element and lane centerline topology reasoning module, the second point-level lane centerline feature information and the second instance-level lane centerline feature information.

[0020] The first traffic element characteristic information is also the first traffic element characteristic information corresponding to the target driving scene.

[0021] The above technical solution achieves a precise and semantically rich understanding of topological structures through fine-grained point-level interactions and high-level instance-level relationships. A recurrent interaction mechanism facilitates iterative refinement, where attention weights from the lane centerline decoding module (also known as the detector) are fed into the lane centerline topology inference module, while topological inference information from the lane centerline topology inference module (also known as instance topological relationships) is fed back into the lane centerline decoding module, creating a mutually reinforcing loop to improve performance. This recurrently enhanced reasoning of the output information from the lane centerline decoding module and the lane centerline topology inference module effectively increases the accuracy of the topological relationships between lane centerlines and between traffic elements and lane centerlines obtained during vehicle driving.

[0022] According to another specific implementation of the present application, each lane centerline decoding module includes a point-level processing unit, an instance-level processing unit, a point-to-instance integration unit, and a relationship decoder unit. The first lane centerline decoding module obtains the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight based on the first point-level lane centerline feature information, the first instance-level lane centerline feature information, and the first lane centerline topology reasoning information input into the first lane centerline decoding module, including: the point-level processing unit obtains the second point-level lane centerline feature information and the first instance-level lane centerline feature information according to the first point-level lane centerline feature information and the preset point-to-point relationship matrix. The second point-level lane centerline feature information is obtained; the relation decoder unit obtains point-to-instance relationship information and instance-to-instance relationship information based on the first lane centerline topology reasoning information; the point-to-instance integration unit obtains the first point-to-instance attention weight based on the second point-level lane centerline feature information, the first instance-level lane centerline feature information and the point-to-instance relationship information, and determines the distance transformation centerline mask based on the second point-level lane centerline feature information; the instance-level processing unit obtains the second instance-level lane centerline feature information and the first instance-to-instance attention weight based on the first instance-level lane centerline feature information, the distance transformation centerline mask and the instance-to-instance relationship information.

[0023] Using the above technical solution, a point-level processing unit, an instance-level processing unit, a point-to-instance integration unit and a relationship decoder unit are set in the centerline decoding module, and the point-level processing unit obtains the second point-level lane centerline feature information based on the first point-level lane centerline feature information and the preset point-to-point relationship matrix; the relationship decoder unit obtains the point-to-instance relationship information and the instance-to-instance relationship information based on the first lane centerline topology reasoning information; the point-to-instance integration unit obtains the first point-to-instance attention weight based on the second point-level lane centerline feature information, the first instance-level lane centerline feature information and the point-to-instance relationship information, and determines the distance transformation centerline mask based on the second point-level lane centerline feature information; the instance-level processing unit obtains the second instance-level lane centerline feature information and the first instance-to-instance attention weight based on the first instance-level lane centerline feature information, the distance transformation centerline mask and the instance-to-instance relationship information. The system effectively increases the accuracy of the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight obtained by corresponding processing based on the first point-level lane centerline feature information, the first instance-level lane centerline feature information, and the first topological reasoning information input into the first lane centerline decoding module, thereby increasing the accuracy of the traffic element detection result information, the lane centerline detection result information, the topological relationship between lane centerlines, and the topological relationship between lane centerlines and traffic elements obtained during vehicle driving, thereby improving driving safety.

[0024] According to another specific implementation of the present application, the distance transformation centerline mask can be obtained by the following formula:

[0025]

[0026] Among them, b is the pixel point corresponding to the target perspective feature information, C is the lane centerline point c corresponding to the second point-level lane centerline feature information i The set of components, D(b) is the Euclidean distance, L width is the width of the lane centerline, M(b) is the activation function, and RD(b) is the distance transformed centerline mask.

[0027] Using the above technical solution, the distance transform facilitates the extraction of position information, while the rasterized distance transform reduces the learning difficulty of the model and accelerates convergence. Therefore, using a distance transform centerline mask (also known as a rasterized distance transform mask) can better extract image position information (such as the centerline position), further increasing the accuracy of the obtained second-instance-level lane centerline feature information and the first-instance-to-instance attention weights.

[0028] According to another specific implementation method of the present application, the point-level processing unit includes a line perception attention sub-unit and a cross attention sub-unit. The point-level processing unit obtains the second point-level lane centerline feature information based on the first point-level lane centerline feature information and the preset point-to-point relationship matrix, including: the line perception attention sub-unit performs corresponding processing on the first point-level lane centerline feature information and the preset point-to-point relationship matrix to obtain the intermediate point-level lane centerline feature information; the cross attention sub-unit performs corresponding processing on the intermediate point-level lane centerline feature information to obtain the second point-level lane centerline feature information.

[0029] By adopting the above technical solution and setting the line-aware attention sub-unit, it is possible to effectively promote intra-instance interaction between point-level queries and increase the accuracy of the obtained second point-level feature information.

[0030] According to another specific implementation method of the present application, the point-to-instance integration unit includes an integrated attention sub-unit, and the point-to-instance integration unit obtains a first point-to-instance attention weight based on the second point-level lane centerline feature information, the first instance-level lane centerline feature information and the point-to-instance relationship information, including: the integrated attention sub-unit performs corresponding processing on the second point-level lane centerline feature information, the first instance-level lane centerline feature information and the point-to-instance relationship information to obtain the first point-to-instance attention weight.

[0031] Using this technical solution, the integrated attention sub-unit enables each point-level query to selectively focus on relevant instance-level features, enriching its representation with both local and global context. This integrated attention sub-unit further enhances instance-aware feature learning and reasoning, ensuring that point-level queries are enriched with both local location details and global semantic context. This effectively increases the accuracy of the first point-to-instance attention weights obtained.

[0032] According to another specific implementation method of the present application, the instance-level processing unit includes a topology-aware attention subunit and a mask attention subunit. The instance-level processing unit obtains the second instance-level lane centerline feature information and the first instance-to-instance attention weight based on the first instance-level lane centerline feature information, the distance transformation centerline mask and the instance-to-instance relationship information, including: the topology-aware attention subunit performs corresponding processing on the first instance-level lane centerline feature information and the distance transformation centerline mask to obtain the intermediate instance-level lane centerline feature information; the mask attention subunit performs corresponding processing on the intermediate instance-level lane centerline feature information and the instance-to-instance relationship information to obtain the second instance-level lane centerline feature information and the first instance-to-instance attention weight.

[0033] The above technical solution effectively increases the accuracy of the obtained second-instance-level lane centerline feature information and the first-instance-to-instance attention weight.

[0034] According to another specific implementation method of the present application, each lane centerline topology reasoning module includes a point-to-instance topology reasoning unit, an instance-to-instance topology reasoning unit and a connection unit. The first lane centerline topology reasoning module performs corresponding processing based on the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight and the first instance-to-instance attention weight input into the first lane centerline topology reasoning module to obtain the second lane centerline topology reasoning information, including: the point-to-instance topology reasoning unit obtains the point-to-instance topology reasoning information according to the second point-level lane centerline feature information, the first point-to-instance attention weight and the second instance-level lane centerline feature information; the instance-to-instance topology reasoning unit obtains the instance-to-instance topology reasoning information according to the second instance-level lane centerline feature information and the first instance-to-instance attention weight; the connection unit obtains the second lane centerline topology reasoning information according to the point-to-instance topology reasoning information and the instance-to-instance topology reasoning information.

[0035] By adopting the above technical solution, a point-to-instance topology reasoning unit, an instance-to-instance topology reasoning unit (also known as a dual prediction branch) and a connection unit are set in the lane centerline topology reasoning module. The point-to-instance topology reasoning unit combines explicit feature correlation and latent dependency for robust topology reasoning. The instance-to-instance topology reasoning unit uses instance-level queries and their self-attention weights for similarity calculation, effectively integrating instance-level features and their hidden dependencies, realizing comprehensive topological reasoning and increasing the accuracy of the obtained second lane centerline topology reasoning information.

[0036] According to another specific implementation method of the present application, the point-to-instance topology reasoning unit includes at least one first connection subunit and multiple first neural network subunits. The point-to-instance topology reasoning unit obtains point-to-instance topology reasoning information based on the second point-level lane centerline feature information, the first point-to-instance attention weight and the second instance-level lane centerline feature information, including: each first neural network subunit performs corresponding processing on the second point-level lane centerline feature information, the first point-to-instance attention weight and the second instance-level lane centerline feature information to obtain first intermediate feature information; the first connection subunit performs connection processing on the first intermediate feature information to obtain point-to-instance topology reasoning information.

[0037] By adopting the above technical solution, the accuracy of the obtained point-to-instance topological reasoning information is increased by extracting the association relationship through each attention-based first neural network sub-unit.

[0038] According to another specific implementation of the present application, the instance-to-instance topology reasoning unit includes at least one second connection subunit and multiple second neural network subunits. The instance-to-instance topology reasoning unit obtains instance-to-instance topology reasoning information based on the second instance-level lane centerline feature information and the first instance-to-instance attention weight, including: the second neural network subunit performs corresponding processing on the second instance-level lane centerline feature information and the first instance-to-instance attention weight to obtain second intermediate feature information; the second connection unit performs connection processing on the second intermediate feature information to obtain instance-to-instance topology reasoning information.

[0039] By adopting the above technical solution, the accuracy of the obtained instance-to-instance topological reasoning information is increased by extracting the association relationship through each attention-based second neural network sub-unit.

[0040] According to another specific implementation method of the present application, the initial processing module includes a shared backbone network unit and a view conversion unit. The initial processing module performs image perspective feature conversion processing on the target driving scene image to obtain target perspective feature information, including: the shared backbone network unit performs feature extraction processing on the target driving scene image to obtain third intermediate feature information; the view conversion unit performs feature conversion processing on the third intermediate feature information to obtain target perspective feature information.

[0041] By adopting the above technical solution, by sharing the backbone network unit and the view conversion unit, the features of the target driving scene image can be accurately extracted to obtain more accurate target perspective feature information.

[0042] On the second aspect, the implementation method of the present application also discloses a machine learning model, which includes multiple loop modules arranged in a cascade manner, each loop module includes a traffic element detection module and a lane centerline detection module, the traffic element detection module includes a traffic element decoding module and a traffic element and lane centerline topology reasoning module, and the lane centerline detection module includes a lane centerline decoding module and a lane centerline topology reasoning module. Among them, the lane centerline decoding modules at each level obtain updated lane centerline feature information and attention weight information based on the input lane centerline feature information and lane centerline topology reasoning information, and input the updated lane centerline feature information to the lane centerline topology reasoning module at the same level, the traffic element and lane centerline topology reasoning module at the same level, and the lane centerline decoding module at the next level, and input the attention weight information to the lane centerline topology reasoning module at the same level; the lane centerline topology reasoning modules at each level obtain updated lane centerline topology reasoning information based on the input updated lane centerline feature information and attention weight information, and input the updated lane centerline feature information to the lane centerline topology reasoning module at the same level. The centerline topology reasoning information is input into the next-level lane centerline decoding module until the target lane centerline topological relationship is obtained; the traffic element decoding modules at all levels obtain updated traffic element feature information based on the input traffic element feature information, and input the updated traffic element feature information into the next-level traffic element decoding module and the same-level traffic element and lane centerline topological reasoning module; the traffic element and lane centerline topological reasoning modules at all levels obtain updated traffic element and lane centerline topological reasoning information based on the input updated traffic element feature information and the updated lane centerline feature information, until the target traffic element and lane centerline topological relationship is obtained.

[0043] On the third aspect, the implementation method of the present application also discloses a vehicle control method, which includes: obtaining target driving scene topology reasoning result information, such as the topological relationship between lane centerlines and the topological relationship between lane centerlines and traffic elements, through the above-mentioned driving scene topology reasoning method based on the machine learning model; controlling vehicle driving according to the target driving scene topology reasoning result information.

[0044] In a fourth aspect, the implementation method of the present application also discloses an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores a computer program; the processor executes the computer program stored in the memory, so that the electronic device implements the driving scene topology reasoning method based on the machine learning model provided in any one of the implementation methods of the first aspect above.

[0045] In a fifth aspect, the implementation method of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the driving scene topology reasoning method based on the machine learning model provided in any one of the implementation methods of the first aspect above.

[0046] In a sixth aspect, an implementation of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements a driving scene topology reasoning method based on a machine learning model as provided in any one of the implementations of the first aspect above.

[0047] It can be understood that the beneficial effects of the second to sixth aspects mentioned above can also be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 This is a schematic diagram of the principle of a centerline detection and topology reasoning method in the prior art;

[0049] Figure 2 This is a schematic diagram of the principle of another centerline detection and topology reasoning method in the prior art;

[0050] Figure 3 This is a schematic diagram of the structure of a machine learning model provided in an embodiment of the present application;

[0051] Figure 4 This is a schematic structural diagram of a lane centerline decoding module provided in an embodiment of the present application;

[0052] Figure 5 This is a schematic diagram of the structure of a point-level processing unit provided in an embodiment of the present application;

[0053] Figure 6 This is a schematic diagram of the structure of a point-to-instance integration unit provided in an embodiment of the present application;

[0054] Figure 7 This is a schematic diagram of the structure of an instance-level processing unit provided in an embodiment of the present application;

[0055] Figure 8 This is a schematic diagram of the structure of a lane centerline topology reasoning module provided in an embodiment of the present application;

[0056] Figure 9 This is a schematic diagram of the structure of a point-to-instance topology reasoning unit provided in an embodiment of the present application;

[0057] Figure 10 This is a schematic diagram of the structure of an instance-to-instance topology reasoning unit provided in an embodiment of the present application;

[0058] Figure 11 This is a schematic diagram of the principle of a circular reasoning framework corresponding to a machine learning model provided in an embodiment of the present application;

[0059] Figure 12 This is a schematic diagram of the principle of a centerline detection and topology reasoning method provided in an embodiment of the present application;

[0060] Figure 13 This is a schematic diagram of the structure of the inference framework corresponding to a machine learning model provided in an embodiment of the present application;

[0061] Figure 14 This is a schematic diagram of the structure of another reasoning framework corresponding to a machine learning model provided in an embodiment of the present application;

[0062] Figure 15 1 is a schematic structural diagram of a centerline decoder at layer 1 in an inference framework corresponding to a machine learning model provided in an embodiment of the present application;

[0063] Figure 16 1 is a schematic diagram of the structure of a topological reasoning module in the first layer of a cyclic reasoning framework corresponding to a machine learning model provided in an embodiment of the present application;

[0064] Figure 17 is a schematic diagram showing a performance comparison between TopoHR and other state-of-the-art methods in the OpenLane-V2 subset A benchmark test provided by an embodiment of the present application;

[0065] Figure 18 is a schematic diagram of a performance comparison of TopoHR provided by an embodiment of the present application with other state-of-the-art methods in the OpenLane-V2 subset B benchmark test;

[0066] Figure 19 is a schematic diagram of an ablation study result represented by a centerline mask (Mask GT) provided in an embodiment of the present application;

[0067] Figure 20 is a schematic diagram of an ablation study result on the impact of hierarchical attention provided by an embodiment of the present application;

[0068] Figure 21 is a schematic diagram of an ablation study result on the effectiveness of hierarchical topology reasoning provided by an embodiment of the present application;

[0069] Figure 22 1 is a schematic diagram of an ablation study result of adaptive topology loss in TopoHR provided by an embodiment of the present application;

[0070] Figure 23 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0071] As mentioned earlier, the accuracy of the driving scene topology reasoning results will directly or indirectly affect the accuracy of autonomous driving control such as navigation and decision-making, and is crucial to the accuracy and safety of autonomous driving.

[0072] The driving scene topology reasoning results can include topological relationships between driving scene traffic elements and lane centerlines, as well as topological relationships between lane centerlines. Prior art methods for determining these topological relationships typically involve first acquiring a driving scene image of a vehicle in motion, then performing instance-level learning for centerline detection on the image to obtain instance-level information corresponding to the centerline of the driving scene image. Topological reasoning then relies on a sequential module consisting of simplified multi-layer perceptron (MLP) layers to determine the topological relationships between driving scene traffic elements and lane centerlines, as well as the topological relationships between lane centerlines, during the vehicle's driving process. This method determines the topological relationships between driving scene traffic elements and lane centerlines, as well as the topological relationships between lane centerlines, based solely on instance-level information, and employs a pipelined approach to query and process driving scene image information to determine the topological relationships between driving scene traffic elements and lane centerlines, as well as the topological relationships between lane centerlines. This method suffers from inaccurate topological relationships between driving scene traffic elements and lane centerlines, as well as the topological relationships between lane centerlines, during vehicle driving. This impacts the accuracy and safety of autonomous driving.

[0073] Specifically, traditional lane detection methods and online mapping technologies focus on geometric accuracy but fail to capture scene topology. Although high-precision electronic maps (HD maps) can provide topological information about scene topology, they suffer from freshness and scalability issues.

[0074] like Figure 1 As shown in Figure 2, a sequential pipeline approach is provided. In this pipeline, the centerline decoder (i.e., lane centerline detection module, also known as the centerline detector) first extracts instance-level centerline representations, and the topology reasoning module then analyzes these extracted centerline instances.

[0075] Further, if Figure 2As shown in

[15] , most centerline detection methods typically use instance-level query representation in the transformer decoder, defining the task as point set prediction, curve parameter estimation, or binary segmentation (such as 0 / 1 segmentation). For example, TopoNet represents the centerline as a point set and refines it using a scene graph neural network (SGNN). In contrast, TopoBDA models centerline detection as a Bezier curve parameter estimation task with a deformable attention mechanism. Although vectorization-based methods employ point-level attention in local geometric modeling, they often ignore global contextual patterns. Alternatively, segmentation-based methods, such as TopoMask, combine direction prediction to derive vectorization results from segmentation output.

[0076] However, centerlines are inherently invisible, making it challenging to accurately extract their features through direct segmentation modeling. Current implementations either use segmentation as auxiliary supervision or rely on post-processing to generate vectorized point sets, which limits the full utilization of the rich information contained in the segmentation results. Based on instance-level queries of centerline detectors, topology reasoning modules typically apply MLP functions to infer topological relationships.

[0077] For example, TopoNet uses three MLP layers to reason about topological relationships through precise centerline queries. TopoMLP emphasizes the importance of detector performance in a cascaded structure and enhances topological reasoning by integrating position embeddings into the MLP layers. Similarly, TopoLogic improves reasoning performance by incorporating centerline geometric priors to connect two centerlines through their start and end points.

[0078] While these methods have achieved promising results in topological reasoning for driving scenes, significant gaps remain in the accuracy of topological relationship prediction. These methods typically rely on a cascaded architecture, where centerline detection and relationship reasoning modules are optimized separately, resulting in inconsistent feature representations. Furthermore, using a simplified prediction head (typically consisting of only a three-layer MLP) fails to fully capture the complex spatial dependencies inherent in urban road networks.

[0079] Based on this, one implementation of the present application provides a driving scene topology inference method based on a machine learning model. This method can determine target driving scene topology inference result information, including, for example, topological relationships between lane centerlines and between lane centerlines and traffic elements, based on point-level and instance-level lane centerline feature information corresponding to the driving scene image in the image scene, as well as point-to-instance and instance-to-instance attention weights. This method captures fine-grained point-to-instance relationships and global topological connections. Furthermore, by looping and enhancing inference through the output information of the lane centerline decoding module and the lane centerline topology inference module, the accuracy or precision of the driving scene topology inference results is effectively improved, thereby improving the accuracy of autonomous driving control aspects such as autonomous driving navigation and decision-making, and further improving the accuracy and safety of autonomous driving.

[0080] Next, with reference to the accompanying drawings, the steps and advantages of the driving scene topology reasoning method based on the machine learning model provided by this application are described in detail.

[0081] In one implementation of this application, Figure 3 As shown, the machine learning model includes multiple loop modules such as loop module 1, loop module 2...loop module n, which are arranged in a cascade manner (for example, n levels are set), wherein each loop module includes a traffic element detection module and a lane centerline detection module, the traffic element detection module includes a traffic element decoding module and a traffic element and lane centerline topology reasoning module, and the lane centerline detection module includes a lane centerline decoding module and a lane centerline topology reasoning module.

[0082] Therefore, the driving scene topology reasoning method based on the machine learning model includes: the lane centerline decoding modules at each level perform corresponding processing based on the point-level lane centerline feature information, instance-level lane centerline feature information and lane centerline topology reasoning information corresponding to the target driving scene image input into the lane centerline decoding module, and obtain updated point-level lane centerline feature information and updated instance-level lane centerline feature information, as well as point-to-instance attention weights and instance-to-instance attention weights, and input the updated point-level lane centerline feature information and updated instance-level lane centerline feature information into the next-level lane centerline decoding module for corresponding processing, input the updated point-level lane centerline feature information and updated instance-level lane centerline feature information, as well as the point-to-instance attention weights and instance-to-instance attention weights into the lane centerline topology reasoning module at the same level for corresponding processing, and input the updated point-level lane centerline feature information and updated instance-level lane centerline feature information into the traffic element and lane centerline topology reasoning module at the same level.

[0083] The lane centerline topology reasoning modules at each level perform corresponding processing based on the updated point-level lane centerline feature information and updated instance-level lane centerline feature information, as well as the point-to-instance attention weight and instance-to-instance attention weight input into the lane centerline topology reasoning module to obtain updated lane centerline topology reasoning information. The updated lane centerline topology reasoning information is input into the next-level lane centerline decoding module for corresponding processing until a topological relationship between lane centerlines corresponding to, for example, the target driving scenario is obtained.

[0084] Traffic element decoding modules at each level obtain updated traffic element feature information based on the traffic element feature information corresponding to the target driving scene image input into the traffic element decoding module, input the updated traffic element feature information into the next level traffic element decoding module, and input the updated traffic element feature information into the traffic element and lane centerline topology reasoning module at the same level.

[0085] The traffic element and lane centerline topology reasoning modules at all levels obtain updated traffic element and lane centerline topology reasoning information based on the updated traffic element feature information in the input traffic element and lane centerline topology reasoning module, as well as the updated point-level lane centerline feature information and the updated instance-level lane centerline feature information, until the topological relationship between the traffic elements and lane centerlines corresponding to the target driving scenario is obtained.

[0086] In other embodiments of the present application, the output of the lane centerline decoding module can also be understood as lane centerline detection result information corresponding to the target driving scenario. That is, through the cyclic processing of the traffic element detection module and the lane centerline detection module arranged in a cascaded manner, the present application can obtain the target driving scenario topology reasoning result information. The target driving scenario topology reasoning result information can include, for example, lane centerline detection result information, topological relationships between lane centerlines, traffic element detection result information, and topological relationships between lane centerlines and traffic elements.

[0087] Point-level lane centerline feature information and instance-level lane centerline feature information can be point-level and instance-level feature information of important elements in the driving scene, such as the lane centerline information during vehicle driving, or information of other important elements.

[0088] That is, in one implementation of the present application, lane centerline decoding modules at all levels can also obtain lane centerline detection result information, such as lane centerline position, based on the point-level lane centerline feature information and instance-level lane centerline feature information corresponding to the target driving scene image input into the lane centerline decoding module.

[0089] In one implementation of the present application, traffic element decoding modules at all levels can also obtain traffic element detection result information, such as the coordinates, width, and height of the bounding box of the traffic element, based on the traffic element feature information corresponding to the target driving scene image input into the traffic element decoding module.

[0090] The driving scene topology reasoning method based on the machine learning model provided in the present application sets up multiple loop modules arranged in a cascade manner, and sets up a traffic element detection module including a traffic element decoding module and a traffic element and lane centerline topology reasoning module in the loop module, and sets up a lane centerline detection module including a lane centerline decoding module and a lane centerline topology reasoning module. First, the lane centerline decoding modules at each level perform corresponding processing based on the point-level lane centerline feature information, instance-level lane centerline feature information and topology reasoning information input therein to obtain updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weight and instance-to-instance attention weight. The updated point-level lane centerline feature information and the updated instance-level lane centerline feature information are then input into the next-level lane centerline decoding module for corresponding processing, the updated point-level lane centerline feature information and the updated instance-level lane centerline feature information, as well as the point-to-instance attention weight and the instance-to-instance attention weight are input into the lane centerline topology reasoning module at the same level for corresponding processing, and the updated point-level lane centerline feature information and the updated instance-level lane centerline feature information are input into the traffic element and lane centerline topology reasoning module at the same level. Each level of lane centerline topology inference module processes the updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weights, instance-to-instance attention weights, and updated traffic element feature information input to it, obtaining updated lane centerline topology inference information. This updated lane centerline topology inference information is then input into the next-level lane centerline decoding module for further processing until lane centerline detection result information corresponding to the target driving scene and the topological relationship between lane centerlines are obtained. Furthermore, each level of traffic element decoding module obtains updated traffic element feature information based on the traffic element feature information corresponding to the target driving scene image input into the traffic element decoding module. This updated traffic element feature information is then input into the next-level traffic element decoding module, and the updated traffic element feature information is then input into the traffic element and lane centerline topology inference module at the same level. The traffic element and lane centerline topology reasoning modules at all levels obtain updated traffic element and lane centerline topology reasoning information based on the input updated traffic element feature information, updated point-level lane centerline feature information, and updated instance-level lane centerline feature information, until the topological relationship between the traffic elements and lane centerlines corresponding to the target driving scenario is obtained.

[0091] This system captures fine-grained point-to-instance relationships and global topological connections based on point-level and instance-level lane centerline feature information, point-to-instance and instance-to-instance attention weights, and traffic element feature information and traffic element and lane centerline topology inference information corresponding to the target driving scene image. Furthermore, through the recurrent enhanced reasoning of the output information from the traffic element decoding module, the traffic element and lane centerline topology inference module, and the lane centerline decoding module and lane centerline topology inference module, the accuracy or precision of the topological relationships between lane centerlines and between traffic elements and lane centerlines obtained during vehicle driving is effectively improved, thereby improving the accuracy of autonomous driving control aspects such as navigation and decision-making, and ultimately the accuracy and safety of autonomous driving.

[0092] In one implementation of this application, Figure 3 As shown, the machine learning model also includes an initial processing module, and the method also includes obtaining point-level lane centerline feature information, instance-level lane centerline feature information, and lane centerline topology reasoning information that are input into the first-level lane centerline decoding module in the following manner: the initial processing module performs perspective feature conversion processing on the target driving scene image to obtain target perspective feature information, and performs query processing based on the target perspective feature information to obtain point-level lane centerline feature information, instance-level lane centerline feature information, and lane centerline topology reasoning information that are input into the first-level lane centerline decoding module.

[0093] The target viewing angle feature information may be, for example, bird's-eye view feature information obtained by extracting bird's-eye view features from the target driving scene image.

[0094] In one implementation of the present application, the method also includes performing second-perspective feature conversion processing on the target driving scene image through an initial processing module to obtain second-perspective feature information, performing traffic element query processing on the second-perspective feature information to obtain traffic element feature information input into the first-level traffic element decoding module.

[0095] The second viewing angle characteristic information may be, for example, perspective view (PV) characteristic information.

[0096] In one implementation of the present application, multiple loop modules include a first loop module and a second loop module, the first loop module is the previous level module of the second loop module, the first loop module includes a first traffic element detection module and a first lane centerline detection module, the second loop module includes a second traffic element detection module and a second lane centerline detection module, the first traffic element detection module includes a first traffic element decoding module and a first traffic element and lane centerline topology reasoning module, the first lane centerline detection module includes a first lane centerline decoding module and a first lane centerline topology reasoning module, the second traffic element detection module includes a second traffic element decoding module and a second traffic element and lane centerline topology reasoning module, and the second lane centerline detection module includes a second lane centerline decoding module. The method includes: a first lane centerline decoding module performs corresponding processing based on the first point-level lane centerline feature information, the first instance-level lane centerline feature information and the first lane centerline topology reasoning information input into the first lane centerline decoding module to obtain the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight and the first instance-to-instance attention weight, inputs the second point-level lane centerline feature information and the second instance-level lane centerline feature information into the second lane centerline decoding module for corresponding processing, inputs the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight and the first instance-to-instance attention weight into the first lane centerline topology reasoning module for corresponding processing, and inputs the second point-level lane centerline feature information and the second instance-level lane centerline feature information into the first traffic element and lane centerline topology reasoning module; the first lane centerline topology reasoning module performs corresponding processing based on the first lane The second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight in the centerline topology reasoning module are processed accordingly to obtain the second lane centerline topology reasoning information, and the second lane centerline topology reasoning information is input into the second lane centerline decoding module; the first traffic element decoding module obtains the second traffic element feature information based on the first traffic element feature information corresponding to the target driving scene image input into the first traffic element decoding module, inputs the second traffic element feature information into the second traffic element decoding module, and inputs the second traffic element feature information into the first traffic element and lane centerline topology reasoning module; the first traffic element and lane centerline topology reasoning module obtains the second traffic element and lane centerline topology reasoning information based on the second traffic element feature information input into the first traffic element and lane centerline topology reasoning module, as well as the second point-level lane centerline feature information and the second instance-level lane centerline feature information.

[0097] Furthermore, the multiple loop modules may further include a third, fourth, etc. loop module, repeating the above process until the topological relationship between the traffic elements and the lane centerlines and the topological relationship between the lane centerlines corresponding to the target driving scene are obtained.

[0098] In one implementation of this application, Figure 4 As shown, each lane centerline decoding module includes a point-level processing unit, an instance-level processing unit, a point-to-instance integration unit, and a relation decoder unit. The first lane centerline decoding module performs corresponding processing based on the first point-level lane centerline feature information, the first instance-level lane centerline feature information and the first lane centerline topology reasoning information input into the first lane centerline decoding module to obtain second point-level lane centerline feature information, second instance-level lane centerline feature information, a first point-to-instance attention weight and a first instance-to-instance attention weight, including: a point-level processing unit obtains the second point-level lane centerline feature information according to the first point-level lane centerline feature information and a preset point-to-point relationship matrix; a relationship decoder unit obtains point-to-instance relationship information and instance-to-instance relationship information according to the first lane centerline topology reasoning information; a point-to-instance integration unit obtains the first point-to-instance attention weight according to the second point-level lane centerline feature information, the first instance-level lane centerline feature information and the point-to-instance relationship information, and determines the distance transformation centerline mask according to the second point-level lane centerline feature information; and an instance-level processing unit obtains the second instance-level lane centerline feature information and the first instance-to-instance attention weight according to the first instance-level lane centerline feature information, the distance transformation centerline mask and the instance-to-instance relationship information.

[0099] In one implementation of the present application, during the detection of the target driving scene centerline, the segmentation-level expression of the target centerline can be obtained by calculating the distance transform centerline mask formula:

[0100]

[0101] Among them, b is the pixel point corresponding to the target perspective feature information, C is the lane centerline point c corresponding to the second point-level lane centerline feature information i The set of components, D(b) is the Euclidean distance, L width is the width of the lane centerline, M(b) is the activation function, and RD(b) is the distance transformed centerline mask.

[0102] In one implementation of this application, Figure 5As shown, the point-level processing unit includes a line-perception attention subunit and a cross-attention subunit. The point-level processing unit obtains second-point-level lane centerline feature information based on the first-point-level lane centerline feature information and a preset point-to-point relationship matrix. The process includes: the line-perception attention subunit processes the first-point-level lane centerline feature information and the preset point-to-point relationship matrix to obtain intermediate-point-level lane centerline feature information; and the cross-attention subunit processes the intermediate-point-level lane centerline feature information to obtain second-point-level lane centerline feature information.

[0103] In one implementation of this application, Figure 6 As shown, the point-to-instance integration unit includes an integrated attention sub-unit. The point-to-instance integration unit obtains a first point-to-instance attention weight based on the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information. The integrated attention sub-unit performs corresponding processing on the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information to obtain the first point-to-instance attention weight.

[0104] In one implementation of this application, Figure 7 As shown, the instance-level processing unit includes a topology-aware attention subunit and a mask attention subunit. The instance-level processing unit obtains second instance-level lane centerline feature information and first instance-to-instance attention weights based on the first instance-level lane centerline feature information, the distance-transformed centerline mask, and the instance-to-instance relationship information. This includes: the topology-aware attention subunit performs corresponding processing on the first instance-level lane centerline feature information and the distance-transformed centerline mask to obtain intermediate instance-level lane centerline feature information; and the mask attention subunit performs corresponding processing on the intermediate instance-level lane centerline feature information and the instance-to-instance relationship information to obtain second instance-level lane centerline feature information and the first instance-to-instance attention weights.

[0105] In one implementation of this application, Figure 8As shown, the lane centerline topology reasoning module includes a point-to-instance topology reasoning unit, an instance-to-instance topology reasoning unit, and a connection unit. The first lane centerline topology reasoning module obtains second lane centerline topology reasoning information based on the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight input into the first lane centerline topology reasoning module. The first lane centerline topology reasoning module obtains the second lane centerline topology reasoning information based on the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information; the instance-to-instance topology reasoning unit obtains the instance-to-instance topology reasoning information based on the second instance-level lane centerline feature information and the first instance-to-instance attention weight; and the connection unit obtains the second lane centerline topology reasoning information based on the point-to-instance topology reasoning information and the instance-to-instance topology reasoning information.

[0106] In one implementation of this application, Figure 9 As shown, the point-to-instance topology reasoning unit includes a first connection subunit and multiple first neural network subunits. The point-to-instance topology reasoning unit obtains point-to-instance topology reasoning information based on the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information, including: each first neural network subunit performs corresponding processing on the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information to obtain first intermediate feature information; the first connection subunit performs connection processing on the first intermediate feature information to obtain point-to-instance topology reasoning information. Of course, it can also include multiple first connection subunits to connect different first intermediate feature information in pairs or in other ways.

[0107] In one implementation of this application, Figure 10 As shown, the instance-to-instance topology reasoning unit includes a second connection subunit and multiple second neural network subunits. The instance-to-instance topology reasoning unit obtains instance-to-instance topology reasoning information based on the second instance-level lane centerline feature information and the first instance-to-instance attention weight, including: each second neural network subunit performs corresponding processing on the second instance-level lane centerline feature information and the first instance-to-instance attention weight to obtain second intermediate feature information; the second connection subunit performs connection processing on the second intermediate feature information to obtain instance-to-instance topology reasoning information. Of course, it can also include multiple second connection subunits to connect different second intermediate feature information in pairs or in other ways.

[0108] In one implementation of the present application, the initial processing module includes a shared backbone network unit and a view conversion unit. The initial processing module performs image perspective feature conversion processing on a target driving scene image to obtain target perspective feature information, including: the shared backbone network unit performs feature extraction processing on the target driving scene image to obtain third intermediate feature information; and the view conversion unit performs feature conversion processing on the third intermediate feature information to obtain target perspective feature information.

[0109] The machine learning model provided by the implementation of this application can be obtained based on model training and used to implement driving scene topology reasoning through the above-mentioned method. As previously mentioned, performing driving scene topology reasoning based on this machine learning model can effectively improve the accuracy or precision of the driving scene topology reasoning results, thereby improving the accuracy of autonomous driving control such as autonomous driving navigation and decision-making, and further improving the accuracy and safety of autonomous driving.

[0110] This application presents a machine learning model and a driving scenario topology inference method based on this model, also known as TopoHR. This is a novel end-to-end topology inference framework based on a recurrent interaction structure and a hierarchical centerline representation. Unlike existing cascaded architectures that rely heavily on detector performance, the detector (i.e., the lane centerline decoding module) and lane centerline topology inference module in TopoHR can enhance each other.

[0111] like Figure 11 As shown, this application introduces a recurrent reasoning framework for machine learning models, also known as the TopoHR framework, in which the self-attention weights (such as point-to-instance attention weights or instance-to-instance attention weights) from the centerline decoder (i.e., the lane centerline decoding module) in the detector (i.e., the lane centerline detection module) are input as feedforward signals into the topology reasoning module (i.e., the lane centerline topology reasoning module), while the topological relationships (such as lane centerline topological relationships) from the topology reasoning module are fed back to the centerline decoder in the detector. In the detector, multiple centerline representations (e.g., point queries, instance queries, and semantic instances) are integrated, and a dedicated module is designed for feature interaction.

[0112] like Figure 12 As shown, first, in order to capture local and global features simultaneously, this application designs a point-to-instance interaction module to facilitate information exchange between point-level queries and instance-level queries. Then, this application introduces a rasterized distance transform (R-DT) segmentation method to extract the spatial information of the centerline, which is then used for instance-level query interaction through a masked attention module. For the topological reasoning module, this application designs a hierarchical topological reasoning module to capture fine-grained point-to-instance and global instance-to-instance topological connections, which are fused to predict the final topological relationship.

[0113] Based on the driving scene topology inference method based on machine learning models provided in this application, the performance of the proposed method is evaluated on the OpenLane-V2 dataset. The results show that under similar configurations, the proposed method has achieved significant improvements compared with existing methods. Specifically, the proposed method has a significant improvement in the TOP ll reached 37.7, in the TOP ll It outperforms the previous state-of-the-art model by 10.1%.

[0114] TopoHR, provided in this application, is a novel end-to-end framework for centerline detection and topological reasoning, in which the detection and reasoning modules iteratively interact to cyclically enhance each other's performance. A hierarchical representation of centerlines is introduced, in which multi-level representations are seamlessly integrated and fused in a hierarchical centerline decoder. In addition, a hierarchical topological reasoning module is proposed to capture fine-grained point-to-instance relationships and global instance-to-instance connections, thereby generating accurate topological results. The method shows good performance on the OpenLane-V2 dataset, far surpassing previous state-of-the-art models in topological reasoning of centerlines.

[0115] Next, the preparatory work related to this application will be described.

[0116] First, online map preparation. Online map solutions can be broadly categorized into rasterization-based and vectorization-based approaches. HDMapNet pioneered segmentation-based approaches by fusing multi-view camera images and LiDAR point clouds using a bird's-eye view (BEV) encoder-decoder architecture to predict instance-level map elements as semantic masks.

[0117] To reduce the dependence on multimodal data, MGMap proposes a camera-only framework with multi-granularity decoding to achieve hierarchical segmentation of map components, Mask2Map enhances point-level feature extraction by incorporating deformable attention into instance segmentation, BLOS-BEV rasterizes navigation data and integrates it with BEV features, P-MapNet utilizes prior information of SD maps and HD maps for pre-fusion, and adopts Mask Auto-Fusion-Encoder (MAE) as a refinement module to address occlusion and artifact problems.

[0118] For vectorization-based methods, VectorMapNet introduces a two-stage approach that combines polyline detection with geometric refinement. MapTR proposes a unified permutation-equivalent modeling strategy to achieve end-to-end vectorized map learning without explicit point ordering constraints. HIMap unifies hierarchical queries for joint detection of lanes and traffic elements, demonstrating robust generalization across different datasets. PivotNet uses pivot points to model lane lines, enhancing the representation of geometric structure. StreamMapNet improves temporal consistency by combining historical BEV features through cross-frame attention.

[0119] Next comes topology detection and reasoning. Most topology reasoning methods use a sequential pipeline, where a centerline detector first extracts centerlines, followed by a topology reasoning module. For centerline detection, CenterLineDet represents centerlines as vertices and leverages temporal feature fusion for multi-camera perception. In addition to point set formulations, centerline detection also incorporates methods such as instance segmentation and curve parameterization.

[0120] For example, STSU proposed to use Bezier curves for centerline detection, and TopoMask enhanced the centerline detection by introducing instance segmentation and adding direction prediction to parse the vectorized results. On this basis, TopoMaskV2 combined instance-level Bezier curves with segmentation modeling and used the centerline segmentation results to optimize Bezier curve predictions. TopoBDA further extended this approach by introducing a deformable attention mechanism based on Bezier curve modeling and introducing auxiliary losses through instance segmentation. TopoFormer uses the geometric distance between centerlines to guide global information aggregation and models a reasonable road structure under the counterfactual intervention layer. In addition, SMERF and TopoSD proposed to use SD-Map to enhance the centerline detector.

[0121] For topological reasoning, TopoNet first uses two MLP layers to reduce the embedding dimensionality of each instance, then sends the concatenated features to another MLP and uses sigmoid activation to predict the relationship between them. TopoMLP emphasizes the "detect first, reason later" strategy and designs a high-performance detector and MLP that incorporates implicit positional embedding features.

[0122] TopoLogic proposes a topology-centric loss function that explicitly preserves lane connections in the segmentation output. LaneSegNet embeds the topological affinity field into instance segmentation to achieve real-time lane map extraction. Topo2Seq converts the graph topology relationship of the traffic scene into a serial representation and uses a hierarchical transformer to achieve multi-scale topological reasoning.

[0123] Finally, there is the distance transform. Distance Transform (DT) has been widely used in various computer vision tasks. In semantic segmentation, DT guides the segmentation of tubular structures by utilizing geometric features.

[0124] Similarly, in object detection, DATNet enhances the model's ability to perceive spatial information by predicting the distance to the nearest instance boundary for each pixel. Extended to the 3D domain, 3DSC utilizes the Euclidean 3D distance transform for action recognition to better capture motion changes. DT has also shown significant practicality in geospatial applications. For example, POSTPROC combines DT-based edge maps with CNN to improve topological continuity in aerial road detection. Recently, MapVR applied differentiable rasterization to vectorized output and used distance changes on raster maps to achieve accurate geometric-aware supervision without introducing additional computation during inference.

[0125] This application applies DT to enhance centerline representation and improve topological reasoning in driving scenarios.

[0126] Next, combine the attached Figure 13-16 , further explain and illustrate the method provided in this application.

[0127] In one implementation of this application, Figure 13 As shown, the TopoHR framework proposed in this application includes a recurrent architecture, including an L-layer (i.e., L-level) traffic element decoder (i.e., traffic element decoding module) and a topology reasoning module (i.e., traffic element and lane centerline topology reasoning module), as well as an L-layer hierarchical centerline decoder and a hierarchical topology reasoning module. The L-layer traffic element decoder and topology reasoning module constitute the aforementioned traffic element detection module, and the L-layer hierarchical centerline decoder and hierarchical topology reasoning module constitute the aforementioned lane centerline detection module. The recurrent architecture consists of the following core components:

[0128] (1) The initial processing module includes a backbone network (i.e., a shared backbone network unit, Backbone) and a view conversion module (i.e., PV-to-BEV, view conversion unit). The backbone network is used to first convert the multi-view image (i.e., the target driving scene image, Input) into PV features as the input of the traffic element decoder. The PV features are then converted into BEV features through the view conversion module (i.e., PV-to-BEV, view conversion unit) as the input of the hierarchical centerline decoder.

[0129] Further, if Figure 14 As shown in Figure 2, the backbone network and view conversion module can also be called the BEV feature extractor.

[0130] (2) Hierarchical Centerline Decoder (i.e., lane centerline decoding module), which introduces a hierarchical query representation Q hcl ∈ R N×(P+1)×C , where N, P, and C represent the maximum number of centerline instances, the number of points per centerline, and the number of hidden channels, respectively. It consists of two interrelated components: point-level query Q p ∈R P×C and instance-level query Q i ∈R C , enhancing point- and instance-level feature representations through a series of attention mechanisms and point-to-instance integrators. The hierarchical centerline decoder includes the aforementioned point-level module (Pts-level Module), the point-to-instance integrator (Pts-Ins Integrator), the relation encoder, the instance-level module (Ins-level Module), and the R-DT mask. These modules work together to produce the lane centerline output (i.e., lane centerline detection result information, such as classification, segmentation, and coordinates).

[0131] (3) Hierarchical Topology Module (i.e., Lane Centerline Topology Reasoning Module), including Point-to-Instance Topology Reasoning Module (Pts-Ins Topo Module) and Instance-to-Instance Topology Reasoning Module (Ins-Ins Topo Module), performs comprehensive topological reasoning by integrating updated hierarchical queries and point-to-instance and instance-to-instance. p2i ∈R (N×P)×N , W i2i ∈R N×N Through fine-grained point-level interactions and high-level instance-level relationships, it achieves a precise and semantically rich understanding of topology. A recurrent interaction mechanism facilitates iterative refinement, where attention weights from the detector are fed into the topology reasoning module, and instance topology relationships from the reasoning module are fed back into the detector, creating a mutually reinforcing cycle that improves performance and ultimately results in the centerline topology output (i.e., lane centerline topology relationships).

[0132] Unlike uniform segmentation labels for all positive samples, the distance transform encodes spatial proximity by mapping Euclidean distances. This approach effectively addresses the challenge of centerline segmentation without relying on specific visual features. In our approach, rather than using the distance transform directly, we convert the distance field into a structured grid rasterization to improve computational efficiency and feature representation.

[0133] Formally, given the center line C={c1, c2, ..., c P}, where c i =(x i ,y i ) represents the vectorized point set, that is, the coordinates of the points on the center line. First, calculate the distance from each pixel point b to the nearest center line point c in the BEV feature (i.e., target perspective feature information). i ∈C. The distance value is limited to the lane width L width Then, these values ​​are normalized to the range [0,1] and rasterized into 11 categories according to the step size △ = 0.1, generating the final rasterized distance transform centerline mask RD (b). The specific process can be achieved by the following formula:

[0134]

[0135] Where b is the pixel point in the BEV image, C is the given center line point c i The set of components, D(b) is the Euclidean distance, L width is the width of the lane centerline (i.e., lane width), M(b) is the activation function, and RD(b) is the distance transform centerline mask (i.e., rasterized distance transform centerline mask).

[0136] The rasterized distance transform mask is applied to the mask attention module and participates in instance matching and loss calculation. To this end, the rasterized distance transform mask can better extract the position information of the centerline and is easier to converge than the standard regression method.

[0137] (4) Traffic element detection module, including the Traffic Element Decoder (i.e., the Traffic Element Decoder) and the Topology Reasoning Module (i.e., the Traffic Element and Lane Centerline Topology Reasoning Module). The Traffic Element Decoder processes the PV feature information corresponding to the target driving scene image to obtain the Traffic Element Output (i.e., the Traffic Element detection result information, such as the traffic element classification (class), the coordinates (x, y) and width (w) and height (h) of the bounding box (bbox) surrounding the traffic element). The Topology Reasoning Module also obtains the Traffic Element and Centerline Topology Output (i.e., the topological relationship between the lane centerline and the traffic element). In addition, there can be multiple bounding boxes, so they can be referred to as bboxes.

[0138] Further, if Figure 15 As shown in the figure, the input and output of the first layer of the hierarchical centerline decoder in the TopoHR framework. Figure 15 As shown in part (a), a line-aware attention module (also known as a line-aware attention sub-unit) is applied in the point-level processing unit (Pts-level Module) to promote intra-instance interactions between point-level queries. Specifically, for each set of point-level queries (e.g., 11 point-level queries per instance), the line-aware attention operation is limited to operating within each centerline instance. To enforce this constraint, the point-to-point relationship matrix (P2PRelation) M is introduced. p2p ∈R (N×P)×(N×P) , which is used as an attention mask when computing line-aware attention. This mask effectively limits cross-instance interactions by focusing attention on point-level queries within the same centerline instance, maintaining instance-specific feature learning and preventing information exchange between different centerlines. In one implementation of this application, the point-level module also introduces a cross-attention module (i.e., a cross-attention subunit, Cross-Attention).

[0139] Furthermore, referring to relational DETR, feature representation is enhanced by modeling the relationships between object instances using intermediate attention maps in the decoder. This relational learning framework is naturally consistent with the core requirements of centerline detection and topological reasoning tasks. The inherent structural dependencies and topological constraints in centerlines make them particularly suitable for such relation-aware feature learning.

[0140] In order to use topological reasoning to improve centerline detection, e.g. Figure 15 As shown, this application introduces a relation encoder to process the topology prediction from the previous layer (i.e., the l-1 layer topology reasoning information). This application uses the MLP layer to generate two different topological relations: (1) point-to-instance topological relation (i.e., point-to-instance relation, P2IRelation) p2i ∈R (N×P)×N and (2) topological relationships between instances (i.e., instance-to-instance relationships, I2I Relations) M i2i ∈R N ×N The point-to-instance and instance-to-instance relationships are then integrated into attention masks and fed into the integrator attention module (i.e., the integrator attention subunit, Integrator Attention) and the topology-aware attention module (i.e., the topology-aware attention subunit, Topo-Aware Attention). This application establishes a mutually reinforcing loop between the centerline detection and topology reasoning modules—one component iteratively improves the performance of the other through information exchange and joint optimization.

[0141] The hierarchical framework promotes multi-scale interaction between point-by-point geometric features and instance-level semantics in centerline modeling. Through the local-global attention architecture, feature representation is iteratively improved by fusing positional details with semantic context, thus achieving coherent feature propagation across hierarchical spaces. Figure 15 As shown in part (b) (point-to-instance integration unit, Pts-InsIntegrator), the hierarchical integration attention module operates in two cross-level steps. Given the Figure 15 The updated point-level query (i.e., the second point-level lane centerline feature information) Q p,l ∈R N×P×C and the instance-level query Q generated from the (l-1)th layer (i.e., the first instance-level lane centerline feature information) i,l-1 ∈R NxC , the integrator attention adopts the cross attention mechanism. Specifically, the point-level query Q p,l acts as Q, while instance-level query Q i,l-1 Acts as a key-value. The specific formula for point-to-instance integration is as follows:

[0142]

[0143]

[0144] Among them, M p2iRepresenting point-to-instance relationships, the attention module enables each point-level query to selectively focus on relevant instance-level features, thereby enriching its representation with local and global context. Then, the learnable coefficient W Agg ∈R P Aggregate the improved point-level query into an updated instance-level query:

[0145]

[0146] To further enhance instance-aware feature learning and reasoning, the updated instance-level query is then input into the instance-level module ( Figure 15 Part (c) shows the instance-level processing unit (Ins-level Module). This approach ensures that point-level queries simultaneously enrich local location details and global semantic context, while instance-level queries are dynamically updated by aggregating and refining point-level features.

[0147] In one implementation of this application, Figure 15 As shown in part (c), the instance-level module includes a topology-aware attention subunit (Topo-Aware Attention) and a masked-Attention subunit (Masked-Attention). It processes the updated instance-level query, R-DT mask, and instance-to-instance relationship to obtain the updated instance-level query and instance-to-instance attention weights.

[0148] Existing topological reasoning methods mainly rely on instance-level query interactions modeled by simplified MLP layers. This application proposes a hierarchical representation paradigm that explicitly integrates point-to-instance and instance-to-instance topological relationships. This application emphasizes the hierarchical establishment of topological relationships: when considering two centerlines C i and C j When there is a topological relationship between the two centerlines, this relationship is reflected not only in the instance-level representation of the two centerlines, but also in the C i Point-level representation and C j This is crucial between instance-level expressions of

[0149] Specifically, if Figure 16As shown in the figure, the hierarchical module proposed in this application includes two prediction branches: (1) The instance-to-instance topology module (i.e., instance-to-instance topology reasoning unit, Ins-Ins Topology Reasoning) uses dual MLP encoding to calculate similarity (such as inner product calculation through the second connection subunit) using instance-level queries and their self-attention weights, and extracts association relationships through attention-based MLP. (2) The point-to-instance topology module (i.e., point-to-instance topology reasoning unit, Pts-Ins Topology Reasoning) combines explicit feature correlation and latent dependency for robust topology reasoning. This design achieves comprehensive topology reasoning by effectively integrating instance-level features and their hidden dependencies.

[0150] Specifically, the instance-to-instance topology module performs a concatenation (e.g., inner product operation) based on the instance-to-instance attention weights and the instance-level query to obtain an updated instance-level query. The point-to-instance topology module performs a concatenation (e.g., inner product operation) based on the point-to-instance attention weights, the instance-level query, and the click query to obtain an updated second instance-level query. Finally, hierarchical topological reasoning is performed based on the updated instance-level query and the updated second instance-level query to obtain topological reasoning information.

[0151] The aforementioned instance-to-instance topology reasoning process is shown in the following formula:

[0152]

[0153]

[0154]

[0155] Point-to-instance topology prediction follows a similar workflow, applying an average operation along the point dimension to generate a point-to-instance topology prediction:

[0156]

[0157]

[0158]

[0159] The final hierarchical topology reasoning result comes from two hierarchical components T i2i and T p2i results.

[0160] Next, the training loss of this application is explained.

[0161] This paper proposes an adaptive topological loss for topological reasoning supervision, replacing the traditional focal loss used in TopoLogic. The framework adopts a dynamic weighting strategy based on reparameterized cross entropy, where the weights of negative samples follow an exponential scaling e λneg·xi , x i represents the predicted positive probability, while the positive sample keeps a fixed weight λ pos This mechanism creates an adaptive gradient modulation that proportionally amplifies the penalty for negative samples that exhibit high confidence scores, effectively mitigating false positive predictions while maintaining topological consistency in the feature space. Overall, the training loss formula can be written as:

[0162]

[0163] Among them, the center line detection loss It consists of a focus loss for centerline instance classification and an L1 loss for regression of vectorized centerline results. The dice loss and cross entropy loss are combined to guide the learning of instance-level features based on the centerline mask of the rasterized distance transform. is the proposed adaptive topology loss.

[0164] Next, the advantages of the method provided in this application are explained in combination with specific experimental data.

[0165] The TopoHR presented in this application is evaluated on the OpenLane-V2 benchmark, which is a comprehensive dataset that integrates Argoverse 2 and nuScenes. The dataset contains 2000 scenes, divided into subsets A and B, with multi-view Figure 2 Hz images and annotations of 3D centerlines, traffic elements, and their topological relationships. Subset A contains seven camera views, while subset B contains six camera views. This application uses the following official evaluation metrics. Evaluation metrics include:

[0166] DET l (average Frechet distance across matching thresholds), DET t (traffic element similarity based on IoU), TOP ll (centerline topological matrix similarity) and TOP lt (centerline traffic element topological similarity), the OLS metric calculates the average of these multi-task metrics.

[0167] This application uses ResNet50 as the backbone and a Feature Pyramid Network (FPN) to obtain multi-scale features with an input resolution of 1550×2048. 200 centerline and 100 traffic element queries are initialized for detection and topological relationship reasoning. Following the BEVFormer, this application projects image features into a predefined BEV space with a grid resolution of 200×100. In the centerline detector, the regression head consists of a 3-layer MLP with LayerNorm and ReLU activations, outputting an 11×3 3D position offset for each centerline.

[0168] The topological reasoning method proposed in this application focuses on modeling the topological relationship of centerlines. The settings in TopoLogic are followed to detect traffic elements and calculate the topological relationship between centerlines and traffic elements. The TopoHR model is trained on 8 NVIDIA 4090 GPUs with a total batch size of 8 for 24 iterations. For optimization, the AdamW optimizer is used with an initial learning rate of , with a weight decay of 0.01.

[0169] Our proposed TopoHR achieves competitive performance compared to previous state-of-the-art methods such as TopoMaskV2 and TopoBDA in either centerline detector or traffic element detector without using any additional training tricks or optimizations (one-to-many centerline query; different backbones and denoising strategies for traffic element detection).

[0170] Specific as Figure 17 Figure 2 shows the performance comparison of TopoHR presented in this application with other state-of-the-art methods on the OpenLane-V2 subset A benchmark. The best result is highlighted in bold, the second best result is underlined, and the third best result is highlighted in italics.

[0171] Other state-of-the-art methods include: Method 1 (Map Transformer (MapTR), corresponding to Conference 1 (International Conference on Learning Representations, ICLR 2023), Method 2 (TopoNet (TopoNet), corresponding to Conference 2 (Arxiv 2023), Method 3 (SMERF (SD Map Encoding Representation from Transformers), corresponding to Conference 3 (IEEE International Conference on Robotics and Automation, ICRA 2024), Method 4 (TopoLogic, corresponding to Conference 4 (Conference and Workshop on Neural Information Processing Systems, NeurIPS 2024), Method 5 (TopoFormer, corresponding to Conference 5 (Arxiv 2024)). 2024)), Method 6 (i.e., TopoMaskV2, the corresponding conference is Conference 5 (i.e., Arxiv2024)), Method 7 (i.e., TopoBDA, the corresponding conference is Conference 5 (i.e., Arxiv 2024)).

[0172] The Standard Definition Map (SDMap) is a one-to-many query in the Lane Centerline Detection Branch, and Object-to-Map (O2M) is an additional one-to-many query in the Lane Centerline Detection Branch. Dynamic Threshold Detection (i.e., Differentiable Binarization (DB)) and Optimized Clustering (K-means improvement), or DBKB, are different backbones in the Traffic Element Detection Branch. DN is the denoising strategy for the Traffic Element Detection Branch. This Solution-1 method refers to the TopoHR method, and this Solution-2 method refers to the TopoHR-L method, where -L expands the number of TopoHR centerline queries to 600.

[0173] The TopoHR-L provided by this application has the highest classification accuracy of TOP ll and TOP ltThe best results were 37.7 and 38.2 respectively, and the second best result was 51.6 on the Ordinary Least Squares (OLS). The TopoHR provided in the evaluation parameter detection task (DET) DET l and TOP ll The second best results were 36.0 and 32.7 respectively.

[0174] TopoHR achieves state-of-the-art (SOTA) performance without specialized training strategies: outperforming TopoFormer by 3.7 in OLS and 1.0 in TOP ll It is 8.2 higher than the average (increased by 34.0%), and even in the TOP ll The scaled variant TopoHR-L (with 600 centerline queries) further improves the TOP ll Improved by 10.1 (36.6%). Excellent DET achieved by TopoMaskV2 and TopoBDA t and OLS scores are mainly due to its advanced traffic element decoder, which adopts DAB-DETR and DINO (Distillation in Noise) architecture.

[0175] like Figure 18 Figure 2. Performance comparison of our TopoHR presented in this application with other state-of-the-art methods on the OpenLane-V2 subset B benchmark. The best result is highlighted in bold, the second best result is underlined, and the third best result is highlighted in italics.

[0176] Other state-of-the-art methods include: Method 2 (i.e., TopoNet, the corresponding conference is Conference 2 (i.e., Arxiv2023)), Method 4 (i.e., TopoLogic, the corresponding conference is Conference 4 (i.e., NeurIPS 2024)), Method 5 (i.e., TopoFormer, the corresponding conference is Conference 5 (i.e., Arxiv 2024)), Method 6 (i.e., TopoMaskV2, the corresponding conference is (i.e., Arxiv 2024)), and Method 7 (i.e., TopoBDA, the corresponding conference is Conference 5 (i.e., Arxiv 2024)).

[0177] The TopoHR-L provided by this application evaluates the parameters TOP ll and TOP ltThe best results were 38.9 and 28.2 respectively. The proposed solution-1 (i.e., TopoHR) achieved the second best result of 51.9 on the evaluation parameter OLS. TopoHR performed well in topological reasoning and ll The score is 32.7 points, which is better than method 6 (i.e. TopoMaskV2). The model of this solution-2 (i.e. TopoHR-L) has significantly improved over TopoBDA in terms of evaluation parameters TOP ll The increase of 4.9 indicates the robustness and adaptability of the TopoHR framework provided by this application on different datasets.

[0178] Furthermore, this application conducted an ablation study to evaluate the impact of the key components proposed in this application on the OpenLane-V2 Subset A dataset.

[0179] The first is the design of the hierarchical centerline representation. Figure 19 As shown in the figure, the ablation study of the centerline mask (Mask GT) representation is performed, where MA represents instance-level masked attention; Distance Transform (DT) represents distance transform mask; R-DT represents rasterizing distance transform mask. The ablation study results are as follows: (1) Adding 0 / 1 instance segmentation loss to evaluate the parameters TOP ll and TOP lt The evaluation results of the proposed method are improved by 2.0 and 1.8 (rows 1-2); (2) instance-level masked attention enhances query feature interaction (row 3); compared with the baseline, the rasterized distance transform GT mask provided by this application is in TOP ll and TOP lt We achieve maximum gains of 4.8 and 3.5 on the dataset, respectively, confirming the effectiveness of the hierarchical representation in capturing both local and global structures.

[0180] Then there is the effect of hierarchical attention relation modeling. Figure 20 As shown, an ablation study verifies the impact of hierarchical attention: the baseline model (without relation encoder, masked attention or topology module) implements e.g. DET l =28.6 and TOP ll =26.3. Introducing point-level constrained attention (i.e., point-to-point, P2P) (enabling intra-instance interaction) will transform DET l Improved to 33.3 (+4.7 improvement). Gradually integrate the point-to-instance (i.e. point-to-instance, P2I) and instance-level (i.e. instance-to-instance, I2I) topological relations generated by the relation encoder to DET l The improvement is further increased to 33.6 (5.0 over the baseline), demonstrating incremental accuracy gains through layer-wise inference.

[0181] Secondly, the design of hierarchical topological reasoning. Figure 21 As shown in Figure 2, the effectiveness of hierarchical topological reasoning is demonstrated: integrating instance-level attention in the topology head (TopoHead) and attention weight (Attn Weight) can improve DET l and TOP ll Indicators. Hierarchical modeling combines instance-level and point-level interactions in DET l and TOP ll Achieve gains of 1.4 and 0.7 over the baseline, respectively, confirming the benefits of hierarchical modeling for centerline detection and topology.

[0182] Finally, the impact of adaptive topology loss. The effectiveness of adaptive topology loss (ATL) in TopoHR is shown in the following example: Figure 22 As shown in the figure, the adaptive topology loss (ATL) proposed in this application is used to replace the standard focal loss (FC). ll and TOP lt The evaluation results improved by 0.4 and 1.3 respectively.

[0183] It should be noted that Figures 17-20 In , “√” indicates that the target branch was experimented with using the corresponding existing method. For example, Figure 17 The “√” in the first row indicates that the MapTR method published in the ICLR 2023 conference was used to experiment on the O2M branch in the lane centerline detection branch and obtain evaluation indicators.

[0184] TopoHR, proposed in this application, is an end-to-end topological reasoning method that integrates cycle detector topological interaction and hierarchical centerline representation. This architecture overcomes the limitations of sequential pipelines by enabling the co-evolution of detection modules and topological reasoning modules, while the hierarchical representation fuses multi-level centerline position features through rasterized distance transforms. The unified hierarchical topological module simultaneously captures fine-grained point-to-instance relationships and global topological connectivity. Extensive experiments show that TopoHR achieves excellent topological reasoning performance on the OpenLane-V2 benchmark.

[0185] Therefore, the TopoHR proposed in this application effectively improves the accuracy or precision of driving scene topology reasoning results. Therefore, when applied to autonomous driving systems, based on more accurate or higher-precision driving scene topology reasoning results, the accuracy of autonomous driving control such as autonomous driving navigation and decision-making can be improved, thereby improving the accuracy and safety of autonomous driving.

[0186] Furthermore, the implementation method of the present application also proposes a vehicle control method, which includes: obtaining the target driving scene topology reasoning result information through the above-mentioned TopoHR, and controlling the vehicle driving according to the target driving scene topology reasoning result information to achieve more accurate and safe autonomous driving.

[0187] See Figure 23 , Figure 23 The figure shows a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 23 As shown, the electronic device may include: a transceiver 121 , a processor 122 , and a memory 123 .

[0188] Processor 122 executes the computer program / instructions stored in the memory, causing processor 122 to implement a portion of the technical solutions of the driving scenario topology inference method based on a machine learning model in the above-mentioned embodiment. Processor 122 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NPU), etc.; it can also be a digital data processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0189] The memory 123 is connected to the processor 122 via a system bus and communicates with the processor 122. The memory 123 is used to store computer program instructions.

[0190] By way of example and not limitation, the memory 123 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 123 may include removable or non-removable (or fixed) media. Where appropriate, the memory 123 may be internal or external to the integrated gateway device. In a specific embodiment, the memory 123 is a non-volatile solid-state memory. In a specific embodiment, the memory 123 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0191] The transceiver 121 may be used to obtain tasks to be executed and configuration information of the tasks to be executed.

[0192] The implementation method of the present application also provides a chip for running computer instructions / programs, which is used to execute the technical solution of the above-mentioned driving scene topology reasoning method based on the machine learning model.

[0193] The implementation method of the present application also provides a computer-readable storage medium, which stores computer instructions / programs. When the computer instructions / programs are run on the processor of an electronic device, the processor of the electronic device executes the technical solution of the driving scene topology reasoning method based on the machine learning model.

[0194] The implementation method of the present application also provides a computer program product, which includes a computer program / instructions stored in a computer-readable storage medium. At least one processor can read the computer program / instructions from the computer-readable storage medium. When the at least one processor executes the computer program / instructions, the technical solution of the driving scene topology reasoning method based on the machine learning model can be implemented.

[0195] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.

[0196] It should be noted that, in addition to the implementation of the present application described in the above-mentioned specific embodiments, those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. Although the description of the present application is introduced in conjunction with the preferred embodiment, this does not mean that the features of this invention are limited to this implementation. On the contrary, the purpose of introducing the invention in conjunction with the implementation is to cover other options or modifications that may be extended based on the implementation method of the present application. In order to provide an in-depth understanding of the present application, the above description contains many specific details, and the present application can also be implemented without using these details. In addition, in order to avoid confusion or blurring the focus of the present application, some specific details will be omitted in the description. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other in the absence of conflict.

[0197] It should be noted that in this specification, similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0198] Although the present application has been illustrated and described with reference to certain preferred implementations of the present application, those skilled in the art should understand that the above description is provided as a further detailed explanation of the present application in conjunction with specific implementations, and that the specific implementation of the present application should not be limited to these descriptions. Those skilled in the art may make various changes in form and detail, including simple deductions or substitutions, without departing from the spirit and scope of the present application.

Claims

1. A driving scene topology reasoning method based on a machine learning model, characterized in that: The machine learning model includes a plurality of loop modules arranged in a cascade manner, each of the loop modules including a traffic element detection module and a lane centerline detection module, the traffic element detection module including a traffic element decoding module and a traffic element and lane centerline topology reasoning module, the lane centerline detection module including a lane centerline decoding module and a lane centerline topology reasoning module, and the method including: The lane centerline decoding module at each level obtains updated lane centerline feature information and attention weight information based on the input lane centerline feature information and lane centerline topology reasoning information, and inputs the updated lane centerline feature information into the lane centerline topology reasoning module at the same level, the traffic element and lane centerline topology reasoning module at the same level, and the lane centerline decoding module at the next level, and inputs the attention weight information into the lane centerline topology reasoning module at the same level. The lane centerline feature information includes point-level lane centerline feature information and instance-level lane centerline feature information corresponding to the target driving scene image, and the updated lane centerline feature information includes updated point-level lane centerline feature information and updated instance-level lane centerline feature information. The lane centerline topology reasoning module at each level obtains updated lane centerline topology reasoning information based on the input updated lane centerline feature information and the attention weight information, and inputs the updated lane centerline topology reasoning information into the lane centerline decoding module at the next level until the target lane centerline topology relationship is obtained; The traffic element decoding modules at each level obtain updated traffic element feature information based on the input traffic element feature information, and input the updated traffic element feature information to the traffic element decoding module at the next level and the traffic element and lane centerline topology reasoning module at the same level; The traffic element and lane centerline topology reasoning modules at each level obtain updated traffic element and lane centerline topology reasoning information based on the input updated traffic element feature information and the updated lane centerline feature information until the target traffic element and lane centerline topology relationship is obtained.

2. The driving scene topology reasoning method based on a machine learning model according to claim 1, characterized in that: The attention weight information includes point-to-instance attention weight and instance-to-instance attention weight.

3. The driving scene topology reasoning method based on a machine learning model according to claim 2, characterized in that: The machine learning model further includes an initial processing module. The method further includes obtaining the point-level lane centerline feature information, the instance-level lane centerline feature information, and the lane centerline topology inference information input to the first-level lane centerline decoding module by: The initial processing module performs perspective feature conversion processing on the target driving scene image to obtain target perspective feature information, and performs query processing based on the target perspective feature information to obtain the point-level lane centerline feature information, the instance-level lane centerline feature information, and the lane centerline topology reasoning information that are input into the first-level lane centerline decoding module.

4. The driving scene topology reasoning method based on a machine learning model according to claim 3, characterized in that: The multiple loop modules include a first loop module and a second loop module, the first loop module is a previous module of the second loop module, the first loop module includes a first traffic element detection module and a first lane centerline detection module, the second loop module includes a second traffic element detection module and a second lane centerline detection module, the first traffic element detection module includes a first traffic element decoding module and a first traffic element and lane centerline topology reasoning module, the first lane centerline detection module includes a first lane centerline decoding module and a first lane centerline topology reasoning module, the second traffic element detection module includes a second traffic element decoding module and a second traffic element and lane centerline topology reasoning module, and the second lane centerline detection module includes a second lane centerline decoding module. The method includes: The first lane centerline decoding module obtains second point-level lane centerline feature information, second instance-level lane centerline feature information, first point-to-instance attention weight, and first instance-to-instance attention weight based on the input first point-level lane centerline feature information, first instance-level lane centerline feature information, and first lane centerline topology reasoning information, inputs the second point-level lane centerline feature information and the second instance-level lane centerline feature information into the first lane centerline topology reasoning module, the first traffic element and lane centerline topology reasoning module, and the second lane centerline decoding module, and inputs the first point-to-instance attention weight and the first instance-to-instance attention weight into the first lane centerline topology reasoning module; The first lane centerline topology inference module obtains second lane centerline topology inference information based on the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight, and inputs the second lane centerline topology inference information into the second lane centerline decoding module; The first traffic element decoding module obtains second traffic element feature information based on the input first traffic element feature information, and inputs the second traffic element feature information into the second traffic element decoding module and the first traffic element and lane centerline topology inference module; The first traffic element and lane centerline topology reasoning module obtains second traffic element and lane centerline topology reasoning information based on the input second traffic element feature information, the second point-level lane centerline feature information, and the second instance-level lane centerline feature information.

5. The driving scene topology reasoning method based on a machine learning model according to claim 4, characterized in that: Each lane centerline decoding module includes a point-level processing unit, an instance-level processing unit, a point-to-instance integration unit, and a relation decoder unit. The first lane centerline decoding module obtains second point-level lane centerline feature information, second instance-level lane centerline feature information, first point-to-instance attention weight, and first instance-to-instance attention weight based on the first point-level lane centerline feature information, first instance-level lane centerline feature information, and first lane centerline topology reasoning information input into the first lane centerline decoding module, including: The point-level processing unit obtains the second point-level lane centerline feature information based on the first point-level lane centerline feature information and a preset point-to-point relationship matrix; The relation decoder unit obtains point-to-instance relation information and instance-to-instance relation information based on the first lane centerline topology reasoning information; The point-to-instance integration unit obtains the first point-to-instance attention weight based on the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information, and determines a distance transformation centerline mask based on the second point-level lane centerline feature information; The instance-level processing unit obtains the second instance-level lane centerline feature information and the first instance-to-instance attention weight based on the first instance-level lane centerline feature information, the distance transformation centerline mask and the instance-to-instance relationship information.

6. The driving scene topology reasoning method based on a machine learning model according to claim 5, characterized in that: The distance transform centerline mask is obtained by the following formula: Where b is the pixel point corresponding to the target perspective feature information, C is the lane centerline point c corresponding to the second point-level lane centerline feature information i The set of components, D(b) is the Euclidean distance, L width is the width of the lane centerline, M(b) is the activation function, and RD(b) is the distance transform centerline mask.

7. The driving scene topology reasoning method based on a machine learning model according to claim 6, characterized in that: The point-level processing unit includes a line perception attention subunit and a cross attention subunit. The point-level processing unit obtains the second point-level lane centerline feature information based on the first point-level lane centerline feature information and a preset point-to-point relationship matrix, including: The line perception attention sub-unit performs corresponding processing on the first point-level lane centerline feature information and the preset point-to-point relationship matrix to obtain intermediate point-level lane centerline feature information; The cross-attention sub-unit performs corresponding processing on the intermediate point-level lane centerline feature information to obtain the second point-level lane centerline feature information.

8. The driving scene topology reasoning method based on a machine learning model according to claim 7, characterized in that: The point-to-instance integration unit includes an integrated attention subunit, and the point-to-instance integration unit obtains the first point-to-instance attention weight according to the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information, including: The integrated attention sub-unit performs corresponding processing on the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information to obtain the first point-to-instance attention weight.

9. The driving scene topology reasoning method based on a machine learning model according to claim 8, characterized in that: The instance-level processing unit includes a topology-aware attention subunit and a mask attention subunit. The instance-level processing unit obtains the second instance-level lane centerline feature information and the first instance-to-instance attention weight based on the first instance-level lane centerline feature information, the distance-transformed centerline mask, and the instance-to-instance relationship information, including: The topology-aware attention subunit performs corresponding processing on the first instance-level lane centerline feature information and the distance-transformed centerline mask to obtain intermediate instance-level lane centerline feature information; The masked attention sub-unit performs corresponding processing on the intermediate instance-level lane centerline feature information and the instance-to-instance relationship information to obtain the second instance-level lane centerline feature information and the first instance-to-instance attention weight.

10. The driving scenario topology reasoning method based on a machine learning model according to any one of claims 4 to 9, characterized in that: Each of the lane centerline topology reasoning modules includes a point-to-instance topology reasoning unit, an instance-to-instance topology reasoning unit, and a connection unit. The first lane centerline topology reasoning module obtains second lane centerline topology reasoning information based on the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight input into the first lane centerline topology reasoning module, including: The point-to-instance topology reasoning unit obtains point-to-instance topology reasoning information based on the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information; The instance-to-instance topology reasoning unit obtains instance-to-instance topology reasoning information according to the second instance-level lane centerline feature information and the first instance-to-instance attention weight; The connection unit obtains the second lane centerline topology reasoning information based on the point-to-instance topology reasoning information and the instance-to-instance topology reasoning information.

11. The driving scene topology reasoning method based on a machine learning model according to claim 10, characterized in that: The point-to-instance topology reasoning unit includes at least one first connection subunit and multiple first neural network subunits. The point-to-instance topology reasoning unit obtains point-to-instance topology reasoning information based on the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information, including: Each of the first neural network subunits processes the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information accordingly to obtain first intermediate feature information; The first connection subunit performs connection processing on the first intermediate feature information to obtain the point-to-instance topology reasoning information.

12. The driving scene topology reasoning method based on a machine learning model according to claim 11, characterized in that: The instance-to-instance topology reasoning unit includes at least one second connection subunit and a plurality of second neural network subunits. The instance-to-instance topology reasoning unit obtains instance-to-instance topology reasoning information based on the second instance-level lane centerline feature information and the first instance-to-instance attention weight, including: Each of the second neural network sub-units processes the second instance-level lane centerline feature information and the first instance-to-instance attention weight accordingly to obtain second intermediate feature information; The second connection subunit performs connection processing on the second intermediate feature information to obtain the instance-to-instance topology reasoning information.

13. The driving scene topology reasoning method based on a machine learning model according to claim 3, characterized in that: The initial processing module includes a shared backbone network unit and a view conversion unit. The initial processing module performs view feature conversion processing on the target driving scene image to obtain target view feature information, including: The shared backbone network unit performs feature extraction processing on the target driving scene image to obtain third intermediate feature information; The view conversion unit performs feature conversion processing on the third intermediate feature information to obtain the target viewing angle feature information.

14. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores a computer program; The processor executes the computer program stored in the memory so that the electronic device implements the driving scene topology reasoning method based on the machine learning model as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Topological reasoning method, system and equipment for driving scene and storage medium

    CN116386009A