Driving scene topology reasoning method based on machine learning model and electronic equipment

Through the cascading loop module based on machine learning model, the accuracy of topological inference in driving scenarios is improved, the problem of inaccurate topological inference results in the existing technology is solved, and the accuracy and safety of navigation and decision-making of autonomous driving is improved.

CN120258153AActive Publication Date: 2025-07-04NULLMAX INC

Patent Information

Application Number
CN202510732754.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In the prior art, the accuracy of the topological inference results of driving scenarios is insufficient, which affects the accuracy and safety of autonomous driving navigation and decision-making.

Method used

The topological inference method of driving scenarios based on machine learning models is adopted, and multiple cyclic modules are set up in a cascade manner, including traffic element detection module and lane centerline detection module. The cyclic enhancement inference of the lane centerline decoding module and topological inference module are used to capture fine-grained point-to-instance relationship and global topological connection, and improve the accuracy of topological inference results.

Benefits of technology

It improves the accuracy of the topological relationship between the lane center line and between traffic elements and the lane center line during vehicle driving, and improves the accuracy and safety of autonomous driving navigation and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258153A_ABST
    Figure CN120258153A_ABST
Patent Text Reader

Abstract

The invention discloses a driving scene topology reasoning method based on a machine learning model and electronic equipment, the model comprises a plurality of circulation modules arranged in a cascade mode, and each circulation module comprises a traffic element detection module and a lane center line detection module. The lane center line detection module comprises a lane center line decoding module and a lane center line topology reasoning module. In the method, each level of lane center line decoding module obtains updated lane center line feature information and attention weight information which are input to the next level of lane center line decoding module and the same level of lane center line topology reasoning module based on input lane center line feature information and lane center line topology reasoning information; therefore, circular interaction is realized until the topological relation corresponding to the target driving scene is obtained. The accuracy or precision of driving scene topological relation reasoning is improved, so that the accuracy of automatic driving control is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of autonomous driving, and particularly to a driving scenario topology inference method and an electronic device based on a machine learning model. Background Art

[0002] Driving scenario topology inference refers to a technology in an autonomous driving system that infers topology inference results such as the topology structure of a driving scenario by analyzing the spatial relationships and logical structures of elements such as lanes, intersections, vehicles, and pedestrians in the driving scenario. This kind of inference can help the autonomous driving system better understand road layouts, driving paths, and other information related to autonomous driving, thereby improving the accuracy of autonomous driving controls such as navigation and decision-making. Therefore, the accuracy or precision of the driving scenario topology inference results directly or indirectly affects the accuracy of autonomous driving controls such as navigation and decision-making, and is crucial for the accuracy and safety of autonomous driving.

[0003] How to improve the accuracy or precision of the driving scenario topology inference results, thereby improving the accuracy of autonomous driving controls such as autonomous driving navigation and decision-making, to improve the accuracy and safety of autonomous driving, is a direction being explored in the current field. Summary of the Invention

[0004] The present application provides a driving scenario topology inference method based on a machine learning model, which is used to solve the problem in the prior art of how to improve the accuracy or precision of the driving scenario topology inference results, thereby improving the accuracy of autonomous driving controls such as autonomous driving navigation and decision-making, to improve the accuracy and safety of autonomous driving.

[0005] To solve the above technical problems, in a first aspect, an implementation manner of the present application discloses a driving scenario topology inference method based on a machine learning model. The machine learning model includes a plurality of recurrent modules arranged in a cascaded manner. Each recurrent module includes a traffic element detection module and a lane centerline detection module. The traffic element detection module includes a traffic element decoding module and a traffic element and lane centerline topology inference module. The lane centerline detection module includes a lane centerline decoding module and a lane centerline topology inference module. The method includes: each level of lane centerline decoding module obtains updated lane centerline feature information and attention weight information based on the input lane centerline feature information and lane centerline topology inference information, inputs the updated lane centerline feature information to the lane centerline topology inference module at the same level, the traffic element and lane centerline topology inference module at the same level, and the lane centerline decoding module at the next level, and inputs the attention weight information to the lane centerline topology inference module at the same level; each level of lane centerline topology inference module obtains updated lane centerline topology inference information based on the input updated lane centerline feature information and attention weight information, and inputs the updated lane centerline topology inference information to the lane centerline decoding module at the next level until the target lane centerline topological relationship is obtained; each level of traffic element decoding module obtains updated traffic element feature information based on the input traffic element feature information, and inputs the updated traffic element feature information to the traffic element decoding module at the next level and the traffic element and lane centerline topology inference module at the same level; each level of traffic element and lane centerline topology inference module obtains updated traffic element and lane centerline topology inference information based on the input updated traffic element feature information and updated lane centerline feature information until the target traffic element and lane centerline topological relationship is obtained.

[0006] Among them, the target lane centerline topological relationship is the topological relationship between the lane centerlines corresponding to the target driving scenario, and the target traffic element and lane centerline topological relationship is the topological relationship between the lane centerlines and traffic elements corresponding to the target driving scenario.

[0007] With the above technical solution, a plurality of loop modules are provided in a cascaded manner, and a traffic element detection module including a traffic element decoding module and a traffic element and lane centerline topology inference module is provided in the loop module, and a lane centerline detection module including a lane centerline decoding module and a lane centerline topology inference module is provided. Thus, according to the lane centerline feature information and attention weight information in the driving scene, as well as the traffic element feature information and traffic element and lane centerline topology inference information, fine-grained point-to-instance relationships and global topological connections are captured. Moreover, through the cyclic enhanced inference of the output information of the traffic element decoding module and the traffic element and lane centerline topology inference module, and the lane centerline decoding module and the lane centerline topology inference module, the accuracy or precision of the topological inference results of the driving scene, such as the topological relationship between lane centerlines and the topological relationship between traffic elements and lane centerlines obtained during vehicle driving, is effectively improved, thereby improving the accuracy of autonomous driving navigation, decision-making, etc. in autonomous driving control, and further improving the accuracy and safety of autonomous driving.

[0008] According to another specific implementation manner of the present application, the lane centerline feature information includes point-level lane centerline feature information and instance-level lane centerline feature information corresponding to the target driving scene image, and the updated lane centerline feature information includes updated point-level lane centerline feature information and updated instance-level lane centerline feature information.

[0009] According to another specific implementation manner of the present application, the attention weight information includes point-to-instance attention weight and instance-to-instance attention weight.

[0010] That is, according to another specific implementation manner of the present application, the method includes: each level of lane centerline decoding module obtains updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weights, and instance-to-instance attention weights based on the point-level lane centerline feature information, instance-level lane centerline feature information, and lane centerline topology inference information corresponding to the target driving scene image in the input lane centerline decoding module, inputs the updated point-level lane centerline feature information and updated instance-level lane centerline feature information to the next-level lane centerline decoding module, inputs the updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weights, and instance-to-instance attention weights to the same-level lane centerline topology inference module, and inputs the updated point-level lane centerline feature information and updated instance-level lane centerline feature information to the same-level traffic element and lane centerline topology inference module; each level of lane centerline topology inference module obtains updated lane centerline topology inference information based on the updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weights, and instance-to-instance attention weights in the input lane centerline topology inference module, and inputs the updated lane centerline topology inference information to the next-level lane centerline decoding module until the lane centerline detection result information corresponding to the target driving scene and the topological relationship between lane centerlines are obtained; each level of traffic element decoding module obtains updated traffic element feature information based on the traffic element feature information corresponding to the target driving scene image in the input traffic element decoding module, inputs the updated traffic element feature information to the next-level traffic element decoding module, and inputs the updated traffic element feature information to the same-level traffic element and lane centerline topology inference module; each level of traffic element and lane centerline topology inference module obtains updated traffic element and lane centerline topology inference information based on the updated traffic element feature information, updated point-level lane centerline feature information, and updated instance-level lane centerline feature information in the input traffic element and lane centerline topology inference module until the topological relationship between the lane centerline and traffic elements corresponding to the target driving scene is obtained.

[0011] With the above technical solution, a plurality of loop modules are provided in a cascaded manner, and a traffic element detection module including a traffic element decoding module and a traffic element and lane centerline topology inference module is provided in the loop module, and a lane centerline detection module including a lane centerline decoding module and a lane centerline topology inference module is provided. Among them, each level (i.e., each layer) of the lane centerline decoding module processes the corresponding point-level lane centerline feature information, instance-level lane centerline feature information, and topology inference information corresponding to the target driving scene image input thereto to obtain updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weights, and instance-to-instance attention weights, and then inputs the updated point-level lane centerline feature information and updated instance-level lane centerline feature information to the next-level lane centerline decoding module for corresponding processing, inputs the updated point-level lane centerline feature information and updated instance-level lane centerline feature information, as well as the point-to-instance attention weights and instance-to-instance attention weights to the same-level lane centerline topology inference module for corresponding processing, and inputs the updated point-level lane centerline feature information and updated instance-level lane centerline feature information to the same-level traffic element and lane centerline topology inference module. Each level (i.e., each layer) of the lane centerline topology inference module processes the corresponding updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weights, and instance-to-instance attention weights input thereto to obtain updated lane centerline topology inference information, and inputs the updated lane centerline topology inference information to the next-level lane centerline decoding module for corresponding processing until the topological relationship between the lane centerlines corresponding to the target driving scene is obtained. And each level of the traffic element decoding module obtains updated traffic element feature information based on the traffic element feature information corresponding to the target driving scene image input to the traffic element decoding module, inputs the updated traffic element feature information to the next-level traffic element decoding module, and inputs the updated traffic element feature information to the same-level traffic element and lane centerline topology inference module; each level of the traffic element and lane centerline topology inference module obtains updated traffic element and lane centerline topology inference information based on the input updated traffic element feature information, as well as the updated point-level lane centerline feature information and updated instance-level lane centerline feature information, until the topological relationship between the traffic elements and lane centerlines corresponding to the target driving scene is obtained.

[0012] Thus, it is possible to capture fine-grained point-to-instance relationships and global topological connections based on the point-level lane centerline feature information and instance-level lane centerline feature information corresponding to the target driving scene image in the driving scene, as well as the point-to-instance attention weights and instance-to-instance attention weights, and the traffic element feature information and the topological inference information between traffic elements and lane centerlines. Moreover, through the cyclic enhanced inference of the output information of the traffic element decoding module, the traffic element and lane centerline topological inference module, the lane centerline decoding module, and the lane centerline topological inference module, the accuracy or precision of the topological inference results of the driving scene, such as the topological relationship between lane centerlines obtained during vehicle driving and the topological relationship between traffic elements and lane centerlines, is effectively improved, thereby improving the accuracy of autonomous driving navigation, decision-making, and other aspects of autonomous driving control, and further enhancing the accuracy and safety of autonomous driving.

[0013] According to another specific implementation manner of the present application, the machine learning model further includes an initial processing module, and the method further includes obtaining the point-level lane centerline feature information, instance-level lane centerline feature information, and lane centerline topological inference information input to the first-level lane centerline decoding module in the following manner: The initial processing module performs perspective feature conversion processing on the target driving scene image to obtain target perspective feature information, and based on the target perspective feature information, performs query processing to obtain the point-level lane centerline feature information, instance-level lane centerline feature information, and lane centerline topological inference information input to the first-level lane centerline decoding module.

[0014] According to another specific implementation manner of the present application, the target perspective feature information is bird's-eye view perspective feature information.

[0015] Adopting the above technical solution, through the initial processing module performing image perspective feature conversion processing on the target driving scene image to obtain target perspective feature information (such as converting to bird's-eye view perspective feature information), and performing query processing based on the target perspective feature information to obtain the point-level lane centerline feature information, instance-level lane centerline feature information, and topological inference information input to the first-level lane centerline decoding module, it can more accurately reflect the local and global information of the target driving scene image, effectively increasing the accuracy of the traffic element detection result information, lane centerline detection result information, topological relationship between lane centerlines, and topological relationship between lane centerlines and traffic elements obtained during vehicle driving.

[0016] According to another specific implementation manner of the present application, the method further includes the initial processing module performing second perspective feature conversion processing on the target driving scene image to obtain second perspective feature information, and performing traffic element query processing on the second perspective feature information to obtain the traffic element feature information input to the first-level traffic element decoding module.

[0017] Among them, the second perspective feature information may be, for example, perspective view (PV) feature information.

[0018] According to another specific implementation manner of the present application, the method further includes performing feature conversion processing on the perspective view feature information to obtain bird's-eye view perspective feature information.

[0019] According to another specific implementation manner of the present application, the multiple loop modules include a first loop module and a second loop module. The first loop module is the previous-level module of the second loop module. The first loop module includes a first traffic element detection module and a first lane centerline detection module. The second loop module includes a second traffic element detection module and a second lane centerline detection module. The first traffic element detection module includes a first traffic element decoding module and a first traffic element and lane centerline topology inference module. The first lane centerline detection module includes a first lane centerline decoding module and a first lane centerline topology inference module. The second traffic element detection module includes a second traffic element decoding module and a second traffic element and lane centerline topology inference module. The second lane centerline detection module includes a second lane centerline decoding module. The method includes: The first lane centerline decoding module, based on the first point-level lane centerline feature information, the first instance-level lane centerline feature information, and the first lane centerline topology inference information input into the first lane centerline decoding module, obtains second point-level lane centerline feature information, second instance-level lane centerline feature information, a first point-to-instance attention weight, and a first instance-to-instance attention weight, inputs the second point-level lane centerline feature information and the second instance-level lane centerline feature information into the first lane centerline topology inference module, the first traffic element and lane centerline topology inference module, and the second lane centerline decoding module, and inputs the first point-to-instance attention weight and the first instance-to-instance attention weight into the first lane centerline topology inference module; The first lane centerline topology inference module, based on the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight input into the first lane centerline topology inference module, obtains second lane centerline topology inference information, and inputs the second lane centerline topology inference information into the second centerline decoding module; The first traffic element decoding module, based on the first traffic element feature information input into the first traffic element decoding module, obtains second traffic element feature information, inputs the second traffic element feature information into the second traffic element decoding module, and the first traffic element and lane centerline topology inference module; The first traffic element and lane centerline topology inference module, based on the second traffic element feature information input into the first traffic element and lane centerline topology inference module, and the second point-level lane centerline feature information and the second instance-level lane centerline feature information, obtains second traffic element and lane centerline topology inference information.

[0020] Wherein, the first traffic element feature information is also the first traffic element feature information corresponding to the target driving scenario.

[0021] With the above technical solution, through fine-grained point-level interaction and high-level instance-level relationships, an accurate and semantically rich understanding of the topological structure is achieved. The cyclic interaction mechanism promotes iterative refinement, where the attention weights from the lane centerline decoding module (which can also be referred to as the detector) are fed into the lane centerline topology inference module, and the topological inference information (which can also be referred to as the instance topological relationship) from the lane centerline topology inference module is fed back into the lane centerline decoding module, thus creating a complementary cycle to improve performance. Through the cyclic enhancement of inference by the output information of the lane centerline decoding module and the lane centerline topology inference module, the accuracy of the topological relationships between the lane centerlines obtained during vehicle driving and the topological relationships between traffic elements and lane centerlines is effectively increased.

[0022] According to another specific implementation manner of the present application, each lane centerline decoding module includes a point-level processing unit, an instance-level processing unit, a point-to-instance integration unit, and a relationship decoder unit. Then, based on the first point-level lane centerline feature information, the first instance-level lane centerline feature information, and the first lane centerline topology inference information input into the first lane centerline decoding module, the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight are obtained, including: the point-level processing unit obtains the second point-level lane centerline feature information according to the first point-level lane centerline feature information and a preset point-to-point relationship matrix; the relationship decoder unit obtains the point-to-instance relationship information and the instance-to-instance relationship information according to the first lane centerline topology inference information; the point-to-instance integration unit obtains the first point-to-instance attention weight according to the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information, and determines the distance-transformed centerline mask according to the second point-level lane centerline feature information; the instance-level processing unit obtains the second instance-level lane centerline feature information and the first instance-to-instance attention weight according to the first instance-level lane centerline feature information, the distance-transformed centerline mask, and the instance-to-instance relationship information.

[0023] With the above technical solution, a point-level processing unit, an instance-level processing unit, a point-to-instance integration unit, and a relationship decoder unit are set in the centerline decoding module. The point-level processing unit obtains second point-level lane centerline feature information according to the first point-level lane centerline feature information and a preset point-to-point relationship matrix. The relationship decoder unit obtains point-to-instance relationship information and instance-to-instance relationship information according to the first lane centerline topological inference information. The point-to-instance integration unit obtains a first point-to-instance attention weight according to the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information, and determines a distance transform centerline mask according to the second point-level lane centerline feature information. The instance-level processing unit obtains second instance-level lane centerline feature information and a first instance-to-instance attention weight according to the first instance-level lane centerline feature information, the distance transform centerline mask, and the instance-to-instance relationship information. It effectively increases the accuracy of the corresponding processing based on the first point-level lane centerline feature information, the first instance-level lane centerline feature information, and the first topological inference information input into the first lane centerline decoding module, to obtain the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight. Furthermore, it increases the accuracy of the traffic element detection result information, the lane centerline detection result information, the topological relationship between lane centerlines, and the topological relationship between lane centerlines and traffic elements obtained during the vehicle driving process, improving driving safety.

[0024] According to another specific implementation manner of the present application, the distance transform centerline mask can be obtained through the following formula:

[0025]

[0026] where b is the pixel point corresponding to the target perspective feature information, C is the set composed of the lane centerline points c corresponding to the second point-level lane centerline feature information, D(b) is the Euclidean distance, L i is the width of the lane centerline, M(b) is the activation function, and RD(b) is the distance transform centerline mask. width

[0027] With the above technical solution, distance transform is beneficial for extracting position information, and rasterized distance transform can reduce the learning difficulty of the model and accelerate convergence. Therefore, using the distance transform centerline mask (which can also be called the rasterized distance transform mask) can better extract the position information of the image (such as the position information of the centerline), and can further increase the accuracy of the obtained second instance-level lane centerline feature information and the first instance-to-instance attention weight.

[0028] According to another specific implementation manner of the present application, the point-level processing unit includes a line-aware attention sub-unit and a cross-attention sub-unit. The point-level processing unit obtains second point-level lane centerline feature information according to the first point-level lane centerline feature information and a preset point-to-point relationship matrix, including: the line-aware attention sub-unit performs corresponding processing on the first point-level lane centerline feature information and the preset point-to-point relationship matrix to obtain intermediate point-level lane centerline feature information; the cross-attention sub-unit performs corresponding processing on the intermediate point-level lane centerline feature information to obtain second point-level lane centerline feature information.

[0029] By adopting the above technical solution, by setting the line-aware attention sub-unit, it is possible to effectively promote the intra-instance interaction between point-level queries and increase the accuracy of the obtained second point-level feature information.

[0030] According to another specific implementation manner of the present application, the point-to-instance integration unit includes an integration attention sub-unit. The point-to-instance integration unit obtains a first point-to-instance attention weight according to the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information, including: the integration attention sub-unit performs corresponding processing on the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information to obtain a first point-to-instance attention weight.

[0031] By adopting the above technical solution, the integration attention sub-unit enables each point-level query to selectively focus on relevant instance-level features, thereby enriching its representation with local and global contexts. By setting the integration attention sub-unit, it is possible to further enhance instance-aware feature learning and reasoning, ensuring that the point-level queries are enriched with both local position details and global semantic contexts. The accuracy of the obtained first point-to-instance attention weight is effectively increased.

[0032] According to another specific implementation manner of the present application, the instance-level processing unit includes a topology-aware attention sub-unit and a mask attention sub-unit. The instance-level processing unit obtains second instance-level lane centerline feature information and a first instance-to-instance attention weight according to the first instance-level lane centerline feature information, the distance-transformed centerline mask, and the instance-to-instance relationship information, including: the topology-aware attention sub-unit performs corresponding processing on the first instance-level lane centerline feature information and the distance-transformed centerline mask to obtain intermediate instance-level lane centerline feature information; the mask attention sub-unit performs corresponding processing on the intermediate instance-level lane centerline feature information and the instance-to-instance relationship information to obtain second instance-level lane centerline feature information and a first instance-to-instance attention weight.

[0033] Adopting the above technical solution effectively improves the accuracy of the obtained second instance-level lane centerline feature information and the first instance-to-instance attention weight.

[0034] According to another specific implementation manner of the present application, each lane centerline topology reasoning module includes a point-to-instance topology reasoning unit, an instance-to-instance topology reasoning unit, and a connection unit. The first lane centerline topology reasoning module performs corresponding processing based on the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight input into the first lane centerline topology reasoning module to obtain the second lane centerline topology reasoning information, including: the point-to-instance topology reasoning unit obtains the point-to-instance topology reasoning information according to the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information; the instance-to-instance topology reasoning unit obtains the instance-to-instance topology reasoning information according to the second instance-level lane centerline feature information and the first instance-to-instance attention weight; the connection unit obtains the second lane centerline topology reasoning information according to the point-to-instance topology reasoning information and the instance-to-instance topology reasoning information.

[0035] Adopting the above technical solution, by setting a point-to-instance topology reasoning unit, an instance-to-instance topology reasoning unit (which can also be called a dual prediction branch), and a connection unit in the lane centerline topology reasoning module, the point-to-instance topology reasoning unit combines explicit feature correlation and potential dependence for robust topology reasoning, and the instance-to-instance topology reasoning unit uses instance-level queries and their self-attention weights for similarity calculation, effectively integrating instance-level features and their hidden dependencies, realizing comprehensive topology reasoning, and improving the accuracy of the obtained second lane centerline topology reasoning information.

[0036] According to another specific implementation manner of the present application, the point-to-instance topology reasoning unit includes at least one first connection sub-unit and multiple first neural network sub-units. The point-to-instance topology reasoning unit obtains the point-to-instance topology reasoning information according to the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information, including: each first neural network sub-unit performs corresponding processing on the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information to obtain the first intermediate feature information; the first connection sub-unit performs connection processing on the first intermediate feature information to obtain the point-to-instance topology reasoning information.

[0037] Adopting the above technical solution, by extracting the association relationship based on attention by each first neural network sub-unit, the accuracy of the obtained point-to-instance topology reasoning information is improved.

[0038] According to another specific implementation manner of the present application, the instance-to-instance topology inference unit includes at least one second connection sub-unit and multiple second neural network sub-units. The instance-to-instance topology inference unit obtains instance-to-instance topology inference information based on the second instance-level lane centerline feature information and the first instance-to-instance attention weight, including: the second neural network sub-unit performs corresponding processing on the second instance-level lane centerline feature information and the first instance-to-instance attention weight to obtain second intermediate feature information; the second connection unit performs connection processing on the second intermediate feature information to obtain instance-to-instance topology inference information.

[0039] With the above technical solution, by extracting the association relationship through each second neural network sub-unit based on attention, the accuracy of the obtained instance-to-instance topology inference information is increased.

[0040] According to another specific implementation manner of the present application, the initial processing module includes a shared backbone network unit and a view conversion unit. The initial processing module performs image perspective feature conversion processing on the target driving scene image to obtain target perspective feature information, including: the shared backbone network unit performs feature extraction processing on the target driving scene image to obtain third intermediate feature information; the view conversion unit performs feature conversion processing on the third intermediate feature information to obtain target perspective feature information.

[0041] With the above technical solution, through the shared backbone network unit and the view conversion unit, the features of the target driving scene image can be accurately extracted to obtain more accurate target perspective feature information.

[0042] Second aspect, the implementation of the present application also discloses a machine learning model, which includes a plurality of recurrent modules arranged in a cascaded manner. Each recurrent module includes a traffic element detection module and a lane centerline detection module. The traffic element detection module includes a traffic element decoding module and a traffic element and lane centerline topology inference module. The lane centerline detection module includes a lane centerline decoding module and a lane centerline topology inference module. Among them, each level of lane centerline decoding module, based on the input lane centerline feature information and lane centerline topology inference information, obtains updated lane centerline feature information and attention weight information, inputs the updated lane centerline feature information to the same-level lane centerline topology inference module, the same-level traffic element and lane centerline topology inference module, and the next-level lane centerline decoding module, and inputs the attention weight information to the same-level lane centerline topology inference module; each level of lane centerline topology inference module, based on the input updated lane centerline feature information and attention weight information, obtains updated lane centerline topology inference information, and inputs the updated lane centerline topology inference information to the next-level lane centerline decoding module until the target lane centerline topological relationship is obtained; each level of traffic element decoding module, based on the input traffic element feature information, obtains updated traffic element feature information, and inputs the updated traffic element feature information to the next-level traffic element decoding module and the same-level traffic element and lane centerline topology inference module; each level of traffic element and lane centerline topology inference module, based on the input updated traffic element feature information and updated lane centerline feature information, obtains updated traffic element and lane centerline topology inference information until the target traffic element and lane centerline topological relationship is obtained.

[0043] Third aspect, the implementation of the present application also discloses a vehicle control method, which includes: obtaining target driving scenario topology inference result information, such as the topological relationship between lane centerlines, and the topological relationship between lane centerlines and traffic elements, through the above-mentioned driving scenario topology inference method based on a machine learning model; controlling the vehicle to drive according to the target driving scenario topology inference result information.

[0044] Fourth aspect, the implementation of the present application also discloses an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores a computer program; the processor executes the computer program stored in the memory so that the electronic device implements the driving scenario topology inference method based on a machine learning model provided by any one of the implementations in the first aspect above.

[0045] Fifth aspect, an implementation manner of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the driving scenario topology inference method based on a machine learning model provided by any one of the implementation manners in the first aspect above.

[0046] Sixth aspect, an implementation manner of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the driving scenario topology inference method based on a machine learning model provided by any one of the implementation manners in the first aspect above.

[0047] It can be understood that the beneficial effects of the second aspect to the sixth aspect above can also refer to the relevant descriptions in the first aspect above, and will not be elaborated here. Description of the Drawings

[0048] Figure 1 is a schematic diagram of the principle of a centerline detection and topology inference method in the prior art;

[0049] Figure 2 is a schematic diagram of the principle of another centerline detection and topology inference method in the prior art;

[0050] Figure 3 is a schematic diagram of the structure of a machine learning model provided by an embodiment of the present application;

[0051] Figure 4 is a schematic diagram of the structure of a lane centerline decoding module provided by an embodiment of the present application;

[0052] Figure 5 is a schematic diagram of the structure of a point-level processing unit provided by an embodiment of the present application;

[0053] Figure 6 is a schematic diagram of the structure of a point-to-instance integration unit provided by an embodiment of the present application;

[0054] Figure 7 is a schematic diagram of the structure of an instance-level processing unit provided by an embodiment of the present application;

[0055] Figure 8 is a schematic diagram of the structure of a lane centerline topology inference module provided by an embodiment of the present application;

[0056] Figure 9 is a schematic diagram of the structure of a point-to-instance topology inference unit provided by an embodiment of the present application;

[0057] Figure 10 is a schematic diagram of the structure of an instance-to-instance topology inference unit provided by an embodiment of the present application;

[0058] Figure 11 It is a schematic diagram of the principle of a cyclic inference framework corresponding to a machine learning model provided by an embodiment of the present application;

[0059] Figure 12 It is a schematic diagram of the principle of a centerline detection and topology inference method provided by an embodiment of the present application;

[0060] Figure 13 It is a schematic diagram of the structure of an inference framework corresponding to a machine learning model provided by an embodiment of the present application;

[0061] Figure 14 It is a schematic diagram of the structure of another inference framework corresponding to a machine learning model provided by an embodiment of the present application;

[0062] Figure 15 It is a schematic diagram of the structure of the centerline decoder of the l-th layer in an inference framework corresponding to a machine learning model provided by an embodiment of the present application;

[0063] Figure 16 It is a schematic diagram of the structure of the topology inference module of the l-th layer in a cyclic inference framework corresponding to a machine learning model provided by an embodiment of the present application;

[0064] Figure 17 It is a schematic diagram of the performance comparison between TopoHR and other state-of-the-art methods in the OpenLane-V2 subset A benchmark test provided by an embodiment of the present application;

[0065] Figure 18 It is a schematic diagram of the performance comparison between TopoHR and other state-of-the-art methods in the OpenLane-V2 subset B benchmark test provided by an embodiment of the present application;

[0066] Figure 19 It is a schematic diagram of the ablation study results of the centerline mask (Mask GT) representation provided by an embodiment of the present application;

[0067] Figure 20 It is a schematic diagram of the ablation study results of the influence of hierarchical attention provided by an embodiment of the present application;

[0068] Figure 21 It is a schematic diagram of the ablation study results of the effectiveness of hierarchical topology inference provided by an embodiment of the present application;

[0069] Figure 22 It is a schematic diagram of the ablation study results of the adaptive topology loss in TopoHR provided by an embodiment of the present application;

[0070] Figure 23 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0071] As described above, the accuracy of the driving scenario topology inference result directly or indirectly affects the accuracy of autonomous driving control aspects such as navigation and decision-making, which is crucial for the accuracy and safety of autonomous driving.

[0072] The driving scenario topology inference result may include the topological relationship between the driving scenario traffic elements and the lane centerline, as well as the topological relationship between the lane centerlines, etc. In the prior art, the determination method for the topological relationship between the driving scenario traffic elements and the lane centerline and the topological relationship between the lane centerlines is usually to first obtain the driving scenario image of the vehicle's travel, then perform instance-level learning of centerline detection on the image to obtain the instance-level information corresponding to the centerline of the driving scenario image, and then rely on an ordered module composed of a simplified multi-layer perceptron (MLP) layer to perform topology inference to obtain the topological relationship between the driving scenario traffic elements and the lane centerline and the topological relationship between the lane centerlines during the vehicle driving process. This method only determines the topological relationship between the driving scenario traffic elements and the lane centerline and the topological relationship between the lane centerlines from the instance-level information, and queries and processes the information of the driving scenario image in a pipeline manner to obtain the topological relationship between the driving scenario traffic elements and the lane centerline and the topological relationship between the lane centerlines, resulting in inaccurate topological relationships between the driving scenario traffic elements and the lane centerline and the topological relationship between the lane centerlines during the vehicle driving process, which affects the accuracy and safety of autonomous driving.

[0073] Specifically, traditional lane detection methods and online mapping techniques focus on geometric accuracy but fail to capture the scene topology relationship. Although a high definition map (HD Map) can provide topological information of the scene topology relationship, there are freshness and scalability problems.

[0074] As Figure 1 shown, a sequential pipeline approach is provided. In this pipeline, the centerline decoder (i.e., the lane centerline detection module, which can also be called the Centerline Decoder) first extracts the instance-level centerline representation, and the topology inference module then analyzes these extracted centerline instances.

[0075] Furthermore, as Figure 2As shown, most centerline detection methods usually use instance-level query representations in the transformer decoder, defining the task as point set (Point Set) prediction, curve parameter (Curve Param) estimation, or binary segmentation (such as 0 / 1 segmentation, 0 / 1 Segment). For example, TopoNet represents the centerline as a point set and refines it using a scene graph neural network (SGNN). In contrast, TopoBDA models centerline detection as a Bezier curve parameter estimation task with a deformable attention mechanism. Although vectorization-based methods adopt point-level attention in local geometry modeling, they tend to ignore global context patterns. Alternatively, segmentation-based methods, such as TopoMask, combine orientation prediction to derive vectorized results from the segmentation output.

[0076] However, the centerline is inherently invisible, making it challenging to accurately extract its features through direct segmentation modeling. Current implementations either use segmentation as auxiliary supervision or rely on post-processing to generate vectorized point sets, which limits the full utilization of the rich information contained in the segmentation results. Based on the instance-level queries of the centerline detector, the topology inference module usually applies an MLP function to infer topological relationships.

[0077] For example, TopoNet uses three MLP layers to infer topological relationships through precise centerline queries. TopoMLP emphasizes the importance of detector performance in the cascade structure and enhances topological inference by integrating position embeddings into the MLP layer. Similarly, TopoLogic improves the inference performance by combining centerline geometric priors and connecting two centerlines through start and end points.

[0078] Although these methods have achieved promising results in topological inference for driving scenarios, there are still significant gaps in the accuracy of topological relationship prediction. These methods usually rely on cascade architectures where the centerline detection and relationship inference modules are optimized separately, resulting in inconsistent feature representations. In addition, using a simplified prediction head (usually consisting of only 3 layers of MLP) cannot fully capture the complex spatial dependencies inherent in urban road networks.

[0079] Based on this, an implementation manner of the present application provides a driving scenario topology inference method based on a machine learning model, which can determine target driving scenario topology inference result information including, for example, the topological relationship between lane centerlines and the topological relationship between lane centerlines and traffic elements according to the point-level lane centerline feature information, instance-level lane centerline feature information, point-to-instance attention weights, and instance-to-instance attention weights corresponding to the driving scenario images in the image scenario, capturing fine-grained point-to-instance relationships and global topological connections. Moreover, through the mutual cyclic enhancement inference of the information output by the lane centerline decoding module and the lane centerline topology inference module, the accuracy or precision of the driving scenario topology inference result is effectively improved, thereby improving the accuracy of autonomous driving navigation, decision-making, and other aspects of autonomous driving control, and further improving the accuracy and safety of autonomous driving.

[0080] Next, with reference to the accompanying drawings, the steps and advantages of the driving scenario topology inference method based on a machine learning model provided by the present application will be described in detail.

[0081] In an implementation manner of the present application, as Figure 3 shown, the machine learning model includes a plurality of recurrent modules, such as recurrent module 1, recurrent module 2... recurrent module n, etc., arranged in a cascaded manner (for example, set to n levels). Among them, each recurrent module includes a traffic element detection module and a lane centerline detection module. The traffic element detection module includes a traffic element decoding module and a traffic element and lane centerline topology inference module, and the lane centerline detection module includes a lane centerline decoding module and a lane centerline topology inference module.

[0082] Therefore, the driving scenario topology inference method based on the machine learning model includes: each level of lane centerline decoding module performs corresponding processing based on the point-level lane centerline feature information, instance-level lane centerline feature information, and lane centerline topology inference information corresponding to the target driving scenario image in the input lane centerline decoding module, obtaining updated point-level lane centerline feature information and updated instance-level lane centerline feature information, as well as point-to-instance attention weights and instance-to-instance attention weights, inputting the updated point-level lane centerline feature information and updated instance-level lane centerline feature information to the next-level lane centerline decoding module for corresponding processing, inputting the updated point-level lane centerline feature information and updated instance-level lane centerline feature information, as well as point-to-instance attention weights and instance-to-instance attention weights, to the same-level lane centerline topology inference module for corresponding processing, and inputting the updated point-level lane centerline feature information and updated instance-level lane centerline feature information to the same-level traffic element and lane centerline topology inference module.

[0083] The topological inference module for lane centerlines at all levels processes the updated point-level lane centerline feature information and updated instance-level lane centerline feature information in the input lane centerline topological inference module, as well as the point-to-instance attention weights and instance-to-instance attention weights, to obtain updated lane centerline topological inference information. The updated lane centerline topological inference information is input to the next-level lane centerline decoding module for corresponding processing until the topological relationship between the lane centerlines corresponding to the target driving scenario is obtained, for example.

[0084] The traffic element decoding module at all levels obtains updated traffic element feature information based on the traffic element feature information corresponding to the target driving scenario image input to the traffic element decoding module, inputs the updated traffic element feature information to the next-level traffic element decoding module, and inputs the updated traffic element feature information to the traffic element and lane centerline topological inference module at the same level.

[0085] The traffic element and lane centerline topological inference module at all levels obtains updated traffic element and lane centerline topological inference information based on the updated traffic element feature information, as well as the updated point-level lane centerline feature information and updated instance-level lane centerline feature information in the input traffic element and lane centerline topological inference module, until the topological relationship between the traffic elements and lane centerlines corresponding to the target driving scenario is obtained.

[0086] In other embodiments of the present application, the output of the lane centerline decoding module can also be understood as the lane centerline detection result information corresponding to the target driving scenario. That is, through the cyclic processing of the traffic element detection module and the lane centerline detection module arranged in a cascaded manner in the present application, the target driving scenario topological inference result information can be obtained. The target driving scenario topological inference result information can include, for example, lane centerline detection result information, the topological relationship between lane centerlines, traffic element detection result information, and the topological relationship between lane centerlines and traffic elements.

[0087] The point-level lane centerline feature information and instance-level lane centerline feature information can be the point-level and instance-level feature information of important elements in the driving scenario, such as the information of the lane centerline during vehicle driving, or the information of other important elements.

[0088] That is, in one implementation manner of the present application, the lane centerline decoding module at all levels can also obtain lane centerline detection result information, such as the lane centerline position, etc., based on the point-level lane centerline feature information and instance-level lane centerline feature information corresponding to the target driving scenario image input to the lane centerline decoding module.

[0089] In an implementation manner of the present application, each level of traffic element decoding module can also obtain traffic element detection result information based on the traffic element feature information corresponding to the target driving scene image in the input traffic element decoding module, such as the coordinates, width, height, etc. of the bounding box of the traffic element.

[0090] The driving scene topology inference method based on a machine learning model provided by the present application sets multiple loop modules arranged in a cascaded manner, and a traffic element detection module including a traffic element decoding module and a traffic element and lane centerline topology inference module is set in the loop module, and a lane centerline detection module including a lane centerline decoding module and a lane centerline topology inference module is set. First, each level of lane centerline decoding module performs corresponding processing based on the point-level lane centerline feature information, instance-level lane centerline feature information, and topology inference information input thereto, to obtain updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weights, and instance-to-instance attention weights, and then inputs the updated point-level lane centerline feature information and updated instance-level lane centerline feature information to the next-level lane centerline decoding module for corresponding processing, inputs the updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weights, and instance-to-instance attention weights to the same-level lane centerline topology inference module for corresponding processing, and inputs the updated point-level lane centerline feature information and updated instance-level lane centerline feature information to the same-level traffic element and lane centerline topology inference module. Each level of lane centerline topology inference module performs corresponding processing based on the updated point-level lane centerline feature information, updated instance-level lane centerline feature information, point-to-instance attention weights, instance-to-instance attention weights, and updated traffic element feature information input thereto, to obtain updated lane centerline topology inference information, and inputs the updated lane centerline topology inference information to the next-level lane centerline decoding module for corresponding processing until the lane centerline detection result information corresponding to the target driving scene and the topological relationship between the lane centerlines are obtained. And each level of traffic element decoding module obtains updated traffic element feature information based on the traffic element feature information corresponding to the target driving scene image input to the traffic element decoding module, inputs the updated traffic element feature information to the next-level traffic element decoding module, and inputs the updated traffic element feature information to the same-level traffic element and lane centerline topology inference module. Each level of traffic element and lane centerline topology inference module obtains updated traffic element and lane centerline topology inference information based on the input updated traffic element feature information, as well as the updated point-level lane centerline feature information and updated instance-level lane centerline feature information, until the topological relationship between the traffic element and the lane centerline corresponding to the target driving scene is obtained.

[0091] Thus, it is possible to capture fine-grained point-to-instance relationships and global topological connections based on the point-level lane centerline feature information and instance-level lane centerline feature information corresponding to the target driving scene image in the driving scene, as well as the point-to-instance attention weights and instance-to-instance attention weights, and the traffic element feature information and the traffic element and lane centerline topological inference information. Moreover, through the cyclic enhanced inference of the output information of the traffic element decoding module, the traffic element and lane centerline topological inference module, the lane centerline decoding module, and the lane centerline topological inference module, the accuracy or precision of the topological relationships between the lane centerlines obtained during the vehicle driving process and the topological relationships between the traffic elements and the lane centerlines is effectively improved, thereby improving the accuracy of autonomous driving navigation, decision-making, and other aspects of autonomous driving control, and further enhancing the accuracy and safety of autonomous driving.

[0092] In an implementation manner of the present application, as Figure 3 shown, the machine learning model further includes an initial processing module, and the method further includes obtaining the point-level lane centerline feature information, instance-level lane centerline feature information, and lane centerline topological inference information input to the first-level lane centerline decoding module in the following manner: The initial processing module performs perspective feature conversion processing on the target driving scene image to obtain target perspective feature information, and performs query processing based on the target perspective feature information to obtain the point-level lane centerline feature information, instance-level lane centerline feature information, and lane centerline topological inference information input to the first-level lane centerline decoding module.

[0093] The target perspective feature information can, for example, be the bird's-eye view perspective feature information obtained by extracting the bird's-eye view feature of the target driving scene image.

[0094] In an implementation manner of the present application, the method further includes performing second perspective feature conversion processing on the target driving scene image by the initial processing module to obtain second perspective feature information, and performing traffic element query processing on the second perspective feature information to obtain the traffic element feature information input to the first-level traffic element decoding module.

[0095] Among them, the second perspective feature information can, for example, be the perspective view (PV) feature information.

[0096] In an implementation of the present application, multiple loop modules include a first loop module and a second loop module. The first loop module is the previous-level module of the second loop module. The first loop module includes a first traffic element detection module and a first lane centerline detection module. The second loop module includes a second traffic element detection module and a second lane centerline detection module. The first traffic element detection module includes a first traffic element decoding module and a first traffic element and lane centerline topology inference module. The first lane centerline detection module includes a first lane centerline decoding module and a first lane centerline topology inference module. The second traffic element detection module includes a second traffic element decoding module and a second traffic element and lane centerline topology inference module. The second lane centerline detection module includes a second lane centerline decoding module. The method includes: The first lane centerline decoding module performs corresponding processing based on the first point-level lane centerline feature information, the first instance-level lane centerline feature information, and the first lane centerline topology inference information input into the first lane centerline decoding module, to obtain second point-level lane centerline feature information, second instance-level lane centerline feature information, a first point-to-instance attention weight, and a first instance-to-instance attention weight. The second point-level lane centerline feature information and the second instance-level lane centerline feature information are input into the second lane centerline decoding module for corresponding processing. The second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight are input into the first lane centerline topology inference module for corresponding processing, and the second point-level lane centerline feature information and the second instance-level lane centerline feature information are input into the first traffic element and lane centerline topology inference module; The first lane centerline topology inference module performs corresponding processing based on the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight input into the first lane centerline topology inference module, to obtain second lane centerline topology inference information, and the second lane centerline topology inference information is input into the second lane centerline decoding module; The first traffic element decoding module obtains second traffic element feature information based on the first traffic element feature information corresponding to the target driving scene image input into the first traffic element decoding module, the second traffic element feature information is input into the second traffic element decoding module, and the second traffic element feature information is input into the first traffic element and lane centerline topology inference module; The first traffic element and lane centerline topology inference module obtains second traffic element and lane centerline topology inference information based on the second traffic element feature information input into the first traffic element and lane centerline topology inference module, and the second point-level lane centerline feature information and the second instance-level lane centerline feature information.

[0097] Furthermore, the multiple loop modules may also include a third, fourth, etc. loop module, and the above process is repeated until the topological relationship between the traffic elements corresponding to the target driving scene and the lane centerline and the topological relationship between the lane centerlines are obtained.

[0098] In one implementation of the present application, Figure 4 As shown, each lane centerline decoding module includes a point-level processing unit, an instance-level processing unit, a point-to-instance integration unit, and a relation decoder unit. The first lane centerline decoding module performs corresponding processing based on the first point-level lane centerline feature information, the first instance-level lane centerline feature information and the first lane centerline topological reasoning information input into the first lane centerline decoding module to obtain the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight and the first instance-to-instance attention weight, including: the point-level processing unit obtains the second point-level lane centerline feature information according to the first point-level lane centerline feature information and the preset point-to-point relationship matrix; the relationship decoder unit obtains the point-to-instance relationship information and the instance-to-instance relationship information according to the first lane centerline topological reasoning information; the point-to-instance integration unit obtains the first point-to-instance attention weight according to the second point-level lane centerline feature information, the first instance-level lane centerline feature information and the point-to-instance relationship information, and determines the distance transformation centerline mask according to the second point-level lane centerline feature information; the instance-level processing unit obtains the second instance-level lane centerline feature information and the first instance-to-instance attention weight according to the first instance-level lane centerline feature information, the distance transformation centerline mask and the instance-to-instance relationship information.

[0099] In one implementation of the present application, during the detection of the centerline of the target driving scene, the segmentation level expression of the target centerline can be obtained by calculating the formula of the distance transformation centerline mask:

[0100]

[0101] Among them, b is the pixel point corresponding to the target perspective feature information, C is the lane centerline point c corresponding to the second point-level lane centerline feature information i The set of components, D(b) is the Euclidean distance, L width is the width of the lane centerline, M(b) is the activation function, and RD(b) is the distance transformed centerline mask.

[0102] In one implementation of the present application, Figure 5As shown in the figure, the point-level processing unit includes a line perception attention sub-unit and a cross-attention sub-unit. The point-level processing unit obtains second point-level lane centerline feature information according to the first point-level lane centerline feature information and a preset point-to-point relationship matrix, including: the line perception attention sub-unit performs corresponding processing on the first point-level lane centerline feature information and the preset point-to-point relationship matrix to obtain intermediate point-level lane centerline feature information; the cross-attention sub-unit performs corresponding processing on the intermediate point-level lane centerline feature information to obtain second point-level lane centerline feature information.

[0103] In one implementation manner of the present application, as Figure 6 shown in the figure, the point-to-instance integration unit includes an integration attention sub-unit. The point-to-instance integration unit obtains a first point-to-instance attention weight according to the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information, including: the integration attention sub-unit performs corresponding processing on the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information to obtain the first point-to-instance attention weight.

[0104] In one implementation manner of the present application, as Figure 7 shown in the figure, the instance-level processing unit includes a topology perception attention sub-unit and a mask attention sub-unit. The instance-level processing unit obtains second instance-level lane centerline feature information and a first instance-to-instance attention weight according to the first instance-level lane centerline feature information, the distance transformation centerline mask, and the instance-to-instance relationship information, including: the topology perception attention sub-unit performs corresponding processing on the first instance-level lane centerline feature information and the distance transformation centerline mask to obtain intermediate instance-level lane centerline feature information; the mask attention sub-unit performs corresponding processing on the intermediate instance-level lane centerline feature information and the instance-to-instance relationship information to obtain second instance-level lane centerline feature information and the first instance-to-instance attention weight.

[0105] In one implementation manner of the present application, as Figure 8As shown in the figure, the lane centerline topology inference module includes a point-to-instance topology inference unit, an instance-to-instance topology inference unit, and a connection unit. The first lane centerline topology inference module obtains the second lane centerline topology inference information based on the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight input into the first lane centerline topology inference module, including: the point-to-instance topology inference unit obtains the point-to-instance topology inference information according to the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information; the instance-to-instance topology inference unit obtains the instance-to-instance topology inference information according to the second instance-level lane centerline feature information and the first instance-to-instance attention weight; the connection unit obtains the second lane centerline topology inference information according to the point-to-instance topology inference information and the instance-to-instance topology inference information.

[0106] In an implementation manner of the present application, as Figure 9 shown in the figure, the point-to-instance topology inference unit includes a first connection subunit and multiple first neural network subunits. The point-to-instance topology inference unit obtains the point-to-instance topology inference information according to the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information, including: each first neural network subunit performs corresponding processing on the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information to obtain the first intermediate feature information; the first connection subunit performs connection processing on the first intermediate feature information to obtain the point-to-instance topology inference information. Of course, it may also include multiple first connection subunits, and perform pairwise connection or connection in other ways on different first intermediate feature information.

[0107] In an implementation manner of the present application, as Figure 10 shown in the figure, the instance-to-instance topology inference unit includes a second connection subunit and multiple second neural network subunits. The instance-to-instance topology inference unit obtains the instance-to-instance topology inference information according to the second instance-level lane centerline feature information and the first instance-to-instance attention weight, including: each second neural network subunit performs corresponding processing on the second instance-level lane centerline feature information and the first instance-to-instance attention weight to obtain the second intermediate feature information; the second connection subunit performs connection processing on the second intermediate feature information to obtain the instance-to-instance topology inference information. Of course, it may also include multiple second connection subunits, and perform pairwise connection or connection in other ways on different second intermediate feature information.

[0108] In an implementation manner of the present application, the initial processing module includes a shared backbone network unit and a view transformation unit. The initial processing module performs image perspective feature transformation processing on the target driving scenario image to obtain target perspective feature information, including: the shared backbone network unit performs feature extraction processing on the target driving scenario image to obtain third intermediate feature information; the view transformation unit performs feature transformation processing on the third intermediate feature information to obtain target perspective feature information.

[0109] The machine learning model provided by the implementation manner of the present application can be obtained based on model training and is used to implement driving scenario topology reasoning in the above manner. As described above, performing driving scenario topology reasoning based on this machine learning model can effectively improve the accuracy or precision of the driving scenario topology reasoning result, thereby improving the accuracy of aspects such as autonomous driving navigation and decision-making in autonomous driving control, and further improving the accuracy and safety of autonomous driving.

[0110] The machine learning model provided by the present application and the driving scenario topology reasoning method based on the machine learning model can also be called TopoHR, which is a new end-to-end topology reasoning framework based on a cyclic interaction structure and hierarchical centerline representation. Different from the cascade architecture in the prior art that heavily relies on the performance of the detector, in the TopoHR provided by the present application, the detector (i.e., the lane centerline decoding module) and the lane centerline topology reasoning module can enhance each other.

[0111] As Figure 11 shown, the present application introduces a cyclic reasoning framework for a machine learning model, which can also be called the TopoHR framework. Among them, the self-attention weights (such as point-to-instance attention weights or instance-to-instance attention weights) from the centerline decoder (i.e., the lane centerline decoding module) in the detector (i.e., the lane centerline detection module) are input into the topology reasoning module (i.e., the lane centerline topology reasoning module) as feedforward signals, and the topological relationships (such as lane centerline topological relationships) from the topology reasoning module are fed back to the centerline decoder in the detector. In the detector, multiple centerline representations (for example, point queries, instance queries, and semantic instances) are integrated, and a dedicated module is designed for feature interaction.

[0112] As Figure 12 shown, first, in order to capture both local and global features simultaneously, the present application designs a point-to-instance interaction module to facilitate information exchange between point-level queries and instance-level queries. Then, the present application introduces a rasterized distance transform (R-DT) segmentation method to extract the spatial information of the centerline, and then uses it for instance-level query interaction through a masked attention module. For the topology reasoning module, the present application designs a hierarchical topology reasoning module to capture fine-grained point-to-instance and global instance-to-instance topological connections, and these connections are fused to predict the final topological relationship.

[0113] Based on the driving scenario topology inference method based on a machine learning model provided in this application, the performance of the proposed method was evaluated on the OpenLane-V2 dataset. The results show that there are significant improvements compared with existing methods under similar configurations. Specifically, the method achieved 37.7 in the TOP ll in subset A segmentation and was 10.1 higher than the previous state-of-the-art model in the TOP ll .

[0114] TopoHR provided in this application is a novel end-to-end framework for centerline detection and topology inference, where the detection and inference modules iteratively interact to enhance each other's performance in a loop. A hierarchical representation of the centerline is introduced, in which multi-level representations are seamlessly integrated and fused in the hierarchical centerline decoder. In addition, a hierarchical topology inference module is proposed to capture fine-grained point-to-instance relationships and global instance-to-instance connections, thereby generating accurate topology results. The method has shown good performance on the OpenLane-V2 dataset and far exceeds the previous state-of-the-art models in terms of topology inference of the centerline.

[0115] Next, the preparatory work related to this application will be described.

[0116] First is the preparation of the online map. Online map solutions can be roughly divided into rasterization-based and vectorization-based methods. HDMapNet pioneered the segmentation-based method by fusing multi-view camera images and LiDAR point clouds through a bird's-eye view (BEV) encoder-decoder architecture and predicting instance-level map elements as semantic masks.

[0117] To reduce the dependence on multi-modal data, MGMap proposed a pure camera framework with multi-granularity decoding, achieving hierarchical segmentation of map components. Mask2Map enhanced point-level feature extraction by incorporating deformable attention into instance segmentation. BLOS-BEV rasterized navigation data and integrated it with BEV features. P-MapNet pre-fused prior information from SD maps and HD maps and adopted Masked Autoencoders (MAE) as a refinement module to address occlusion and artifact problems.

[0118] For vectorization-based methods, VectorMapNet introduces a two-stage approach that combines polyline detection with geometric refinement. MapTR proposes a unified permutation-equivariant modeling strategy to achieve end-to-end vectorized map learning without explicit point ordering constraints. HIMap unifies hierarchical queries for joint detection of lanes and traffic elements, demonstrating robust generalization across different datasets. PivotNet models lane lines with pivot points, enhancing the representation of geometric structures. StreamMapNet improves temporal consistency by combining historical BEV features with cross-frame attention.

[0119] Then comes topological detection and reasoning. Most topological reasoning methods use sequential pipelines where a centerline detector first extracts the centerline, followed by a topological reasoning module. For centerline detection, CenterLineDet represents the centerline as vertices and utilizes temporal feature fusion for multi-camera perception. Besides point set formulations, centerline detection also combines methods such as instance segmentation and curve parameterization.

[0120] For example, STSU suggests using Bezier curves for centerline detection, and TopoMask enhances centerline detection by introducing instance segmentation and adding direction prediction to parse the vectorization results. On this basis, TopoMaskV2 combines instance-level Bezier curves with segmentation modeling and uses centerline segmentation results to optimize Bezier curve prediction. TopoBDA further extends this method by introducing a deformable attention mechanism based on Bezier curve modeling and introducing auxiliary losses through instance segmentation. TopoFormer utilizes the geometric distance between centerlines to guide global information aggregation and models reasonable road structures under a counterfactual intervention layer. Additionally, SMERF and TopoSD suggest using SD-Map to enhance the centerline detector.

[0121] In terms of topological reasoning, TopoNet first uses two MLP layers to reduce the embedding dimension of each instance, then sends the concatenated features into another MLP and uses sigmoid activation to predict the relationships between them. TopoMLP emphasizes the "detect first, reason later" strategy and designs a high-performance detector and MLP that contains implicit position embedding features.

[0122] TopoLogic proposes a topology-centric loss function that explicitly preserves lane connections in the segment output, LaneSegNet embeds topological affinity fields into instance segmentation to achieve real-time lane graph extraction, and Topo2Seq converts the graph topological relationships of traffic scenes into a serial representation and uses hierarchical transformers for multi-scale topological reasoning.

[0123] Finally, there is the distance transform. The Distance Transform (DT) has been widely applied in various computer vision tasks. In semantic segmentation, DT guides the segmentation of tubular structures by leveraging geometric properties.

[0124] Similarly, in object detection, DATNet enhances the model's ability to perceive spatial information by predicting the distance to the object, specifically the nearest instance boundary for each pixel. Extended to the 3D domain, 3DSC utilizes the Euclidean 3D distance transform for action recognition to better capture motion changes. In geospatial applications, DT also demonstrates remarkable practicality. For example, POSTPROC combines a DT-based edge map with a CNN to enhance topological continuity in aerial road detection. Recently, MapVR applies differentiable rasterization to vectorized output and uses distance variation on raster maps to achieve precise geometric perception supervision without introducing additional computations during inference.

[0125] This application uses DT to enhance the centerline representation and improve topological reasoning in driving scenarios.

[0126] Next, in combination with the attached Figures 13 - 16 , the method provided by this application will be further explained and described.

[0127] In one implementation of this application, as Figure 13 shown, the proposed TopoHR framework in this application includes a cyclic architecture, which includes a traffic element decoder (i.e., traffic element decoding module) and a topological reasoning module (i.e., traffic element and lane centerline topological reasoning module) at the L-th layer (i.e., L-level), as well as a hierarchical centerline decoder and a hierarchical topological reasoning module at the L-th layer. The traffic element decoder and topological reasoning module at the L-th layer constitute the aforementioned traffic element detection module, and the hierarchical centerline decoder and hierarchical topological reasoning module at the L-th layer constitute the aforementioned lane centerline detection module. This cyclic architecture consists of the following core components:

[0128] (1) Initial processing module, including a backbone network (i.e., shared backbone network unit, Backbone) and a view transformation module (i.e., PV-to-BEV, view transformation unit). First, the backbone network is used to convert the multi-view image (i.e., target driving scenario image, Input) into PV features as the input to the traffic element decoder, and then the PV features are converted into BEV features through the view transformation module (i.e., PV-to-BEV, view transformation unit) as the input to the hierarchical centerline decoder.

[0129] Furthermore, as Figure 14 shown, the backbone network and the view transformation module can also be referred to as the BEV feature extractor (i.e., BEV Feature Extractor).

[0130] (2) Hierarchical Centerline Decoder (i.e., the lane centerline decoding module), which introduces a hierarchical query representation Q hcl ∈ R N×(P+1)×C , where N, P, and C represent the maximum number of centerline instances, the number of points per centerline, and the number of hidden channels respectively. It consists of two interrelated components: the point-level query Q p ∈ R P×C and the instance-level query Q i ∈ R C . The feature representations at the point and instance levels are enhanced through a series of attention mechanisms and point-to-instance integrators. The hierarchical centerline decoder includes the aforementioned point-level module (Pts-level Module), point-to-instance integration module (Pts-Ins Integrator), relation encoder, instance-level module (Ins-levelModule), and R-DT mask, etc. Through the coordinated use of these modules, the lane centerline output (Centerline Output, i.e., the lane centerline detection result information, such as classification (Class), segmentation (Seg), coordinates (Coords), etc.) is obtained.

[0131] (3) Hierarchical Topology Module (i.e., the lane centerline topology reasoning module), including the point-to-instance topology reasoning module (Pts-Ins Topo Module) and the instance-to-instance topology reasoning module (Ins-Ins Topo Module), which performs comprehensive topology reasoning by integrating the updated hierarchical queries and point-to-instance and instance-to-instance. For the instance attention weights W p2i ∈ R (N×P)×N , W i2i ∈ R N×N , through fine-grained point-level interactions and high-level instance-level relationships, an accurate and semantically rich understanding of the topology structure is achieved. The cyclic interaction mechanism promotes iterative refinement, where the attention weights from the detector are fed into the topology reasoning module, and the instance topology relationships from the reasoning module are fed back into the detector, thus creating a complementary cycle to improve performance, and thus obtaining the centerline topology output (CenterlineTopology Output, that is, the lane centerline topology relationship).

[0132] Unlike the segmentation labels that are unified for all positive samples, the distance transform encodes spatial proximity by mapping Euclidean distances. This method effectively addresses the challenges of centerline segmentation without relying on specific visual features. In our method, instead of directly using the distance transform, the distance field is converted to a structured grid rasterization in the following way to improve computational efficiency and feature representation.

[0133] Formally, given a centerline C = {c1, c2, …, c P}, where c i = (x i , y i ) represents the vectorized point set, that is, the coordinates of the points on the centerline. First, calculate the Euclidean distance from each pixel point b in the BEV feature (i.e., the target perspective feature information) to the nearest centerline point c i ∈ C. The distance value is restricted to within half of the lane width L width . Then, these values are normalized to the range [0, 1] and rasterized into 11 categories with a step size of △ = 0.1 to generate the final rasterized distance transform centerline mask RD(b). The specific process can be implemented by the following formula:

[0134]

[0135] where b is the pixel point in the BEV map, C is the set composed of the given centerline points c i , D(b) is the Euclidean distance, L width is the width of the lane centerline (i.e., the lane width), M(b) is the activation function, and RD(b) is the distance transform centerline mask (i.e., the rasterized distance transform centerline mask).

[0136] Apply the rasterized distance transform mask to the mask attention module and participate in instance matching and loss calculation. For this purpose, the rasterized distance transform mask can better extract the position information of the centerline and converge more easily than the standard regression method.

[0137] (4) Traffic element detection module, including a traffic element decoder (Traffic Element Decoder, i.e., traffic element decoding module) and a topology reasoning module (Topology Reasoning Module, i.e., traffic element and lane centerline topology reasoning module). The traffic element decoder processes the corresponding PV feature information of, for example, the target driving scene image to obtain a traffic element output (Traffic Element Output, i.e., traffic element detection result information, such as traffic element classification (class), coordinates (x, y) of the bounding box (bounding box, bbox) enclosing the traffic element, and width (w) and height (h)). And through the topology reasoning module, a traffic element and centerline topology output (Traffic Element &CenterlineTopology Output, i.e., the topological relationship between the lane centerline and the traffic element) is obtained. Additionally, there can be multiple bounding boxes, so they can be abbreviated as bboxs for short.

[0138] Further, as Figure 15 shown, it is the input and output of the l-th layer of the hierarchical centerline decoder in the TopoHR framework. In the part (a) of the point-level module ( Figure 15 , that is, the point-level processing unit, Pts-level Module), a line-aware attention module (i.e., line-aware attention sub-unit, Line-Aware Attention) is applied to promote intra-instance interaction between point-level queries. Specifically, for each group of point-level queries (e.g., 11 point-level queries for each instance), the line-aware attention operation is limited to operate within each centerline instance. To enforce this constraint, a point-to-point relationship matrix (P2PRelation) M p2p ∈R (N×P)×(N×P) is introduced and used as an attention mask when calculating the line-aware attention. This mask effectively restricts cross-instance interaction by focusing the attention on point-level queries within the same centerline instance, maintains instance-specific feature learning, and prevents information exchange between different centerlines. In one implementation manner of this application, the point-level module also introduces a cross-attention module (i.e., cross-attention sub-unit, Cross-Attention).

[0139] Further, referring to the relationship DETR, the feature representation is enhanced by using intermediate attention maps in the decoder to model the relationships of object instances. There is a natural consistency between this relationship learning framework and the core requirements of the centerline detection and topology reasoning tasks. The inherent structural dependencies and topological constraints in the centerline make it particularly suitable for this relationship-aware feature learning.

[0140] To improve centerline detection using topological reasoning, as Figure 15 shown, this application introduces a Relation Encoder to process the topological predictions from the previous layer (i.e., the topological reasoning information of layer l-1). This application uses an MLP layer to generate two different topological relationships: (1) the point-to-instance topological relationship (i.e., the point-to-instance relationship, P2IRelation) M p2i ∈R (N×P)×N and (2) the inter-instance topological relationship (i.e., the instance-to-instance relationship, I2I Relation) M i2i ∈R N ×N . The point-to-instance relationship and the instance-to-instance relationship are then integrated into an attention mask and fed into the Integrator Attention (i.e., the integrator attention subunit) and the Topo-Aware Attention (i.e., the topo-aware attention subunit). This application establishes a complementary cycle between the centerline detection and the topological reasoning modules - the components iteratively improve the performance of each other through information exchange and joint optimization.

[0141] The hierarchical framework facilitates multi-scale interactions between the point-wise geometric features and the instance-level semantics in centerline modeling. Through the local-global attention architecture, the feature representation is iteratively improved by fusing the position details with the semantic context, thus enabling coherent feature propagation across hierarchical spaces. As Figure 15 shown in part (b) (the Pts-InsIntegrator, the point-to-instance integration unit), the hierarchical integrator attention module operates in two cross-level steps. Given the point-level query (i.e., the second point-level lane centerline feature information) Q Figure 15 updated from the point-level module (as shown in part (a) of p,l ∈R N×P×C and the instance-level query (i.e., the first instance-level lane centerline feature information) Q i,l-1 ∈R NxC generated from the (l-1)th layer, the integrator attention adopts a cross-attention mechanism. Specifically, the point-level query Q p,l acts as Q, while the instance-level query Q i,l-1 acts as KV. The specific formula for point-to-instance integration is as follows:

[0142]

[0143]

[0144] where M p2iIndicating the point-to-instance relationship, the attention module enables each point-level query to selectively focus on relevant instance-level features, thereby enriching its representation with local and global contexts. Then, through the learnable coefficient W Agg ∈R P aggregate the refined point-level queries into updated instance-level queries:

[0145]

[0146] To further enhance instance-aware feature learning and reasoning, the updated instance-level queries are then fed into the instance-level module ( Figure 15 as shown in part (c) of, the instance-level processing unit, Ins-level Module). This approach ensures that the point-level queries enrich both local position details and global semantic contexts, while the instance-level queries are dynamically updated by aggregating the refined point-level features.

[0147] In one implementation of the present application, as Figure 15 shown in part (c) of, the instance-level module includes a Topo-Aware Attention sub-unit and a Masked-Attention sub-unit. By processing the updated instance-level queries, the R-DT mask, and the instance-to-instance relationship, updated instance-level queries and instance-to-instance attention weights are obtained.

[0148] Existing topology inference methods mainly rely on instance-level query interactions modeled through simplified MLP layers. The present application proposes a hierarchical representation paradigm that explicitly integrates point-to-instance and instance-to-instance topology relationships. The present application emphasizes the hierarchical establishment of topology relationships: when considering the topology relationship between two centerlines C i and C j this relationship is not only reflected in the instance-level representations on the two centerlines, but also in the point-level representation of C i and the instance-level representation of C j which is crucial.

[0149] Specifically, as Figure 16As shown, the hierarchical module proposed in this application includes a dual-prediction branch: (1) The instance-to-instance topology module (i.e., the instance-to-instance topology reasoning unit, Ins-Ins Topology Reasoning) uses a dual MLP encoding to calculate similarity using instance-level queries and their self-attention weights (such as calculating the inner product through the second connection subunit), and extracts the association relationship through an attention-based MLP. (2) The point-to-instance topology module (i.e., the point-to-instance topology reasoning unit, Pts-Ins Topology Reasoning) combines explicit feature correlation and latent dependence for robust topology reasoning. This design enables comprehensive topology reasoning by effectively integrating instance-level features and their hidden dependencies.

[0150] Specifically, the instance-to-instance topology module performs a connection (such as an inner product operation) based on the instance-to-instance attention weight and the instance-level query to obtain an updated instance-level query. The point-to-instance topology module performs a connection (such as an inner product operation) based on the point-to-instance attention weight, the instance-level query, and the click query to obtain an updated second instance-level query. Finally, hierarchical topology reasoning is performed based on the updated instance-level query and the updated second instance-level query to obtain topology reasoning information.

[0151] The aforementioned instance-to-instance topology reasoning process is shown as follows:

[0152]

[0153]

[0154]

[0155] The point-to-instance topology prediction follows a similar workflow, applying an average operation along the point dimension to generate the point-to-instance topology prediction:

[0156]

[0157]

[0158]

[0159] The final hierarchical topology reasoning result comes from the results of two hierarchical components T i2i and T p2i of.

[0160] Next, the training loss of this application will be described.

[0161] This application proposes an adaptive topological loss for topological inference supervision, replacing the traditional focal loss used in TopoLogic. The framework adopts a dynamic weighting strategy based on reparameterized cross-entropy, where the negative sample weights follow an exponential scaling e λneg·xi , x i represents the predicted positive probability, while the positive samples maintain a fixed weight λ pos . This mechanism creates an adaptive gradient modulation, which proportionally amplifies the penalty for negative samples showing high confidence scores, effectively reducing false positive predictions while maintaining topological consistency in the feature space. Overall, the training loss formula can be written as:

[0162]

[0163] where the centerline detection loss consists of a focal loss for centerline instance classification and an L1 loss for vectorizing the centerline result regression. Combines dice loss and cross-entropy loss to guide the learning of instance-level features based on the rasterized distance transform centerline mask. The topological inference loss is the proposed adaptive topological loss.

[0164] Next, combined with specific experimental data, the advantages of the method provided by this application are described.

[0165] TopoHR provided by this application was evaluated on the OpenLane-V2 benchmark, which is a comprehensive dataset integrating Argoverse 2 and nuScenes. The dataset contains 2000 scenes, divided into subset A and subset B, with multi-view Figure 2 Hz images and annotations of 3D centerlines, traffic elements, and their topological relationships. Subset A contains seven camera views, while subset B contains six camera views. This application uses the following official evaluation metrics. The evaluation metrics include:

[0166] DET l (average Frechet distance across matching thresholds), DET t (IoU-based traffic element similarity), TOP ll (centerline topological matrix similarity), and TOP lt (centerline traffic element topological similarity). The OLS metric calculates the average of these multi-task metrics.

[0167] This application uses ResNet50 as the backbone and utilizes the Feature Pyramid Network (FPN) to obtain multi-scale features with an input resolution of 1550×2048. 200 centerlines and 100 traffic element queries are initialized for detection and topological relationship reasoning. Following BEVFormer, this application projects the image features into a predefined BEV space with a grid resolution of 200×100. In the centerline detector, the regression head consists of a 3-layer MLP with LayerNorm and ReLU activation, and each centerline outputs a 3D position offset of 11×3.

[0168] The topological reasoning method proposed in this application mainly focuses on modeling the topological relationships of centerlines. The settings in TopoLogic are followed to detect traffic elements and calculate the topological relationships between centerlines and traffic elements. The TopoHR model is trained on 8 NVIDIA 4090 GPUs with a total batch size of 8 for 24 epochs. For optimization, the AdamW optimizer is adopted with an initial learning rate of , and a weight decay of 0.01.

[0169] Compared with previous state-of-the-art methods such as TopoMaskV2 and TopoBDA, TopoHR provided by this application achieves competitive performance without using any additional training tricks or optimizations (one-to-many centerline queries; different backbones and denoising strategies for traffic element detection) in the centerline detector or traffic element detector.

[0170] Specifically, as Figure 17 shown, it is the performance comparison between TopoHR provided by this application and other state-of-the-art methods in the OpenLane-V2 subset A benchmark. The best results are highlighted in bold, the second-best results are underlined, and the third-best results are italicized.

[0171] Among other state-of-the-art methods, they include: Method 1 (i.e., Map Transformer (MapTR), corresponding conference is Conference 1 (i.e., The International Conference on Learning Representations (ICLR) 2023)), Method 2 (i.e., TopoNet, corresponding conference is Conference 2 (i.e., Arxiv) 2023), Method 3 (i.e., SMERF (SD Map Encoding Representation from Transformers), corresponding conference is Conference 3 (i.e., International Conference on Robotics and Automation (ICRA) 2024)), Method 4 (i.e., TopoLogic, corresponding conference is Conference 4 (i.e., Conference and Workshop on Neural Information Processing Systems (NeurIPS) 2024)), Method 5 (i.e., TopoFormer, corresponding conference is Conference 5 (i.e., Arxiv 2024)), Method 6 (i.e., TopoMaskV2, corresponding conference is Conference 5 (i.e., Arxiv 2024)), Method 7 (i.e., TopoBDA, corresponding conference is Conference 5 (i.e., Arxiv 2024)).

[0172] Among them, the Standard Definition Map (SDMap) is a one-to-many query in the Lane Branch, and the Object-to-Map (O2M) is an additional one-to-many query in the Lane Branch; the dynamic threshold detection (i.e., Differentiable Binarization (DB) and optimized clustering (improved K-means), i.e., DBKB is a different backbone in the Traffic Branch; DN is the denoising strategy for the Traffic Branch. The method of this solution - 1 refers to the TopoHR method, and the method of this solution - 2 refers to the TopoHR-L method, where -L extends the number of TopoHR centerline queries to 600.

[0173] The TopoHR-L provided by this application has the highest classification accuracy TOP among the evaluation parameters ll and TOP ltThe best results of 37.7 and 38.2 were respectively obtained on [specific items], and the second-best result of 51.6 was obtained on Ordinary Least Squares (OLS). The provided TopoHR achieved state-of-the-art (SOTA) performance without a dedicated training strategy: it was 3.7 higher than TopoFormer in OLS, 8.2 higher in [TOP item 1] (a 34.0% improvement), and even 4.7 higher than TopoBDA in [TOP item 2] (a 17.0% improvement). The scaled variant TopoHR-L (with 600 centerline queries) further improved [TOP item 2] by 10.1 (a 36.6% improvement) compared to the previous SOTA TopoBDA. The excellent DET and OLS scores achieved by TopoMaskV2 and TopoBDA are mainly attributed to their advanced traffic element decoders, which adopt the DAB-DETR and DINO (Distillation in Noise) architectures. l and TOP ll The second-best results of 36.0 and 32.7 were respectively obtained on [items].

[0174] TopoHR achieved state-of-the-art (SOTA) performance without a dedicated training strategy: it was 3.7 higher than TopoFormer in OLS, and in [TOP item 1] ll it was 8.2 higher (a 34.0% improvement), and even in [TOP item 2] ll it was 4.7 higher than TopoBDA (a 17.0% improvement). The scaled variant TopoHR-L (with 600 centerline queries) further improved [TOP item 2] ll by 10.1 (a 36.6% improvement) compared to the previous SOTA TopoBDA. The excellent DET t and OLS scores achieved by TopoMaskV2 and TopoBDA are mainly attributed to their advanced traffic element decoders, which adopt the DAB-DETR and DINO (Distillation in Noise) architectures.

[0175] As Figure 18 shown, the performance comparison of the TopoHR provided in this application (i.e., ours) with other state-of-the-art methods in the OpenLane-V2 subset B benchmark test. The best results are highlighted in bold, the second-best results are underlined, and the third-best results are italicized.

[0176] Among them, other state-of-the-art methods include: Method 2 (i.e., TopoNet, corresponding conference is Conference 2 (i.e., Arxiv2023)), Method 4 (i.e., TopoLogic, corresponding conference is Conference 4 (i.e., NeurIPS 2024)), Method 5 (i.e., TopoFormer, corresponding conference is Conference 5 (i.e., Arxiv 2024)), Method 6 (i.e., TopoMaskV2, corresponding conference is (i.e., Arxiv 2024)), Method 7 (i.e., TopoBDA, corresponding conference is Conference 5 (i.e., Arxiv 2024)).

[0177] The TopoHR-L provided in this application in the evaluation parameters [TOP item 1] ll and [TOP item 2] ltThe best results of 38.9 and 28.2 were achieved respectively. The proposed Solution-1 (i.e., TopoHR) achieved the second-best result of 51.9 on the evaluation parameter OLS. TopoHR performs well in topological reasoning, and on the evaluation parameter TOP ll scored 32.7 points, outperforming Method 6 (i.e., TopoMaskV2). The model of Solution-2 (i.e., TopoHR-L) has made a significant improvement compared to TopoBDA. On the evaluation parameter TOP ll it increased by 4.9, indicating the robustness and adaptability of the TopoHR framework provided in this application on different datasets.

[0178] Furthermore, this application conducted an ablation study. The impact of the key components proposed in this application on the OpenLane-V2 subset A dataset was evaluated.

[0179] First is the design of the hierarchical centerline representation. As Figure 19 shown, the ablation study of the centerline mask (Mask GT) representation, where MA represents the instance-level masked attention; the distance transform (DT) represents the distance transform mask; R-DT represents the rasterized distance transform mask. The ablation study results are as follows: (1) Adding 0 / 1 instance segmentation losses improved the evaluation results of the evaluation parameters TOP ll and TOP lt by 2.0 and 1.8 respectively (lines 1-2); (2) The instance-level masked attention enhanced the query feature interaction (line 3); compared with the baseline, the rasterized distance transform GT mask provided in this application achieved the maximum gains of 4.8 and 3.5 on TOP ll and TOP lt respectively, verifying the effectiveness of the hierarchical representation in capturing local and global structures.

[0180] Then is the effect of the hierarchical attention relationship modeling. As Figure 20 shown, the ablation study verified the impact of hierarchical attention: the baseline model (without a relationship encoder, masked attention, or topological module) achieved, for example, DET l = 28.6 and TOP ll = 26.3. Introducing point-level constrained attention (i.e., point-to-point, P2P) (enabling intra-instance interaction) improved DET l to 33.3 (+4.7 improvement). Gradually integrating the point-to-instance (i.e., point-to-instance, P2I) and instance-level (i.e., instance-to-instance, I2I) topological relationships generated by the relationship encoder further improved DET l to 33.6 (a 5.0 improvement compared to the baseline), demonstrating incremental precision gains through hierarchical reasoning.

[0181] Second, the design of hierarchical topology reasoning. As Figure 21 shown, it demonstrates the effectiveness of hierarchical topology reasoning: integrating instance-level attention in the TopoHead and attention weights can improve DET l and TOP ll metrics. Hierarchical modeling combines instance-level and point-level interactions, achieving gains of 1.4 and 0.7 respectively on the DET l and TOP ll baselines. This confirms the benefits of hierarchical modeling for centerline detection and topology.

[0182] Finally, the impact of adaptive topology loss. The effectiveness of the Adaptive Topology Loss (ATL) in TopoHR is as Figure 22 shown. Replacing the standard Focal Loss (FC) with the adaptive topology loss (i.e., ATL) proposed in this application, the TOP ll and TOP lt evaluation results are improved by 0.4 and 1.3 respectively.

[0183] It should be noted that Figures 17 - 20 in, drawing a "√" indicates that the corresponding existing method is used to experiment on the target branch. For example, Figure 17 the "√" in the first row in indicates that the MapTR method published at the conference ICLR 2023 is used to experiment on the O2M branch in the lane centerline detection branch and obtain the evaluation metrics.

[0184] The TopoHR proposed in this application is an end-to-end topology reasoning method that integrates recurrent detector topology interaction and hierarchical centerline representation. This architecture overcomes the limitations of sequential pipelines by enabling the co-evolution of the detection module and the topology reasoning module, while the hierarchical representation fuses multi-level centerline position features through rasterized distance transformation. The unified hierarchical topology module simultaneously captures fine-grained point-to-instance relationships and global topology connections. Extensive experiments show that TopoHR achieves excellent topology reasoning performance on the OpenLane-V2 benchmark.

[0185] Thus, based on the TopoHR proposed in this application, the accuracy or precision of the topology reasoning results in the driving scenario is effectively improved. Therefore, when applied to the autonomous driving system, based on the more accurate or higher-precision driving scenario topology reasoning results, the accuracy of autonomous driving control such as autonomous driving navigation and decision-making can be improved, thereby enhancing the accuracy and safety of autonomous driving.

[0186] Furthermore, the implementation manner of the present application also proposes a vehicle control method, which includes: obtaining the topological inference result information of the target driving scenario through the above-mentioned TopoHR, and controlling the vehicle to travel according to the topological inference result information of the target driving scenario to achieve more accurate and safe autonomous driving.

[0187] Please refer to Figure 23 , Figure 23 FIG. shows a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 23 shown, the electronic device may include: a transceiver 121, a processor 122, and a memory 123.

[0188] The processor 122 executes the computer program / instructions stored in the memory, so that the processor 122 executes some technical solutions of the driving scenario topological inference method based on the machine learning model in the above embodiment. The processor 122 may be a general-purpose processor, including a central processing unit (CPU), a network processor unit (NPU), etc.; it may also be a digital data processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0189] The memory 123 is connected to the processor 122 through a system bus and completes communication with each other. The memory 123 is used to store computer program instructions.

[0190] By way of example and not limitation, the memory 123 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 123 may include removable or non-removable (or fixed) media. Where appropriate, the memory 123 may be internal or external to the integrated gateway device. In a particular embodiment, the memory 123 is a non-volatile solid-state memory. In a particular embodiment, the memory 123 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0191] The transceiver 121 can be used to obtain the task to be run and the configuration information of the task to be run.

[0192] The implementation mode of this application also provides a chip for running computer instructions / programs, and this chip is used to execute the technical solution of the above-mentioned driving scenario topology inference method based on a machine learning model.

[0193] The implementation mode of this application also provides a computer-readable storage medium, and computer instructions / programs are stored in this computer-readable storage medium. When the computer instructions / programs run on the processor of an electronic device, the processor of the electronic device is caused to execute the technical solution of the above-mentioned driving scenario topology inference method based on a machine learning model.

[0194] The implementation mode of this application also provides a computer program product, and this computer program product includes computer programs / instructions, which are stored in a computer-readable storage medium. At least one processor can read the computer programs / instructions from the computer-readable storage medium, and when at least one processor executes the computer programs / instructions, the technical solution of the above-mentioned driving scenario topology inference method based on a machine learning model can be implemented.

[0195] Those skilled in the art will readily conceive of other embodiments of this application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application, which follow the general principles of this application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in this application.

[0196] It should be noted that, in addition to the implementation manners of the present application described in the above specific embodiments, those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Although the description of the present application is introduced in combination with preferred embodiments, this does not mean that the features of this invention are limited to this implementation manner. On the contrary, the purpose of introducing the invention in combination with the implementation manner is to cover other alternatives or modifications that may be extended based on the implementation manner of the present application. In order to provide a deep understanding of the present application, many specific details are included in the above description, and the present application can also be implemented without using these details. In addition, in order to avoid confusion or obscuring the key points of the present application, some specific details will be omitted in the description. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0197] It should be noted that in this specification, similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0198] Although the present application has been illustrated and described by referring to some preferred implementation manners of the present application, those of ordinary skill in the art should understand that the above content is a further detailed description of the present application in combination with specific implementation manners, and it cannot be determined that the specific implementation of the present application is only limited to these descriptions. Those skilled in the art can make various changes in form and details, including making several simple deductions or substitutions, without departing from the spirit and scope of the present application.

Claims

1. A driving scenario topology inference method based on a machine learning model, characterized in that The machine learning model includes a plurality of recurrent modules arranged in a cascaded manner. Each of the recurrent modules includes a traffic element detection module and a lane centerline detection module. The traffic element detection module includes a traffic element decoding module and a traffic element and lane centerline topology inference module. The lane centerline detection module includes a lane centerline decoding module and a lane centerline topology inference module. The method includes: Each level of the lane centerline decoding module obtains updated lane centerline feature information and attention weight information based on the input lane centerline feature information and lane centerline topology inference information. The updated lane centerline feature information is input into the lane centerline topology inference module at the same level, the traffic element and lane centerline topology inference module at the same level, and the lane centerline decoding module at the next level. The attention weight information is input into the lane centerline topology inference module at the same level. Each level of the lane centerline topology inference module obtains updated lane centerline topology inference information based on the input updated lane centerline feature information and the attention weight information. The updated lane centerline topology inference information is input into the lane centerline decoding module at the next level until the target lane centerline topological relationship is obtained. Each level of the traffic element decoding module obtains updated traffic element feature information based on the input traffic element feature information, and the updated traffic element feature information is input into the traffic element decoding module at the next level and the traffic element and lane centerline topology inference module at the same level. Each level of the traffic element and lane centerline topology inference module obtains updated traffic element and lane centerline topology inference information based on the input updated traffic element feature information and the updated lane centerline feature information until the target traffic element and lane centerline topological relationship is obtained.

2. The method for driving scenario topology reasoning based on a machine learning model according to claim 1, wherein The lane centerline feature information includes point-level lane centerline feature information and instance-level lane centerline feature information corresponding to the target driving scenario image. The updated lane centerline feature information includes updated point-level lane centerline feature information and updated instance-level lane centerline feature information.

3. The method for driving scenario topology inference based on a machine learning model according to claim 2, wherein The attention weight information includes point-to-instance attention weights and instance-to-instance attention weights.

4. The method for driving scenario topology inference based on a machine learning model according to claim 3, wherein The machine learning model further includes an initial processing module. The method further includes obtaining the point-level lane centerline feature information, the instance-level lane centerline feature information, and the lane centerline topology inference information input into the lane centerline decoding module at the first level in the following manner: The initial processing module performs a perspective feature conversion process on the target driving scenario image to obtain target perspective feature information, and performs a query process based on the target perspective feature information to obtain the point-level lane centerline feature information, the instance-level lane centerline feature information, and the lane centerline topology inference information input into the lane centerline decoding module at the first level.

5. The method for driving scenario topology inference based on a machine learning model according to claim 4, characterized in that, The multiple loop modules include a first loop module and a second loop module. The first loop module is the previous-level module of the second loop module. The first loop module includes a first traffic element detection module and a first lane centerline detection module. The second loop module includes a second traffic element detection module and a second lane centerline detection module. The first traffic element detection module includes a first traffic element decoding module and a first traffic element and lane centerline topology inference module. The first lane centerline detection module includes a first lane centerline decoding module and a first lane centerline topology inference module. The second traffic element detection module includes a second traffic element decoding module and a second traffic element and lane centerline topology inference module. The second lane centerline detection module includes a second lane centerline decoding module. The method includes: Based on the input first point-level lane centerline feature information, first instance-level lane centerline feature information, and first lane centerline topology inference information, the first lane centerline decoding module obtains second point-level lane centerline feature information, second instance-level lane centerline feature information, first point-to-instance attention weights, and first instance-to-instance attention weights, and inputs the second point-level lane centerline feature information and the second instance-level lane centerline feature information to the first lane centerline topology inference module, the first traffic element and lane centerline topology inference module, and the second lane centerline decoding module, and inputs the first point-to-instance attention weights and the first instance-to-instance attention weights to the first lane centerline topology inference module; Based on the input second point-level lane centerline feature information, second instance-level lane centerline feature information, first point-to-instance attention weights, and first instance-to-instance attention weights, the first lane centerline topology inference module obtains second lane centerline topology inference information and inputs the second lane centerline topology inference information to the second lane centerline decoding module; Based on the input first traffic element feature information, the first traffic element decoding module obtains second traffic element feature information and inputs the second traffic element feature information to the second traffic element decoding module and the first traffic element and lane centerline topology inference module; Based on the input second traffic element feature information, and the second point-level lane centerline feature information and the second instance-level lane centerline feature information, the first traffic element and lane centerline topology inference module obtains second traffic element and lane centerline topology inference information.

6. The method for driving scenario topology inference based on a machine learning model according to claim 5, characterized in that, Each of the lane centerline decoding modules includes a point-level processing unit, an instance-level processing unit, a point-to-instance integration unit, and a relationship decoder unit. Then, based on the first point-level lane centerline feature information, the first instance-level lane centerline feature information, and the first lane centerline topology inference information input into the first lane centerline decoding module, the first lane centerline decoding module obtains the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight, including: The point-level processing unit obtains the second point-level lane centerline feature information according to the first point-level lane centerline feature information and a preset point-to-point relationship matrix; The relationship decoder unit obtains point-to-instance relationship information and instance-to-instance relationship information according to the first lane centerline topology inference information; The point-to-instance integration unit obtains the first point-to-instance attention weight according to the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information, and determines a distance-transformed centerline mask according to the second point-level lane centerline feature information; The instance-level processing unit obtains the second instance-level lane centerline feature information and the first instance-to-instance attention weight according to the first instance-level lane centerline feature information, the distance-transformed centerline mask, and the instance-to-instance relationship information.

7. The method for driving scenario topology inference based on a machine learning model according to claim 6, characterized in that The distance-transformed centerline mask is obtained through the following formula: Among them, b is the pixel point corresponding to the target perspective feature information, and C is the lane center line point c corresponding to the second point-level lane center line feature information i is the set composed of, D(b) is the Euclidean distance, and L width is the width of the lane center line, M(b) is the activation function, and RD(b) is the distance transformation center line mask.

8. The driving scenario topology inference method based on a machine learning model according to claim 7, wherein The point-level processing unit includes a line-aware attention sub-unit and a cross-attention sub-unit. The point-level processing unit obtains the second point-level lane centerline feature information according to the first point-level lane centerline feature information and a preset point-to-point relationship matrix, including: The line-aware attention sub-unit performs corresponding processing on the first point-level lane centerline feature information and the preset point-to-point relationship matrix to obtain intermediate point-level lane centerline feature information; The cross-attention sub-unit performs corresponding processing on the intermediate point-level lane centerline feature information to obtain the second point-level lane centerline feature information.

9. The method for driving scenario topology inference based on a machine learning model according to claim 8, wherein, The point-to-instance integration unit includes an integration attention sub-unit. The point-to-instance integration unit obtains the first point-to-instance attention weight according to the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information, including: The integration attention sub-unit performs corresponding processing on the second point-level lane centerline feature information, the first instance-level lane centerline feature information, and the point-to-instance relationship information to obtain the first point-to-instance attention weight.

10. The method for driving scenario topology inference based on a machine learning model according to claim 9, wherein The instance-level processing unit includes a topology-aware attention sub-unit and a mask attention sub-unit. The instance-level processing unit obtains the second instance-level lane centerline feature information and the first instance-to-instance attention weight according to the first instance-level lane centerline feature information, the distance-transformed centerline mask, and the instance-to-instance relationship information, including: The topological perception attention subunit performs corresponding processing on the first instance-level lane centerline feature information and the distance-transformed centerline mask to obtain intermediate instance-level lane centerline feature information; The masked attention subunit performs corresponding processing on the intermediate instance-level lane centerline feature information and the instance-to-instance relationship information to obtain the second instance-level lane centerline feature information and the first instance-to-instance attention weight.

11. The method for driving scenario topology inference based on a machine learning model according to any one of claims 5-10, characterized in that, Each of the lane centerline topology inference modules includes a point-to-instance topology inference unit, an instance-to-instance topology inference unit, and a connection unit. The first lane centerline topology inference module, based on the second point-level lane centerline feature information, the second instance-level lane centerline feature information, the first point-to-instance attention weight, and the first instance-to-instance attention weight input into the first lane centerline topology inference module, obtains second lane centerline topology inference information, including: The point-to-instance topology inference unit obtains point-to-instance topology inference information according to the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information; The instance-to-instance topology inference unit obtains instance-to-instance topology inference information according to the second instance-level lane centerline feature information and the first instance-to-instance attention weight; The connection unit obtains the second lane centerline topology inference information according to the point-to-instance topology inference information and the instance-to-instance topology inference information.

12. The method for driving scenario topology inference based on a machine learning model according to claim 11, characterized in that, The point-to-instance topology inference unit includes at least one first connection subunit and multiple first neural network subunits. The point-to-instance topology inference unit obtains point-to-instance topology inference information according to the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information, including: Each of the first neural network subunits performs corresponding processing on the second point-level lane centerline feature information, the first point-to-instance attention weight, and the second instance-level lane centerline feature information to obtain first intermediate feature information; The first connection subunit performs connection processing on the first intermediate feature information to obtain the point-to-instance topology inference information.

13. The method for driving scenario topology inference based on a machine learning model according to claim 12, characterized in that, The instance-to-instance topology inference unit includes at least one second connection subunit and multiple second neural network subunits. The instance-to-instance topology inference unit obtains instance-to-instance topology inference information according to the second instance-level lane centerline feature information and the first instance-to-instance attention weight, including: Each of the second neural network subunits performs corresponding processing on the second instance-level lane centerline feature information and the first instance-to-instance attention weight to obtain second intermediate feature information; The second connection subunit performs connection processing on the second intermediate feature information to obtain the instance-to-instance topology inference information.

14. The method for driving scenario topology reasoning based on a machine learning model according to claim 4, wherein The initial processing module includes a shared backbone network unit and a view conversion unit. The initial processing module performs perspective feature conversion processing on the target driving scenario image to obtain target perspective feature information, including: The shared backbone network unit performs feature extraction processing on the target driving scenario image to obtain third intermediate feature information; The view conversion unit performs feature conversion processing on the third intermediate feature information to obtain the target perspective feature information.

15. An electronic device, characterized in that, Including: A processor, and a memory communicatively connected to the processor; The memory stores a computer program; The processor executes the computer program stored in the memory, so that the electronic device implements the driving scenario topology inference method based on a machine learning model according to any one of claims 1-14.

Citation Information

Patent Citations

  • Topological reasoning method, system and equipment for driving scene and storage medium

    CN116386009A

  • Static scene representation method and system based on instance level expression

    CN117237918A

  • Method and system for establishing lane line map and extracting topological structure

    CN117671494A

  • Online high-precision vector map generation method based on mask guidance

    CN118247382A

  • Method for realizing BEV local map real-time perception by fusing SD Map

    CN119785311A

Cited By

  • Vehicle driving track determination method and electronic equipment

    CN120840631A