Map detection method and apparatus, and model training method and apparatus

By pre-fusion and decoding of map features and sensor features, the problem of mismatch between map and perception features in existing technologies is solved, achieving higher accuracy in road structure cognition and map detection, adapting to various application scenarios, and reducing the cost of building high-precision maps.

WO2026066420A1PCT designated stage Publication Date: 2026-04-02YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In existing technologies, the fusion of real-time sensing networks and map rules is complex, which affects system performance and limits the accuracy of the fused map. This is especially true for low-precision maps, which can lead to errors or discrepancies in road structure recognition, making it difficult to achieve accurate road structure recognition.

Method used

By pre-fusion of map features and sensor features, road network features are extracted, and a decoder is used to output a higher-precision map. The decoding combines perception features and map features to adapt to various types of scenarios. An attention mechanism is applied to interact and sparsify features, thereby improving decoding efficiency.

Benefits of technology

It achieves higher precision map output, can accurately detect lane and intersection information, reduces the cost of building high-precision maps, adapts to various application scenarios, and improves model training and inference efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025105649_02042026_PF_FP_ABST
    Figure CN2025105649_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a map detection method and apparatus, and a model training method and apparatus, which enable early fusion of map features and sensor features, thus allowing for full extraction of road network features and output of maps having higher accuracy. The map detection method comprises: first, acquiring sensor data, the sensor data comprising data collected by at least one sensor; extracting features from the sensor data to obtain sensing features; simultaneously, acquiring a first map, and extracting features from the first map to obtain first map features; fusing the sensing features and the first map features to obtain a fused feature; and decoding the fused feature to obtain a second map, the accuracy of the second map being higher than the accuracy of the first map, and the second map comprising various types of vectors.
Need to check novelty before this filing date? Find Prior Art

Description

Map detection method and model training method and device

[0001] The present application claims priority to the Chinese patent application No. CN202411402487.6, filed on September 30, 2024, and entitled "A map detection method, a model training method and device", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the field of intelligent vehicles, in particular to a map detection method, a model training method and device. BACKGROUND

[0003] Road structure cognition is very important for intelligent driving. Existing road structure cognition is usually based on real-time perception network and splicing scheme of fusion map rules to model the surrounding scene. However, the fusion of real-time perception network and map rules has complex rule interaction and heavy processing agent, which affects the overall performance of the system. Moreover, the effect after fusion is very limited by the integrity and confidence of the map. Especially for the map with low precision, the fused map may still have little help for road structure cognition. There may also be actual scene changes, resulting in inconsistency between the map and the actual scene, leading to errors in the final road structure cognition result or failure to achieve road structure cognition.

[0004] Therefore, how to improve the accuracy of road structure cognition has become a problem to be solved. SUMMARY

[0005] The present application provides a map detection method, a model training method and device, which can pre-fuse map features and sensor features, fully extract road network features, and output a map with higher precision.

[0006] Therefore, in a first aspect, the present application provides a map detection method, comprising: first, obtaining sensor data, the sensor data comprising data collected by at least one sensor; then extracting features from the sensor data to obtain perception features; at the same time, obtaining a first map and extracting features from the first map to obtain first map features; then fusing the perception features and the first map features to obtain fused features; and decoding the fused features to obtain a second map, the second map having higher precision than the first map and including multiple types of vectors.

[0007] In the embodiments of the present application, before decoding the map, the map and the perception features are fused, that is, pre-fusion, the map road network features can be fully extracted, so that the accuracy of the map obtained by subsequent decoding is higher. Compared with directly decoding after splicing the map features and the perception features after extracting the features from the map, the splicing method is easy to cause inconsistent features, and the actual scene may not match the map. The method provided in the embodiments of the present application can fully extract the environment features in the actual scene by combining the map features and the perception features, realize more accurate road structure cognition, and output more accurate vector maps.

[0008] In a possible implementation, the foregoing decoding according to the fused features to obtain the second map comprises: inputting the fused features and the first map features into at least one decoder, and obtaining the second map according to the output of the at least one decoder, the at least one decoder being respectively used to output at least one type of vector.

[0009] In the embodiments of the present application, one or more decoders can be set, and one or more types of vectors are output according to the one or more decoders, so as to adapt to scenarios with multiple types of vectors.

[0010] In a possible implementation, the foregoing at least one decoder is respectively used for at least one of instance detection, road network completion or semantic segmentation based on the input features. In the embodiments of the present application, the decoder can be set for different application scenarios, and has stronger generalization.

[0011] In a possible implementation, the foregoing second map comprises vectors representing one or more types of entities: lane boundary line, road, intersection or lane center line. In the embodiments of the present application, the map lane detection scenario can be applied, and the lane or intersection information in the actual scene can be detected to obtain more accurate lane and intersection map information.

[0012] In a possible implementation, the foregoing first map data comprises data of a low-precision map. In the embodiments of the present application, a higher-precision map can be output based on the low-precision map, that is, a higher-precision map is constructed at a lower cost.

[0013] In a possible implementation, the foregoing extracting features from the sensor data to obtain perception features includes: extracting features from the sensor data to obtain bird's eye view (BEV) features, and the perception features include the BEV features; and the foregoing fusing the perception features and the first map features to obtain fused features includes: taking the first map features as input of a perception module, outputting second map features, and the perception module is configured to down-sample the first map features at least once and fuse the down-sampled features and the perception features; and fusing the BEV features and the second map features to obtain the fused features. In the implementation of the present application, the map features can be fused into the BEV space, the dimension of the BEV features can be improved, and thus the fused features with richer information can be obtained.

[0014] In a possible implementation, the foregoing fusing the perception features and the first map features to obtain fused features includes: interacting the perception features and the first map features through an attention mechanism to obtain the fused features. In the implementation of the present application, the perception features and the map features can be interacted based on the attention mechanism, and thus the fused features can be obtained in a sparse manner, or the fused features can be understood as sparse features, which helps to reduce the dimension of data, and thus faster and more efficient model training and inference can be achieved.

[0015] In a possible implementation, the foregoing fusing the perception features and the first map features to obtain fused features includes: interacting the perception features and the first map features through an attention mechanism to obtain the fused features. In the implementation of the present application, the perception features and the map features can be interacted based on the attention mechanism, and thus the fused features can be obtained in a sparse manner, or the fused features can be understood as sparse features, which helps to reduce the dimension of data, and thus faster and more efficient model training and inference can be achieved.

[0016] Therefore, in the implementation of the present application, the map detection model that can output a higher-precision map through map feature pre-fusion can be trained, and the map detection model can sufficiently extract map road network features, so that the subsequent decoding can obtain a map with higher precision.

[0017] Optionally, the structure of the map detection model can refer to the description of the foregoing first aspect or any optional implementation of the first aspect, and will not be described here.

[0018] In a possible implementation, the foregoing training data includes low-precision map data. Therefore, in the implementation of the present application, the model training can be implemented at a lower cost, and the map detection model that outputs a map with higher precision can be obtained.

[0019] In a possible implementation, the acquiring the training data includes: acquiring data of an initial low-precision map; and aligning at least one instance in the initial low-precision map, the type of the at least one instance including the type of an instance in a map output by the map detection model. In the implementation of the present application, after the initial high-precision map is acquired, the alignment operation on the types of instances in the initial low-precision map is performed, so as to realize the specification unification between the road vector instance elements in the low-precision map and the vector instance results output by the map detection model.

[0020] In a possible implementation, the acquiring the training data further includes: in the case of acquiring the data of the initial low-precision map, repairing the initial low-precision map to obtain low-precision map data, the repairing including broken line reconnection, shape point correction, or layer deletion. In the implementation of the present application, the low-precision map is repaired, so as to obtain a more accurate low-precision map.

[0021] In a possible implementation, the acquiring the training data further includes: in the case of acquiring the data of the initial low-precision map, verifying the initial low-precision map to obtain a low-precision map that passes the verification. In the implementation of the present application, the low-precision map is further verified, so as to verify whether the map is accurate through the verification, and thus a more accurate low-precision map is obtained.

[0022] In a third aspect, an embodiment of the present application provides a map detection device, including:

[0023] The first acquiring module is configured to acquire sensor data, the sensor data including data collected by at least one sensor;

[0024] The first feature extraction module is configured to extract features from the sensor data to obtain perception features;

[0025] The second acquiring module is configured to acquire a first map;

[0026] The second feature extraction module is configured to extract features from the first map to obtain first map features;

[0027] The fusion module is configured to fuse the perception features and the first map features to obtain fused features;

[0028] The decoding module is configured to decode the fused features to obtain a second map, the precision of the second map being higher than that of the first map, and the second map including a plurality of types of vectors.

[0029] Effects achieved by the third aspect or any optional implementation of the third aspect can be refered to the description of the first aspect or any optional implementation of the first aspect, which will not be described here.

[0030] In a possible implementation, the decoding module is specifically configured to: input the fused feature and the first map feature into at least one decoder, and obtain the second map according to an output of the at least one decoder, the at least one decoder being respectively configured to output at least one type of vector.

[0031] In a possible implementation, the at least one decoder is respectively configured to perform at least one of instance detection, road network completion, or semantic segmentation based on the input feature.

[0032] In a possible implementation, the second map includes vectors representing one or more of the following types of entities: a lane boundary line, a road, an intersection, or a lane center line.

[0033] In a possible implementation, the first map data includes data of a low-definition map.

[0034] In a possible implementation, the first feature extraction module is specifically configured to extract features from sensor data to obtain bird's eye view (BEV) features, and the perception features include the BEV features.

[0035] The fusion module is specifically configured to: input the first map feature as an input of a perception module, and output a second map feature, the perception module being configured to perform at least one time of down-sampling on the first map feature and fuse the at least one time of down-sampled feature with the perception feature; and fuse the BEV feature and the second map feature to obtain the fused feature.

[0036] In a possible implementation, the fusion module is specifically configured to: interact the perception feature and the first map feature through an attention mechanism to obtain the fused feature.

[0037] In a fourth aspect, the present application provides a model training apparatus, comprising:

[0038] The acquisition module is configured to acquire training data, the training data including map data.

[0039] The training module is configured to train a map detection model by using the training data to obtain a trained map detection model, the map detection model being configured to: extract first map features from a first map, extract perception features from sensor data, fuse the perception features and the first map features to obtain fused features, and process the fused features and the first map features to obtain a second map, the second map being higher in accuracy than the first map, and the second map including vectors of multiple types.

[0040] Effects achieved by the fourth aspect or any optional implementation of the fourth aspect can refer to the description of the second aspect or any optional implementation of the second aspect, which will not be described here.

[0041] In a possible implementation, the training data described above comprises low-precision map data.

[0042] In a possible implementation, the obtaining module described above is specifically configured to: obtain data of an initial low-precision map; and align at least one instance in the initial low-precision map, the type of the at least one instance comprising the type of an instance in a map output by the map detection model.

[0043] In a possible implementation, the obtaining module described above is specifically configured to: in the case of obtaining data of an initial low-precision map, repair the initial low-precision map to obtain the low-precision map data, the repairing comprising line reconnection, point correction, or layer deletion.

[0044] In a possible implementation, the obtaining module described above is specifically configured to: in the case of obtaining data of an initial low-precision map, verify the initial low-precision map to obtain a low-precision map that passes the verification.

[0045] In a fifth aspect, an embodiment of the present application provides a map detection device, comprising a processor and a memory, wherein the processor and the memory are interconnected through a circuit, the processor invokes program codes in the memory to perform functions related to processing in the method shown in any of the first aspect.

[0046] In a sixth aspect, an embodiment of the present application provides a model training device, comprising a processor and a memory, wherein the processor and the memory are interconnected through a circuit, the processor invokes program codes in the memory to perform functions related to processing in the method shown in any of the second aspect.

[0047] In a seventh aspect, an embodiment of the present application provides an intelligent driving vehicle, comprising a processor and a memory, wherein the processor and the memory are interconnected through a circuit, the processor invokes program codes in the memory to perform functions related to processing in the positioning method shown in any of the first aspect.

[0048] In an eighth aspect, an embodiment of the present application provides a digital processing chip or a chip, the chip comprising a processing unit and a communication interface, the processing unit obtaining program instructions through the communication interface, the program instructions being executed by the processing unit, the processing unit being configured to perform functions related to processing in any of the optional implementation manners of the first aspect or the second aspect.

[0049] In a ninth aspect, an embodiment of the present application provides a computer-readable storage medium, comprising instructions, when the instructions are executed on a computer, causing the computer to perform the method in any of the optional implementation manners of the first aspect or the second aspect.

[0050] In a tenth aspect, an embodiment of the present application provides a computer program product containing computer programs / instructions, which, when executed by a processor, causes the processor to execute the method in any of the optional implementation forms of the first aspect or the second aspect. BRIEF DESCRIPTION OF DRAWINGS

[0051] FIG. 1 is a schematic diagram of a vehicle structure according to an embodiment of the present application;

[0052] FIG. 2 is a schematic diagram of an application scenario according to an embodiment of the present application;

[0053] FIG. 3 is a schematic diagram of a model training method according to an embodiment of the present application;

[0054] FIG. 4 is a schematic diagram of another model training method according to an embodiment of the present application;

[0055] FIG. 5 is a schematic diagram of a map detection method according to an embodiment of the present application;

[0056] FIG. 6 is a schematic diagram of another map detection method according to an embodiment of the present application;

[0057] FIG. 7 is a schematic diagram of another map detection method according to an embodiment of the present application;

[0058] FIG. 8 is a schematic diagram of another map detection method according to an embodiment of the present application;

[0059] FIG. 9 is a schematic diagram of another map detection method according to an embodiment of the present application;

[0060] FIG. 10 is a schematic diagram of another map detection method according to an embodiment of the present application;

[0061] FIG. 11 is a schematic diagram of another map detection method according to an embodiment of the present application;

[0062] FIG. 12 is a schematic diagram of a map detection model according to an embodiment of the present application;

[0063] FIG. 13 is a schematic diagram of another map detection model according to an embodiment of the present application;

[0064] FIG. 14 is a schematic diagram of a model training apparatus according to an embodiment of the present application;

[0065] FIG. 15 is a schematic diagram of a map detection apparatus according to an embodiment of the present application;

[0066] FIG. 16 is a schematic diagram of another model training apparatus according to an embodiment of the present application;

[0067] FIG. 17 is a structural schematic diagram of another map detection device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0068] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0069] The method provided by the embodiments of the present application relates to a neural network. For the convenience of understanding, some terms or concepts related to the embodiments of the present application are explained as follows.

[0070] (1) Deep neural network

[0071] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple intermediate layers. According to the position of different layers, the neural network in the DNN can be divided into three categories: an input layer, an intermediate layer, and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the number of intermediate layers is the number of hidden layers.

[0072] Although the DNN looks very complex, each layer can be expressed as a linear relationship expression: where, is an input vector, is an output vector, is an offset vector or bias parameter, w is a weight matrix (also called a coefficient), and α() is an activation function. Each layer is only a simple operation on the input vector to obtain the output vector Due to the large number of layers in the DNN, the number of coefficients W and offset vectors is also relatively large. These parameters in the DNN are defined as follows: taking the coefficient w as an example: assuming that in a three-layer DNN, the linear coefficient of the fourth neuron in the second layer to the second neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, and the subscript corresponds to the third layer index 2 of the output and the second layer index 4 of the input.

[0073] In summary, the coefficient of the kth neuron in the L-1th layer to the jth neuron in the Lth layer is defined as

[0074] It is noted that the input layer is without W parameters. In deep neural networks, more intermediate layers allow the network to better capture the complexity of real-world scenarios. In theory, the more parameters a model has, the higher its complexity, and the greater its "capacity" to perform more complex learning tasks. Training a deep neural network is a process of learning the weight matrices, and the ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (weight matrices formed by vectors W of many layers).

[0075] (2) Convolutional Neural Network

[0076] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. The convolutional neural network includes a feature extractor composed of a convolutional layer and a subsampling layer, which can be regarded as a filter. The convolutional layer refers to the neuron layer in the convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of the convolutional neural network, a neuron can be connected only to part of the adjacent layer neurons. In a convolutional layer, there are usually several feature planes, each of which can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights are convolution kernels. Shared weights can be understood as being independent of the way and position of extracting image information. The convolution kernel can be initialized in the form of a matrix of random size, and the convolution kernel can obtain reasonable weights through learning in the training process of the convolutional neural network. In addition, the direct benefit of shared weights is to reduce the connections between layers of the convolutional neural network, while reducing the risk of overfitting.

[0077] (3) transformer

[0078] The transformer structure is a feature extraction network (similar to the convolutional neural network) including an encoder and a decoder. Of course, in some cases, the transformer structure can not include an encoder, but include a decoder.

[0079] Encoder: learn features in a global receptive field through self-attention, such as the features of a pixel point.

[0080] Decoder: learn the features of the required module through self-attention and cross-attention, such as the features of an output frame.

[0081] Exemplarily, the structure of a Transformer layer can include an attention network and a forward network module. In the case of natural language processing, the attention network obtains corresponding weight values based on the attention mechanism to calculate the correlation between words and words to obtain context-related word representations, which is the core part of the Transformer structure. The forward network further transforms the obtained representations to obtain the final output of the Transformer layer. In addition to the two important components, a residual layer (ADD) and linear normalization (Norm) are also stacked on these two components, respectively, to optimize the output of the Transformer layer.

[0082] (4) PV transformer

[0083] A transformer architecture based on attention point-PV space voxel is provided for converting an image into a perspective view. The goal of the PV transformer is to learn an encoding function from points to PV space voxels end-to-end through an attention module.

[0084] (5) Attention mechanism

[0085] The attention mechanism can quickly extract important features of sparse data. The attention mechanism provides an effective modeling method for capturing global context information through QKV. Assuming that the input is Q (query), the context is stored in the form of key-value pair (K, V). Then the attention mechanism is actually a mapping function from query to a series of key-value pairs. The essence of the attention function can be described as a mapping from a query to a series of (key, value) pairs. Attention essentially assigns a weight coefficient to each element in the sequence, which can also be understood as soft addressing. If each element in the sequence is stored in the form of (K, V), then attention completes the addressing by calculating the similarity between Q and K. The similarity calculated by Q and K reflects the importance of the extracted V value, that is, the weight, and then the weighted sum is obtained to obtain the final feature value.

[0086] The calculation of attention mainly includes three steps. The first step is to calculate the similarity between the query and each key to obtain a weight. Common similarity functions include dot product, concatenation, and perception, etc. The second step is to normalize the weights using a softmax function, which can normalize the weights to obtain a probability distribution with a sum of 1. The last step is to sum the weights and the corresponding key values to obtain the final feature value. The specific calculation formula can be as follows:

[0087] where d is the dimension of the matrix Q, K.

[0088] In addition, attention includes self-attention and cross-attention. Self-attention can be understood as a special attention, that is, the input of QKV is consistent. However, the input of QKV in cross-attention is inconsistent. Attention uses the similarity (for example, inner product) between features as a weight to integrate the queried features as the updated value of the current feature. Self-attention is an attention extracted based on the attention of the feature map itself.

[0089] For convolution, the size of the receptive field is limited by the setting of the convolution kernel, which causes the network to often need to be stacked in multiple layers to focus on the entire feature map. The advantage of self-attention is that its attention is global, and it can obtain the global spatial information of the feature map through simple query and assignment.

[0090] Secondly, the method provided by the embodiment of the present application can be applied to various electronic devices, and can perform pre-fusion based on the sensors and map data set in the electronic device, combine the sensor data to more fully extract road network features, and thus obtain a higher-precision output map.

[0091] In a possible implementation, the method provided by the embodiment of the present application can be deployed in a terminal, that is, the aforementioned cushion device can be a terminal, such as a mobile phone, a tablet personal computer (TPC), a media player, a smart television, a laptop computer (LC), an augmented reality (AR) / virtual reality (VR), a vehicle-mounted terminal, a smart driving vehicle, a personal digital assistant (PDA), a personal computer (PC), a camera, a camcorder, a smart watch, a wearable device (WD), etc., which are not limited by the present application.

[0092] For example, the method provided in the embodiments of the present application can be applied to perception and map fusion of a terminal, such as fusion of sensor data collected by a terminal sensor and a map to obtain a more accurate fused map, where the terminal is, for example, a mobile phone, a wearable device, a vehicle, or a robot.

[0093] In a possible implementation, the method provided in the embodiments of the present application can also be deployed in a server, that is, the electronic device described above can be a server, such as a cloud server or a server connected to a terminal. For example, the method provided in the embodiments of the present application can be deployed in a cloud server connected to a terminal, and the cloud server can receive image or point cloud sensor data from the terminal, fuse map data (which can come from the local cloud server or be sent by the terminal) to obtain a more accurate map after fusion.

[0094] The following exemplary describes a case in which the method provided in the embodiments of the present application is applied to a vehicle.

[0095] For example, high-precision positioning of a vehicle is crucial for intelligent driving functions of the vehicle, and the high-precision positioning of the vehicle is highly dependent on a higher-precision map. For example, intelligent driving functions can generally include functions such as lane departure warning (LDW) or lane keeping assist (LKA) in an advanced driving assistance system (ADAS). For example, lane departure warning and lane keeping assist need to detect lane lines in real time and determine whether the ego vehicle deviates from the current lane. For example, when the vehicle deviates from the lane without turning on the turn signal, a warning signal, a vibrating steering wheel, or even active force to pull back the steering wheel can be used to remind the driver to return to the lane. Through the method provided in the embodiments of the present application, sensor data of an environment in which the vehicle is located and an initial map can be fused, which is equivalent to expanding the map in combination with the sensor data, so as to obtain a more accurate map of the environment in which the vehicle is located, thereby enabling the vehicle to achieve higher-precision positioning and environmental perception. For example, when the vehicle determines that the vehicle deviates from the lane according to the high-precision positioning and environmental perception, the lane keeping assist system can control the steering wheel to turn so that the vehicle returns to the current lane.

[0096] For another example, in a robot application scenario, sensor data collected by a robot and initial map data can be fused to obtain a more accurate fused map containing more information, and on this basis, the robot can achieve accurate perception of the environment in which the robot is located and high-precision positioning of the robot, thereby planning a more suitable travel route for the robot based on the high-precision positioning and the fused map.

[0097] The following takes a vehicle as an example, and the method provided in the embodiments of the present application is introduced in the structure of the device.

[0098] For example, taking the vehicle as an example, the structure of the vehicle can be as shown in FIG. 1.

[0099] Referring to FIG. 1, FIG. 1 is a schematic diagram of a structure of a vehicle provided in the embodiments of the present application. The vehicle 100 can be configured in an intelligent driving mode. For example, the vehicle 100 can control itself while being in the intelligent driving mode, and can determine the current state of the vehicle and its surrounding environment through human operation, determine whether there is an obstacle in the surrounding environment, and control the vehicle 100 based on the information of the obstacle. When the vehicle 100 is in the intelligent driving mode, the vehicle 100 can also be operated without human interaction.

[0100] Referring to FIG. 1, FIG. 1 is a schematic diagram of a structure of a vehicle provided in the embodiments of the present application. FIG. 1 is a functional block diagram of the vehicle 100 provided in the embodiments of the present application. The vehicle 100 can be configured in a full or partial intelligent driving mode. For example, the vehicle 100 can obtain the surrounding environment information of the vehicle 100 through a perception system 120, and obtain an intelligent driving strategy based on the analysis of the surrounding environment information to achieve full intelligent driving, or present the analysis result to a user to achieve partial intelligent driving.

[0101] The vehicle 100 can include various subsystems, such as an infotainment system 110, a perception system 120, a decision control system 130, a drive system 140, and a computing platform 150. Optionally, the vehicle 100 can include more or fewer subsystems, and each subsystem can include multiple components. In addition, each subsystem and component of the vehicle 100 can be interconnected through wired or wireless means.

[0102] In some embodiments, the infotainment system 110 can include a communication system 111, an entertainment system 112, and a navigation system 113.

[0103] The communication system 111 can include a wireless communication system that can wirelessly communicate with one or more devices directly or via a communication network. For example, the wireless communication system can use 3G cellular communication, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communication, such as LTE, or 5G cellular communication. The wireless communication system can communicate with a wireless local area network (WLAN) using WiFi. In some embodiments, the wireless communication system can communicate directly with a device using an infrared link, Bluetooth, or ZigBee. The wireless communication system can include one or more dedicated short range communications (DSRC) devices that can include public and / or private data communication between vehicles and / or roadside stations.

[0104] The entertainment system 112 can include a center screen, a microphone, and a sound system, based on which a user can listen to the radio or play music in the vehicle, or connect a mobile phone with the vehicle and realize mobile phone screen projection on the center screen. The center screen can be touch-controlled, and the user can operate the screen by touch. In some cases, the user's voice signal can be obtained through the microphone, and some control of the vehicle 100 by the user can be realized according to the analysis of the user's voice signal, such as adjusting the temperature in the vehicle. In other cases, music can be played for the user through the sound system.

[0105] The navigation system 113 can include a map service to provide navigation for the vehicle 100 to travel along a route. The navigation system 113 can be used in cooperation with the global positioning system 121 and the inertial measurement unit 122 of the vehicle. The map can be a two-dimensional map, a high-definition map, or a map constructed based on data collected during vehicle travel.

[0106] ​The perception system 120 can include several sensors that sense information about the environment surrounding the vehicle 100. For example, the perception system 120 can include a global positioning system 121 (which can include a global navigation satellite system (GNSS), specifically a GPS system, a Beidou system, or other positioning system), an inertial measurement unit (IMU) 122, a LiDAR 123, a millimeter wave radar 124, an ultrasonic radar 125, and a camera 126. The perception system 120 can also include sensors that monitor internal systems of the vehicle 100 (e.g., an in-vehicle air quality monitor, a fuel gauge, an oil temperature gauge, etc.). Sensor data from one or more of these sensors can be used to detect objects and their respective characteristics (location, shape, direction, velocity, etc.). Such detection and identification are key functions for the safe operation of the vehicle 100. The data collected by the sensors in the vehicle, as mentioned below in this application, can include data collected by each unit in the perception system 120, such as point cloud data collected by the LiDAR or environmental images collected by the camera, etc.

[0107] The global positioning system 121 can be used to determine the geographic location of the vehicle 100.

[0108] The inertial measurement unit 122 is used to sense changes in position and orientation of the vehicle 100 based on inertial acceleration. In some embodiments, the inertial measurement unit 122 can be a combination of an accelerometer and a gyroscope.

[0109] The LiDAR 123 can utilize laser light to sense objects in the environment in which the vehicle 100 is located. In some embodiments, the LiDAR 123 can include one or more laser sources, a laser scanner, and one or more detectors, among other system components.

[0110] The millimeter wave radar 124 can utilize radio signals to sense objects within the surrounding environment of the vehicle 100. In some embodiments, in addition to sensing objects, the millimeter wave radar 124 can also be used to sense the speed and / or direction of travel of the objects.

[0111] The ultrasonic radar 125 can utilize ultrasonic signals to sense objects around the vehicle 100.

[0112] The camera 126 can be used to capture image information of the surrounding environment of the vehicle 100. The camera 126 can include a monocular camera, a binocular camera, a structured light camera, and a panoramic camera, etc., and the image information obtained by the camera 126 can include static image information or video stream information.

[0113] The decision control system 130 includes a computing system 131 that makes analytical decisions based on information acquired by the perception system 120. The decision control system 130 also includes a vehicle controller 132 that controls the power system of the vehicle 100, and a steering system 133, a throttle 134, and a braking system 135 that control the vehicle 100.

[0114] The computing system 131 can process and analyze various information acquired by the perception system 120 to identify targets, objects, and / or features in the environment surrounding the vehicle 100. The targets can include pedestrians or animals, and the objects and / or features can include traffic signals, road boundaries, and obstacles. The computing system 131 can use object recognition algorithms, Structure from Motion (SFM) algorithms, video tracking, and other techniques. In some embodiments, the computing system 131 can be used to map the environment, track objects, estimate the speed of objects, and so on. The computing system 131 can analyze the acquired various information and derive a control strategy for the vehicle.

[0115] The vehicle controller 132 can be used to coordinate the control of the power battery and the engine 141 of the vehicle to improve the power performance of the vehicle 100.

[0116] The steering system 133 can be used to adjust the forward direction of the vehicle 100. For example, in one embodiment, the steering system 133 can be a steering wheel system.

[0117] The throttle 134 is used to control the operating speed of the engine 141 and, in turn, the speed of the vehicle 100.

[0118] The braking system 135 is used to control the deceleration of the vehicle 100. The braking system 135 can use friction to slow down the rotation speed of the wheels 144. In some embodiments, the braking system 135 can convert the kinetic energy of the wheels 144 into electric current. The braking system 135 can also take other forms to slow down the rotation speed of the wheels 144 to control the speed of the vehicle 100.

[0119] The drive system 140 includes components that provide power motion for the vehicle 100. In one embodiment, the drive system 140 can include an engine 141, an energy source 142, a transmission system 143, and wheels 144. The engine 141 can be an internal combustion engine, an electric motor, an air compression engine, or other types of engine combinations, such as a hybrid engine composed of a gasoline engine and an electric motor, a hybrid engine composed of an internal combustion engine and an air compression engine. The engine 141 converts the energy source 142 into mechanical energy.

[0120] Examples of energy sources 142 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electrical power. Energy sources 142 can also provide energy for other systems of vehicle 100.

[0121] Transmission system 143 can transmit mechanical power from engine 141 to wheels 144. Transmission system 143 can include a gearbox, a differential, and drive shafts. In one embodiment, transmission system 143 can also include other devices, such as a clutch. Among other things, drive shafts can include one or more shafts that can be coupled to one or more wheels 144.

[0122] Parts or all of the functionality of vehicle 100 is controlled by computing platform 150. Computing platform 150 can include at least one processor 151 that can execute instructions 153 stored in a non-transitory computer readable medium such as memory 152. In some embodiments, computing platform 150 can also be multiple computing devices that control individual components or subsystems of vehicle 100 in a distributed manner.

[0123] Processor 151 can be any conventional processor, such as a commercially available CPU. Alternatively, processor 151 can also include a graphics processing unit (GPU), a field programmable gate array (FPGA), a system on chip (SOC), an application-specific integrated circuit (ASIC), or a combination thereof. Processor 151 can be located on a device that is remote from the vehicle and in wireless communication with the vehicle.

[0124] In some embodiments, memory 152 can contain instructions 153 (e.g., program logic) that can be executed by processor 151 to perform various functions of vehicle 100. Memory 152 can also contain additional instructions, including instructions to send data to, receive data from, interact with, and / or control one or more of infotainment system 110, perception system 120, decision control system 130, and drive system 140.

[0125] The method provided by the embodiments of the present application can be executed by the computing platform 150, for example, the processor 151 can read the program stored in the memory 152, fuse the data collected by the perception system 120 with the initial map, output a higher-precision map, and determine the driving decision of the vehicle based on the higher-precision map, and send an instruction to the decision control system 130 to control the vehicle to realize the intelligent driving function through the decision control system 130.

[0126] Of course, the method provided by the embodiments of the present application can also be directly deployed in the perception system 120 to output the high-precision positioning of the vehicle, or can also be deployed in the decision control system 130, the decision control system 130 fuses the data from the perception system with the initial map to obtain a higher-precision map, realizes the perception of the vehicle environment and the positioning of the vehicle based on the higher-precision map, and controls the vehicle to realize the intelligent driving function of the vehicle based on the perceived target.

[0127] In addition to the instructions 153, the memory 152 can also store data, for example, road maps, route information, the position, direction, speed of the vehicle and other similar vehicle data, and other information, and the higher-precision map obtained by the method provided by the embodiments of the present application can also be stored in the memory 152. Such information can be used by the vehicle 100 and the computing platform 150 during the operation of the vehicle 100 in the autonomous, semi-autonomous and / or manual mode.

[0128] The computing platform 150 can control the functions of the vehicle 100 based on the input received from various subsystems (for example, the driving system 140, the perception system 120 and the decision control system 130). For example, the computing platform 150 can use the input from the decision control system 130 to control the steering system 133 to avoid the obstacles detected by the perception system 120. In some embodiments, the computing platform 150 can be operated to provide control over many aspects of the vehicle 100 and its subsystems.

[0129] Optionally, one or more of the above components can be installed separately from or associated with the vehicle 100. For example, the memory 152 can exist partially or completely separately from the vehicle 100. The above components can be communicatively coupled together in a wired and / or wireless manner.

[0130] Optionally, the above components are only an example, and in actual applications, components in each module can be added or deleted according to actual needs, and FIG. 1 should not be understood as a limitation on the embodiments of the present application.

[0131] The vehicle 100 can be a car, a truck, a motorcycle, a bus, a ship, an airplane, a helicopter, an entertainment vehicle, an amusement park vehicle, construction equipment, a trolley, a golf cart, a train, or the like, which can implement intelligent driving, or a vehicle-mounted terminal, and the like, and embodiments of the present application are not particularly limited.

[0132] The environment perception technology of the intelligent driving system has undergone multiple iterations, and the process of environment perception can be divided into multiple stages. For example, as shown in FIG. 2, it can be divided into “heavy map, light perception”, “light map, heavy perception” and “look at the map, sense perception”, that is, the sensor data of the environment is becoming more and more important in the intelligent driving function of the vehicle. From the intelligent driving vehicle relying only on high-precision maps for driving, to reducing the dependence on high-precision maps and improving real-time perception, with the development of vehicle intelligent driving functions, the map has gradually become auxiliary data for vehicle intelligent driving.

[0133] For intelligent driving vehicles, the road structure recognition system plays a crucial role in the intelligent driving system. The system receives static road element information and traffic light information, and then generates road topology relationships to guide the ego vehicle through various complex road scenes. The road structure recognition system is based on static road elements in the surrounding real-time perception environment, and correct road information helps to generate an accurate road topology network. Taking lane lines and intersections as examples, the road structure recognition system receives original lane lines and intersection information. Each intersection connects multiple roads, which contain multiple lanes, so there are a large number of alternative lane switching routes.

[0134] For example, as vehicle intelligent driving solutions gradually tend to be non-mapping solutions, such as an advanced driving system (ADS), which is an intelligent driving system and of course can have different names in different devices, embodiments of the present application take ADS as an example. In existing road structure cognition implementation solutions, such as the road structure cognition solution in the updated version ADS2.0 of ADS, the road structure cognition solution is a splicing solution based on the output results of a real-time perception network and priori map based on rule logic. Based on this solution, the surrounding scene is modeled, i.e., a map post-fusion solution. The rule splicing map solution based on post-fusion improves the granularity of perception modeling, but is also limited by the post-fusion rule itself. For example, the rule logic of the post-fusion solution is complex, and the post-processing code is heavy, which affects the overall performance of the system. Secondly, the post-fusion effect is completely limited by the completeness and confidence of the map. In existing solutions, the map elements and perception detection elements are usually directly spliced, which can easily lead to inconsistencies. In addition, there can be cases where it cannot resist real changes, such as if the actual scene changes, resulting in a map that does not match the actual scene, making the final fused map inaccurate. Furthermore, if the initial map source is high-precision or standard-precision, the topology and geometric information is not effective, and the features extracted by GNN / CNN / Transformer are simply superimposed on the BEV features, or are directly decoded to predict the segmentation result, which does not greatly improve the road structure elements under full instantiation.

[0135] For example, an existing solution provides a network implemented based on a GNN architecture, called VectorNet, which is a hierarchical graph neural network GNN. The input is a high-precision map, which is implemented as a GNN framework. Road elements such as lane lines, intersections, boundary lines, and zebra crossings are represented by vectors, which are further modeled and encoded into GNN to realize the interaction between road elements, supplemented by high-precision true value supervision to complete the training loop, and output trajectory line detection results. However, the extraction of map road element features is limited to a single modal source, which is an internal modal fusion, and is not fused with multi-source sensors, and the effect is limited. It is limited by the GNN neural network, and the calculation time and deployability are not good. In addition, this solution is mainly used to verify the improvement of vehicle trajectory line accuracy, although it fuses map elements, but it is not verified whether it is helpful for real-time online high-precision map detection.

[0136] For example, a P-MapNet is provided in the prior art, which is an online high-definition map generation model that improves model detection performance by combining the prior information provided by a standard precise map (SD) and a high-definition map (HD), and uses a masked autoencoder (MAE) to capture the prior distribution of the high-definition map to refine the segmentation result. However, this solution is limited to the segmentation task, and the feature extraction and fusion of the SD and the HD are based on the BEV segmentation form, which may cause accuracy loss in the process of modeling and rendering road elements; the MAE is a kind of "beautifying and filling in" module for images, and is also a segmentation task for online high-definition map detection, so the output map contains non-vectors, which may not be convenient for downstream task applications; the input source relies too much and must have both SD and HD inputs, which may introduce noise of the SD and the HD.

[0137] For example, an end-to-end framework of SMERF is provided in the prior art, which inputs Road element information in the SD map, interacts the prior information features in the SD map with the BEV features through a transformer framework, and improves the detection ability of Lane. However, the input source is a single road element Road in the SD map, which contains limited prior information content and is of limited help to real-time high-definition map detection tasks; this solution only analyzes the Lane-Topology subtask, and the real-time high-definition map detection task requires the recognition of a large number of subtasks such as intersections, crosswalks, and boundary lines in addition to Lane detection. If directly applied to the real-time high-definition map generation task, the accuracy and efficiency are difficult to guarantee.

[0138] Most of the problems caused by road structure recognition are that the road element detection is not robust and long-term, so better use of map elements to improve perception performance and range is crucial to the intelligent driving function of vehicles.

[0139] Therefore, the embodiment of the present application provides a high-definition map detection scheme based on a pre-fusion map, which realizes accurate extraction of map road network information by fusing pre-fusion perception features and map features, and obtains a map with higher accuracy.

[0140] The flow of the method provided by the embodiment of the present application is introduced below.

[0141] The method provided in the embodiments of the present application can be divided into a training stage and an inference stage. In the training stage, the map detection model provided in the embodiments of the present application is trained, and in the inference stage, the map detection model provided in the embodiments of the present application is used to perform a downstream task. The training stage and the inference stage can be deployed in the same device or in different devices. For example, the training stage can be deployed in the cloud or a server, and the inference stage can be deployed in a smart terminal, such as a mobile phone, a tablet computer, a smart vehicle, a robot or a wearable device.

[0142] The training stage and the inference stage in the method provided in the embodiments of the present application will be introduced respectively below.

[0143] I. Training stage

[0144] Referring to FIG. 3, a flowchart of a model training method provided in the embodiments of the present application is as follows.

[0145] 301. Obtain training data, wherein the training data includes low-precision map data.

[0146] Specifically, in the training stage, a large amount of low-precision map data can be collected, and the collected low-precision map can be preprocessed to serve as training data. Generally, the collected low-precision map can have noise, mispositioned layers or missing parts of content, etc. In the preprocessing stage, the collected low-precision map can be filtered, corrected, verified or aligned, etc. to obtain a more accurate low-precision map.

[0147] In a possible implementation, the initial low-precision map can be verified to obtain a low-precision map that passes the verification. For example, a large amount of low-precision map data sources can be obtained, and these map data sources can have a large amount of noise, defects or inconsistent content due to low cost, low specification and low precision. In the embodiments of the present application, the initial low-precision map data can be verified to identify whether there are mispositioned layers, map changes, lateral displacement of lost lines, etc., and corresponding labels tag can be generated.

[0148] In a possible implementation, the initial low-precision map can be repaired to reduce part of the noise in the map. The repair operation can specifically include but is not limited to broken line reconnection, shape point correction or layer deletion, etc. The broken line reconnection is to reconnect the lane elements that should be connected but are disconnected in the low-precision map, the shape point correction is to delete unreasonable shape points or adjust the positions of unreasonable shape points, etc., and the layer deletion is to delete mispositioned layers or delete redundant layers, etc. Therefore, in the embodiments of the present application, a more accurate map can be obtained by repairing the low-precision map to a certain extent.

[0149] In a possible implementation, the alignment operation can be performed on at least one road element instance in the low-precision map to achieve specification unification of the road vector instance element in the low-precision map and the vector instance result output by the map detection model. For example, the elements of each road structure cognition in the low-precision map can be classified, and the map elements can include lane, road, intersection, and the like, which can be summarized as a collection of points, lines, and surfaces. For example, the lane line element in the low-precision map is a line segment composed of a series of points, and the intersection is a surface composed of lines. Because the training data and the input of the model can be inconsistent in specification, such as different dimensions or formats, in order to successfully input the map data into the model, the training specification needs to be further aligned. The map data to be sent into the training pipeline is classified according to the road structure cognition element, so as to align the training data and the data specification required by the training model, including but not limited to Lane specification alignment, Road specification alignment, Intersection specification alignment, and the like. Therefore, in the embodiments of the present application, the input required by the model training can be obtained through the specification alignment, so that the data distribution characteristics of the map input can be learned during the model training, thereby more accurately outputting the map detection result.

[0150] Exemplarily, the map preprocessing process can be as shown in FIG. 4. After the low-precision map is collected, first, various verification problems in the data verification pool are verified, such as whether there is a map layer error, map replacement, invalid data, inconsistency between the map and the vehicle trajectory, or positioning failure in the map. Then, the data that fails the verification is repaired, such as broken line reconnection, unreasonable point deletion, or layer error deletion. Then, it is judged whether the repaired data is available. The available data can be preprocessed in the next step. For the unusable data, data cleaning or backflow can be performed, and it is judged whether there is meaningless data, such as data irrelevant to the input required by the model or repeated data, and the meaningless data is discarded, and the meaningful data continues to be repaired. Then, the available data continues to be aligned in specification. This step classifies the road structure cognition elements in the map data to be sent into the training pipeline, and sequentially aligns the training data, including Lane specification alignment, Road specification alignment, Intersection specification alignment, and the like, and adds corresponding labels to each instance. The map data output as the true value is used as the training data. Therefore, in the embodiments of the present application, the low-precision map can be preprocessed by map verification, data repair, or specification alignment, so as to obtain map data similar or consistent with the true value specification, and the high-precision training of the model can be completed without high-precision map.

[0151] 302、training the map detection model using the training data to obtain a trained map detection model.

[0152] The map detection model can be used to: extract first map features from the first map, extract perception features from the sensor data, fuse the perception features and the first map features to obtain fused features, and process the fused features and the first map features to obtain a second map, the second map having higher accuracy than the first map, and the second map including multiple types of vectors.

[0153] During the training process, supervised training can be used, or unsupervised training or contrastive learning can be used, so that the map detection model has the ability to output higher-precision maps.

[0154] Further, the flow of tasks that can be performed by the map detection model can be as described in the inference stage, which will not be described again here.

[0155] Therefore, in the embodiments of the present application, the low-precision map can be preprocessed to obtain more accurate and usable map data based on the low-precision map, so as to implement high-precision training of the map detection model based on the low-precision map data, and enable the model to have the ability to input low-precision maps and output higher-precision maps.

[0156] II. Inference stage

[0157] Referring to FIG. 5, a flowchart of a map detection method provided by an embodiment of the present application is shown as follows.

[0158] 501. Obtain sensor data

[0159] The sensor data can include data collected by one or more sensors in a scene, and can specifically include data collected by image sensors (or cameras), infrared sensors, lidar, millimeter wave radar, and the like. Accordingly, the sensor data can specifically include images, point cloud data, infrared data, and the like.

[0160] In one possible implementation, based on the vehicle architecture shown in FIG. 1, the sensor data can specifically include point cloud or image data collected by lidar, cameras, infrared sensors, and the like in the vehicle. For example, during vehicle travel, sensors provided in the vehicle can collect data of the environment in which the vehicle is located.

[0161] 502. Extract features from the sensor data to obtain perception features.

[0162] The pre-trained feature extraction network can be used to extract features from the sensor data, and the extracted features are referred to as perception features. The feature extraction network can specifically include, but is not limited to, CNN, DNN, or other networks that can be used to extract features.

[0163] Generally, for different types of sensor data, a corresponding feature extraction network can be set. For example, for an image, an image feature extraction network can be set, for laser point cloud data, a point cloud feature extraction network can be set, etc. The type of feature extraction network set can be determined according to the actual application scenario, and the present application does not make any limitation.

[0164] Optionally, when extracting features from sensor data, BEV features, i.e., overhead features in bird's eye view space, can be extracted. Thus, the perception accuracy of environmental perception is improved.

[0165] 503, obtain a first map.

[0166] The first map can be a low-precision map or a high-precision map, etc. In some scenarios, it can also include a high-precision map, etc.

[0167] For example, in the case of obtaining a low-precision map, a higher-precision map can be obtained by using the method provided in the embodiments of the present application, so that the electronic device can implement intelligent functions based on the higher-precision map.

[0168] For another example, in the case of obtaining an incomplete high-precision map, the high-precision map can be expanded by using the method provided in the embodiments of the present application, so as to obtain a high-precision map containing more complete information.

[0169] It should be noted that the present application does not make any limitation on the execution order of steps 501 and 503. Step 501 can be executed first, or step 503 can be executed first. The specific execution order can be determined according to the actual application scenario.

[0170] 504, extract features from the first map to obtain first map features.

[0171] After obtaining the first map, features can be extracted from the first map. A feature extraction network for extracting map features can be set to extract the required features from the map.

[0172] For example, generally, a map can include various vector elements, such as vectors in various lanes or intersections, etc. A feature extraction network for extracting map vectors can be set to extract features from the map, and the extracted features are referred to as first map features.

[0173] 505, fuse the perception features and the first map features to obtain fused features.

[0174] The perception features and the first map features can be fused, and the fused features are referred to as fused features.

[0175] The specific fusion manner can include splicing, weighted fusion, same space projection, or fusion through a neural network, and the like. The specific fusion manner can be selected according to an actual application scenario.

[0176] For example, in a possible implementation, the first map feature can be taken as an input of a perception module, and a second map feature is output, the perception module is configured to down-sample the first map feature at least once and fuse the down-sampled feature with a perception feature; and the BEV feature and the second map feature are fused in a bird's eye view (BEV) space to obtain a fused feature. It can be understood that, in the manner provided in the embodiments of the present application, the number of channels of the map feature can be increased, so that the map features with more channel numbers are fused, the dense feature fusion is realized, and the prior information and the features in the map can be effectively extracted and fused.

[0177] In a possible implementation, the perception feature and the first map feature can also be interacted through an attention mechanism to obtain the fused feature. It can be understood that, the perception feature and the first map feature can be taken as keys and values (k, v) in the input of an attention module, so as to realize the exchange between the perception feature and the first map feature through the attention mechanism, and realize the sparse feature fusion.

[0178] 506. Decoding the fused feature and the first map feature to obtain a second map.

[0179] After the fused feature is obtained, the pre-trained neural network can be used to process the fused feature, and a second map is output, the accuracy of the second map is higher than that of the first map, and the second map includes a plurality of types of vectors obtained by analyzing the fused feature.

[0180] The fused feature and the first map feature can be decoded to obtain a second map with higher accuracy and richer types of vectors.

[0181] Therefore, in the embodiments of the present application, the perception feature used in decoding can contain richer content by fusing the map feature and the perception feature before decoding, so as to improve the accuracy of the output result of the decoder and obtain a second map with higher accuracy. It can be understood that, through the method provided in the embodiments of the present application, the low-precision map feature is extracted using a neural network, and on the basis of the fusion of the multi-source sensor input feature, the online high-precision map detection is realized through the high-precision value supervision training manner, and the real-time high-precision map obtained is a full instantiation result.

[0182] Specifically, the fused features can be input into the decoder to obtain a second map containing more types of vectors. Optionally, the number of decoders can be at least one, i.e., one or more decoders, and at least one decoder is used for at least one of instance detection, road network completion, or semantic segmentation based on the input features.

[0183] In the case of fusing the first map features and the perception features, the perception features can fuse the information contained in the map, such as lane features, lane positions, etc., so that the resulting fused features contain richer features based on the fusion of the perception features and the first map features.

[0184] For example, for different types of required data, corresponding decoding heads can be set in the decoder, such as lane boundaries, roads, intersections, or lane centerlines, etc. respectively, and then the instantiated information such as lane boundaries, roads, intersections, or lane centerlines, etc. can be decoded from the input encoded data.

[0185] In addition, optionally, a neural network can be used to perform the method provided by the embodiments of the present application, such as a map detection network provided by the embodiments of the present application, which can include the aforementioned feature extraction network and decoder. The map detection network can be pre-trained using high-precision data to improve the accuracy of the output results of the map detection network.

[0186] The foregoing describes the method provided by the embodiments of the present application, and the method provided by the embodiments of the present application will be described in detail in combination with a specific application scenario.

[0187] Referring to FIG. 6, a flowchart of another map detection method provided by the embodiments of the present application is shown.

[0188] The input data includes multiple types, which can be divided into multi-source sensor data and low-precision maps. Multi-source sensor data refers to data collected by multiple types or multiple sensors. Of course, the low-precision map here can be replaced by other precision maps or other existing vector maps, etc.

[0189] Subsequently, the multi-source sensor data and the low-precision map are input into the map detection network. The map detection network can be pre-trained using high-precision data to enable the map detection network to output real-time road vectorization mapping results, i.e., the aforementioned second map. That is, by using high-cost high-precision training, the map detection network can output high-precision, high-specification road element vectorization mapping results based on multi-source sensor data and low-cost, low-specification, low-precision maps.

[0190] For example, the structure of the map detection network provided by the embodiments of the present application is shown in FIG. 7, which corresponds to providing a neural network-based map road network information feature fusion framework, which can cover instantiation detection, road network completion, semantic segmentation and other tasks. The vectorized map road elements are modeled as query or bev_feature level features through a neural network. For example, based on the ADS2.0 framework, query or bev_feature level features are integrated to obtain information outside the field of view of the sensor blind area, thereby comprehensively improving the detection capability of the network.

[0191] The following describes each part based on the structure of the map detection network shown in FIG. 7.

[0192] 1. Input data

[0193] The input data can include map data and one or more types of sensor data.

[0194] For example, the map data shown in FIG. 7 can be low-definition map, standard definition map (SD) map, high-definition map, or other crowdsourcing map data, etc.

[0195] The sensor data in FIG. 7 takes the image captured by a camera and the point cloud data collected by a laser radar as an example.

[0196] The image is collected by an image sensor arranged around the vehicle body, such as a long-short focal length camera arranged in front, fisheye cameras arranged on the left and right sides, a rearview camera arranged at the rear, etc. Through these cameras, image data can be captured in real time around the vehicle body in 360 degrees without dead angle and full coverage. The image data can also include depth images, grayscale images, infrared images, etc.

[0197] The laser point cloud can be obtained by a laser radar sensor fixed on the surface of the vehicle body. The laser beam scans the three-dimensional space to obtain the point coordinates, reflectivity and other information of the objects in the three-dimensional space, which can be used to represent the surrounding environment information.

[0198] 2. Feature extraction

[0199] For each type of input data, a corresponding feature extraction network can be set.

[0200] Specifically, in order to more fully extract the features in the input data, different scale feature extraction can be achieved by downsampling or upsampling the input data. For example, as shown in FIG. 8, the input data, such as the input map data, can be downsampled one or more times to obtain one or more downsampled data, and features are extracted from the data of different scales, and the extracted features are upsampled and fused to achieve different feature extraction. Extracting features from different granularities can smooth the noise in the data, and the output features contain more accurate semantics.

[0201] For example, as shown in FIG. 7, for image type input data, an image feature extraction network, i.e. camera backbone, can be set up to extract image features from the input image. And the visual features are converted to BEV space through the PV2BEV network, and the visual BEV features are output. For example, the execution effect of PV2BEV is shown in FIG. 9, and more dense BEV features can be obtained.

[0202] For example, for point cloud type input data, a point cloud feature extraction network, i.e. lidar backbone as shown in FIG. 7, can be set up to output the extracted point cloud features. Optionally, a Dense2Sparse network, i.e. a sparse feature extraction network, can also be set up to output sparse features.

[0203] In addition, if data is collected by multiple sensors, separate feature extraction networks can be set up for each sensor to extract data features from the data collected by each sensor. In order to facilitate processing, multiple features can be fused, such as weighted fusion or splicing, and the fused features are used as perception features.

[0204] For example, for input map data, a map feature extraction network, i.e. map backbone as shown in FIG. 7, can be set up to extract map features from the map, such as vector features of lanes or intersections in the map data, to obtain map features. For example, for map data, multi-scale feature extraction and fusion can be performed to output map features that better represent the characteristics of the map, to improve the accuracy of subsequent encoding and decoding results.

[0205] In addition, for feature encoding of map data, a map encoder can also be used. For example, an encoder for extracting map features can be set up to encode the map data and output map encoding data, i.e. first map features.

[0206] 3. Feature fusion

[0207] After the perception feature and the first map feature are extracted, the perception feature and the first map feature can be fused to output a fused feature. For example, a hybrid map encoder as shown in FIG. 7 can be used to encode the input feature based on an attention mechanism (including a self-attention mechanism and / or a cross-attention mechanism, etc.), map the input feature to an embedding representation in a query manner, so as to facilitate subsequent decoding.

[0208] Specifically, the feature fusion manner can include various manners, such as dense feature fusion or sparse feature fusion, etc. Different feature fusion manners will be introduced below.

[0209] (1) Dense feature fusion

[0210] For example, as shown in FIG. 10, the map data is first processed by a multi-layer perception (MLP). Specifically, the input map feature is first dimensioned, such as dimensioned from <1, 352, 320> to <32, 352, 320>, where <352, 320> represents the feature map resolution, and then the multi-scale feature extraction module is used to down-sample the dimensioned feature map multiple times, such as 3 times. After each down-sampling, the corresponding feature is extracted, which can reduce the resolution of the output feature and increase the number of feature channels of the feature map. The output feature map can be represented as <64, 352, 320>, <128, 352, 320>, and <256, 88, 80>, respectively. After extracting features of different scales, the feature extracted from the last down-sampling is up-sampled and added to the feature map of the same resolution in the down-sampling process. For example, the feature map of <256, 88, 80> is up-sampled and added to the feature map of <128, 352, 320> of the last down-sampling size. The fused feature is up-sampled again. The above operation is repeated to fuse the up-sampled feature map of <64, 352, 320> with the feature map of the same size in the down-sampling process. The output feature dimension is <64, 352, 320>, and the final output map feature has a size of <64, 352, 320>. In the embodiments of the present application, the number of feature channels of the map input is doubled and the size is halved by each up-sampling operation. The purpose of this is to abstract the map features in the BEV segmentation map into high-dimensional features, which can guarantee the interaction between different parts of the map while having high-dimensional semantic features. After up-sampling, the up-sampled feature is added to the map feature of the same size and channel number as the previous down-sampled feature. This can not only retain rich global map semantic information, but also have fine local map semantics.

[0211] For image and point cloud data, features are also extracted respectively, such as extracting features from images using a camera backbone, represented as image features img feat and mask features img feat mask, and extracting features from point cloud data using a voxelizer, represented as lidar feat.

[0212] The map features after multi-scale feature extraction have more rich geometric semantic information, and then the multi-source sensor features, i.e., perception features and map features, can be concatenated in the feature channel dimension using a feature concatenation module. For example, image features (img feat), image feature masks (img feat mask), laser point cloud features (lidar feat), and map features (map feat) <64, 352, 320> are concatenated in the feature channel dimension, and the concatenated features can be represented as bev fusion feature. Then, the output features are adjusted to the required feature size for the subsequent network through a reduce operation.

[0213] In the embodiments of the present application, dense feature fusion is used to obtain dense features, which can contain more rich semantic information, so that the subsequent downstream task can perform decoding based on more rich semantics, thereby obtaining more rich information.

[0214] (2) Sparse feature fusion

[0215] For example, the steps of sparse feature fusion can be as shown in FIG. 11.

[0216] Generally, the vectors in the map not only have vector-shaped points, but also have element types, topological information, and other contents. In the embodiments of the present application, a queue container can be deployed to carry the information of instances in the map. For example, the information of lane lines, the information of roads, or the information of crossroads, etc.

[0217] For example, taking the detection of lane lines as an example, a queue container is deployed to carry the information of lane lines included in the map, with a size of (q, n, c), where q represents the maximum allowed number of lane lines in a single frame of map; n represents the maximum number of shape points contained in each lane line; and c represents the type attribute possessed by each shape point, in addition to the basic x, y, z dimensional geometric coordinates, can also include the direction of the shape point and the lane type of the point, etc. Then input the container into the map encoder module for encoding, and output the map features. Usually, in order to ensure the topological continuity between map elements, in the map encoder, a multi-layer perception network can be used to increase the dimension of the vector map features, and then a self-attention mechanism network (self-attention) is used to make attention interaction between each vector element in the map, and a residual connection and a normalization network layer are also used to help further extract map features.

[0218] Then the multi-head attention mechanism network is used to send the map features as keys and values into the decoding layer, that is, the lane decoder head as shown in FIG. 11.

[0219] At the same time, the image features extracted from the image and the point cloud features extracted from the point cloud data are fused to output the perception features. The perception features are also input into the lane decoder head as (k, v). The query of the lane decoder head is sequentially interacted with the multi-source sensor features and the map features, so as to realize the sparse fusion between the features.

[0220] Further, the structure of the lane decoder head can be as shown in FIG. 12, taking the instance query (i.e., the decoding query representation of the instantiated element) as the input of the lane decoder head. In the lane decoder head, a self-attention layer (self attn), an ADD&Norm layer, a cross-attention layer (cross attn), and an FFN layer can be set, and multiple cross-attention layers are sequentially arranged. The map features and the BEV features are input into the cross-attention layers as (k, v) respectively, so as to realize the interaction with the instance query.

[0221] Generally, sparse features are effective in the case that the feature space is large and most of the features are irrelevant or redundant, therefore, in the embodiments of the present application, the fusion of sparse features helps to reduce the dimensionality of the data, thereby realizing faster and more efficient model training and inference. In addition, in order to improve the accuracy of the representation of visual features, when encoding the image, the visual features can be encoded based on the attention point-voxel mechanism under the PV transformer architecture, thereby improving the accuracy of the output results when decoding based on the visual features, and improving the perception ability of the device to the environment.

[0222] 4. Decoding

[0223] After encoding by the encoder, the features output by the encoder can be decoded by the decoder. Reference is made to the foregoing description of the decoder.

[0224] Generally, in order to improve the richness of the map features, multiple decoders can be provided, which can have different decoding functions. For example, as shown in FIG. 7, multiple decoders are provided for decoding attributes such as lanes, roads, or intersections, respectively, and outputting lanes (such as lane boundary lines or lane center lines, etc.) in the scene, roads in the scene, and intersection conditions in the scene, etc.

[0225] For example, taking the high-definition map detection task as an example, the high-definition map detection task can be decomposed into multiple sub-tasks, as shown in FIG. 13, the process of high-definition supervised training of lane center line sub-task, the BEV features include multi-source sensor features and map features, which are input into the transformer decoder as (k, v), and the centerline query completes information interaction through the Transformer Decoder, and outputs the detection result of the instantiated lane center line. For example, a centerline detection head, an intersection detection head, etc. can be provided. For example, in the centerline detection head, the existence of a vector-shaped point (Point exist), the position regression of a vector-shaped point (Point reg), or the lane type of a vector-shaped point (Point type), etc. can be detected.

[0226] During training, the supervision signal of the model is the high-definition map ground truth, and the loss function obtains network gradients by comparing the difference between the centerline prediction result and the ground truth, thereby realizing forward iteration of the model. The detection tasks of other road elements are similar to the detection process of the lane center line. Therefore, during inference, the model can output a higher-precision lane center line detection result based on the input features.

[0227] For example, specifically, based on the foregoing process, after decoding, the information of the lane, road or intersection in the scene can be output to expand the map. For example, taking the lane as an example, the following one or more can be output:

[0228] centerline coordinates: a point sequence representing the centerline position.

[0229] centerline confidence: indicating the probability of the existence of the centerline. The value range is 0-1.

[0230] centerline type: indicating the lane type, such as ordinary lane, emergency lane, etc.

[0231] lanebound coordinates: lane centerline left and right boundary point sequence.

[0232] lanebound type: lane centerline left and right boundary line type, such as solid line, dashed line, etc.

[0233] Therefore, in the embodiment of the application, the perception features and the map features can be fused before the features are encoded by the encoder, and the fused features are encoded, so that the information contained in the map can be used to supplement the blind area of the sensor, or correct the incorrect perception information, and the map can be expanded by the perception information, so that a higher-precision expanded map is obtained, and the perception ability of the network is improved.

[0234] Especially in the large curvature scene, there may be occlusion. If only the sensor data collected by the sensor of the vehicle is used to perceive the environment, it may lead to the situation that the occluded area cannot be perceived. Especially for the large curvature lane scene, the reduction of the perception ability may reduce the driving safety of the vehicle. However, by using the method provided in the embodiment of the application, the lane line detection ability of the occluded area can be improved by combining the lane information included in the map, so as to improve the map expansion ability of the perception blind area of the vehicle.

[0235] Or, in some scenes, there may be a situation that the map is missing or the map does not match the actual scene road structure. However, by using the method provided in the embodiment of the application, the missing part of the map can be supplemented, or the situation that the map does not match the actual scene road structure can be corrected, so as to obtain a more accurate expanded map, and truly achieve "looking at the map and believing in the perception".

[0236] The foregoing method provided by the application is introduced, and the structure of the device for executing the foregoing method is introduced.

[0237] Referring to FIG. 14, a structural schematic diagram of a model training device is provided in an embodiment of the present application, and the model training device can be used to execute the process of the training stage. The model training device comprises:

[0238] The acquisition module 1401 is configured to acquire training data, and the training data comprises map data.

[0239] The training module 1402 is configured to train the map detection model by using the training data, to obtain a trained map detection model. The map detection model is configured to: extract first map features from a first map, extract perception features from sensor data, fuse the perception features and the first map features to obtain fused features, and process the fused features and the first map features to obtain a second map. The second map has a higher accuracy than the first map, and the second map comprises a plurality of types of vectors.

[0240] In a possible implementation, the training data comprises low-precision map data.

[0241] In a possible implementation, the acquisition module 1401 is specifically configured to: acquire data of an initial low-precision map; and align at least one instance in the initial low-precision map. The types of the at least one instance comprise types of instances in a map output by the map detection model.

[0242] In a possible implementation, the acquisition module 1401 is specifically configured to: in a case where the data of the initial low-precision map is acquired, repair the initial low-precision map to obtain low-precision map data. The repairing comprises line reconnection, point correction or layer deletion.

[0243] In a possible implementation, the acquisition module 1401 is specifically configured to: in a case where the data of the initial low-precision map is acquired, verify the initial low-precision map to obtain a low-precision map that passes the verification.

[0244] Referring to FIG. 15, a structural schematic diagram of a map detection device is provided in an embodiment of the present application, and the map detection device can be used to execute the process of the inference stage. The map detection device comprises:

[0245] The first acquisition module 1501 is configured to acquire sensor data, and the sensor data comprises data collected by at least one sensor.

[0246] The first feature extraction module 1502 is configured to extract features from the sensor data to obtain perception features.

[0247] The second acquisition module 1503 is configured to acquire a first map.

[0248] The second feature extraction module 1504 is configured to extract features from the first map to obtain first map features.

[0249] The fusion module 1505 is configured to fuse the perception features and the first map features to obtain fused features.

[0250] The decoding module 1506 is configured to decode according to the fused features to obtain a second map, the accuracy of the second map being higher than that of the first map, and the second map including vectors of multiple types.

[0251] In a possible implementation, the decoding module 1506 is configured to input the fused features and the first map features into at least one decoder, and obtain the second map according to an output of the at least one decoder, the at least one decoder being configured to output vectors of at least one type respectively.

[0252] In a possible implementation, the at least one decoder is configured to perform at least one of instance detection, road network completion, or semantic segmentation based on the input features.

[0253] In a possible implementation, the second map includes vectors of one or more of the following types of entities: lane boundary lines, roads, intersections, or lane center lines.

[0254] In a possible implementation, the first map data includes data of a low-precision map.

[0255] In a possible implementation, the first feature extraction module 1502 is configured to extract features from sensor data to obtain bird's eye view (BEV) features, and the perception features include the BEV features.

[0256] The fusion module 1505 is configured to: input the first map features as input of a perception module, and output second map features, the perception module being configured to downsample the first map features at least once and fuse the downsampled features with the perception features; and fuse the BEV features and the second map features to obtain the fused features.

[0257] In a possible implementation, the fusion module 1505 is configured to: interact the perception features and the first map features through an attention mechanism to obtain the fused features.

[0258] As shown in FIG. 16, FIG. 16 is a hardware structure schematic diagram of a model training apparatus 160 provided by an embodiment of the present application. The model training apparatus 160 can be used to implement the steps of the foregoing method in the training stage.

[0259] The model training apparatus 160 shown in FIG. 16 can include a processor 1601, a memory 1602, a communication interface 1603, and a bus 1604. The processor 1601, the memory 1602, and the communication interface 1603 can be connected through the bus 1604.

[0260] The processor 1601 is a control center of the model training apparatus 160, and can be a general central processing unit (CPU), or other general-purpose processors, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc., which can specifically include a GPU or an NPU, etc., and can be adaptively set according to actual application scenarios.

[0261] As an example, the processor 1601 can include one or more CPUs, and can also include other processors, such as the CPU, NPU, or GPU shown in FIG. 16, etc.

[0262] The memory 1602 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, and can also be an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program codes in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0263] In one possible implementation, the memory 1602 can exist independently of the processor 1601. The memory 1602 can be connected to the processor 1601 through the bus 1604, for storing data, instructions, or program codes. When the processor 1601 invokes and executes the instructions or program codes stored in the memory 1602, the method provided by the embodiments of the present application can be implemented, for example, the method shown in the training phase.

[0264] In another possible implementation, the memory 1602 can also be integrated with the processor 1601.

[0265] The communication interface 1603 is configured to connect the model training apparatus 160 to other devices through a communication network, which can be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), or the like. The communication interface 1603 can include a receiving unit configured to receive data, and a transmitting unit configured to transmit data.

[0266] The bus 1604 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, only one thick line is shown in FIG. 16, but it does not mean that there is only one bus or only one type of bus.

[0267] It should be noted that the structure shown in FIG. 16 does not constitute a limitation on the model training apparatus 160. In addition to the components shown in FIG. 16, the model training apparatus 160 can include more or fewer components than those shown, or combine certain components, or different arrangement of components.

[0268] As shown in FIG. 17, a hardware structure schematic diagram of a map detection apparatus 170 provided by an embodiment of the present application is shown. The map detection apparatus 170 can be used to implement the steps of the method of the foregoing inference stage.

[0269] The map detection apparatus 170 shown in FIG. 17 can include a processor 1701, a memory 1702, a communication interface 1703, and a bus 1704. The processor 1701, the memory 1702, and the communication interface 1703 can be connected through the bus 1704.

[0270] The processor 1701 is the control center of the map detection apparatus 170, which can be a general central processing unit (CPU), or other general-purpose processors, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc., which can specifically include a GPU or an NPU, etc., and can be adaptively set according to actual application scenarios.

[0271] As an example, the processor 1701 can include one or more CPUs, and can also include other processors, such as the CPU, NPU, or GPU shown in FIG. 17, etc.

[0272] The memory 1702 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited thereto.

[0273] In a possible implementation, the memory 1702 can exist independently of the processor 1701. The memory 1702 can be connected to the processor 1701 through the bus 1704, for storing data, instructions, or program codes. When the processor 1701 invokes and executes the instructions or program codes stored in the memory 1702, the method provided by the embodiments of the present application can be implemented, for example, the method shown in the foregoing reasoning phase.

[0274] In another possible implementation, the memory 1702 can also be integrated with the processor 1701.

[0275] The communication interface 1703 is configured to connect the map detection apparatus 170 to other devices through a communication network, which can be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), or the like. The communication interface 1703 can include a receiving unit configured to receive data, and a sending unit configured to send data.

[0276] The bus 1704 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, only one thick line is shown in FIG. 17, but it does not mean that there is only one bus or only one type of bus.

[0277] It should be noted that the structure shown in FIG. 17 does not constitute a limitation on the map detection device 170, and the map detection device 170 can include more or fewer components than those shown in FIG. 17, or combine certain components, or different component arrangements, in addition to the components shown in FIG. 17.

[0278] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary general hardware, and of course can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., including a plurality of instructions for causing an apparatus (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0279] In the above embodiments, all or part can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in the form of a computer program product in whole or in part.

[0280] The computer readable storage medium in the embodiments of the present application stores a program for training a model or performing an inference task, which, when running on a computer, causes the computer to execute all or part of the steps of the methods described in the foregoing embodiments of FIGS. 3 to 14.

[0281] The embodiments of the present application also provide a digital processing chip. The digital processing chip integrates a circuit for implementing the above processor or the function of the processor and one or more interfaces. When the digital processing chip integrates a memory, the digital processing chip can complete the method steps of any one or more of the foregoing embodiments. When the digital processing chip does not integrate a memory, it can be connected with an external memory through a communication interface. The digital processing chip implements the method steps of any one or more of the above embodiments according to the program code stored in the external memory.

[0282] The embodiments of the present application further provide a computer program product including one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be stored by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0283] The map detection device or model training device provided by the embodiments of the present application can be a chip, which includes a processing unit, for example, a processor, and a communication unit, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute computer execution instructions stored in a storage unit to enable the chip in the device to execute the methods described in the embodiments shown in FIGS. 3 to 14. Alternatively, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit can also be a storage unit outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0284] Specifically, the foregoing processing unit or processor can be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and the like. The general-purpose processor can be a microprocessor or any conventional processor, and the like.

[0285] In addition, it should be noted that the apparatus embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments. In addition, the connection relationship between the modules in the apparatus embodiments provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.

[0286] Those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware, and of course can also be implemented by means of dedicated hardware including special-purpose integrated circuits, special-purpose CPUs, special-purpose memories, special-purpose components, etc. Generally, any function completed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuits, digital circuits, or special-purpose circuits, etc. However, for the present application, software program implementation is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a floppy disk, a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0287] In the above embodiments, all or part can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part can be implemented in the form of a computer program product.

[0288] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0289] The terms "first", "second", and the like in the description and in the claims of the present application and above drawings are used for distinguishing between similar objects and not necessarily for describing a specific sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances such that the embodiments of the application described herein are, for example, capable of orderly or inverse order, depending upon the circumstances. The term "and / or" in the present application is merely used to represent an association between associated objects, and it is possible that three relationships exist, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the associated objects. Furthermore, the terms "include" and "have" and any variations thereof, are intended to cover a non-exclusive inclusion, for example, a process, method, system, product or device that includes a list of steps or modules as an example does not have to be limited to those steps or modules, but can include other steps or modules that are not expressly listed or inherent to such process, method, product or device. The naming or numbering of steps in the present application does not mean that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering, and the named or numbered flow steps can be executed in a different order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of modules in the present application is a logical division, and in actual application, there can be another division manner, for example, multiple modules can be combined or integrated in another system, or some features can be ignored or not executed, in addition, the coupling or direct coupling or communication connection between the displayed or discussed modules can be through some ports, and the indirect coupling or communication connection between the modules can be electrical or other similar forms, which are not limited in the present application. Furthermore, the modules or sub-modules described as separate components can or can not be physically separated, and can or can not be physical modules, or can be distributed to multiple circuit modules, and part or all of the modules can be selected according to actual needs to achieve the purpose of the present application.

Claims

1. A map detection method characterized by, The method comprises: obtaining sensor data, the sensor data comprising data collected by at least one sensor; extracting features from the sensor data to obtain perception features; obtaining a first map; extracting features from the first map to obtain first map features; fusing the perception features and the first map features to obtain fused features; decoding the fused features to obtain a second map, the second map having a higher accuracy than the first map, and the second map comprising vectors of multiple types.

2. The method of claim 1, wherein, The decoding the fused features to obtain a second map comprises: inputting the fused features and the first map features into at least one decoder, and obtaining the second map according to an output of the at least one decoder, the at least one decoder being configured to output vectors of at least one type.

3. The method of claim 2, wherein: the at least one decoder is configured to perform at least one of instance detection, road network completion, or semantic segmentation based on the input features.

4. The method according to any one of claims 1 to 3, characterized in that, the second map comprises vectors representing one or more of the following types of entities: lane boundary lines, roads, intersections, or lane center lines.

5. The method according to any one of claims 1-4, characterized in that, the first map data comprises data of a low-precision map.

6. The method according to any one of claims 1-5, characterized in that, The extracting features from the sensor data to obtain perception features comprises: extracting bird's eye view (BEV) features from the sensor data, and the perception features comprise the BEV features. The fusing the perception features and the first map features to obtain fused features comprises: inputting the first map features as input of a perception module, and outputting second map features, the perception module being configured to downsample the first map features at least once and fuse the downsampled features with the perception features. The fusing the perception features and the first map features to obtain fused features comprises:

7. The method according to any one of claims 1-5, characterized in that, interacting the perception features and the first map features through an attention mechanism to obtain the fused features. The method comprises:

8. A model training method, comprising: obtaining training data, the training data comprising map data; training a map detection model using the training data to obtain a trained map detection model, the map detection model being configured to: extract first map features from a first map, extract perception features from sensor data, fuse the perception features and the first map features to obtain fused features, and process the fused features and the first map features to obtain a second map, the second map having a higher accuracy than the first map, and the second map comprising vectors of multiple types. The training data comprises data of a low-precision map.

9. The method of claim 8, wherein, The obtaining training data comprises:

10. The method of claim 9, wherein, obtaining data of an initial low-precision map; aligning at least one instance in the initial low-precision map, the at least one instance being of a type that is included in instances output by the map detection model. The obtaining training data further comprises:

11. The method according to claim 9 or 10, characterized in that, ​ In the case of obtaining data of an initial low-precision map, repairing the initial low-precision map to obtain the low-precision map data, the repairing including broken line reconnection, shape point correction, or layer deletion.

12. The method according to any one of claims 9-11, characterized in that, The obtaining training data further includes: In the case of obtaining data of an initial low-precision map, repairing the initial low-precision map to obtain the low-precision map data, the repairing including broken line reconnection, shape point correction, or layer deletion.

13. A map detection device, characterized by comprising: Comprise: The first acquisition module is used for acquiring sensor data, and the sensor data includes data collected by at least one sensor. The first feature extraction module is used for extracting features from the sensor data to obtain perception features. The second acquisition module is used for acquiring a first map. The second feature extraction module is used for extracting features from the first map to obtain first map features. The fusion module is used for fusing the perception features and the first map features to obtain fusion features. The decoding module is used for decoding according to the fusion features to obtain a second map, the precision of the second map being higher than that of the first map, and the second map including multiple types of vectors.

14. The apparatus of claim 13, wherein, The decoding module is specifically used for: Inputting the fusion features and the first map features into at least one decoder to obtain the second map according to outputs of the at least one decoder, the at least one decoder being respectively used for outputting at least one type of vector.

15. The apparatus of claim 14, wherein The at least one decoder is respectively used for performing at least one of instance detection, road network completion, or semantic segmentation based on the input features.

16. The apparatus of any one of claims 13-15, wherein, The second map includes vectors representing one or more types of entities: lane boundary lines, roads, intersections, or lane center lines.

17. The apparatus of any one of claims 13-16, wherein, The first map data includes data of a low-precision map.

18. The apparatus of any one of claims 13-17, wherein The first feature extraction module is specifically used for extracting features from the sensor data to obtain bird's eye view (BEV) features, and the perception features include the BEV features. The fusion module is specifically used for: Taking the first map features as inputs of a perception module to output second map features, the perception module being used for performing at least one down-sampling on the first map features and fusing the at least one down-sampled feature and the perception features; Fusing the BEV features and the second map features to obtain the fusion features.

19. The apparatus of any of claims 13-18, wherein, The fusion module is specifically used for, comprising: Interacting the perception features and the first map features through an attention mechanism to obtain the fusion features.

20. A model training apparatus, comprising: Comprise: The acquisition module is used for acquiring training data, and the training data includes map data. The training module is configured to train the map detection model by using the training data, to obtain a trained map detection model, and the map detection model is configured to: extract first map features from a first map, extract perception features from sensor data, fuse the perception features and the first map features to obtain fused features, and process the fused features and the first map features to obtain a second map, the second map having a higher accuracy than the first map, and the second map including multiple types of vectors.

21. The apparatus of claim 20, wherein, The training data includes low-precision map data.

22. The apparatus of claim 21, wherein, The obtaining module is specifically configured to: obtain data of an initial low-precision map; align at least one instance in the initial low-precision map, the type of the at least one instance including a type of an instance in a map output by the map detection model.

23. The apparatus of claim 21 or 22, wherein, The obtaining module is specifically configured to: in a case where the data of the initial low-precision map is obtained, repair the initial low-precision map to obtain the low-precision map data, and the repairing includes line reconnection, shape point correction, or layer deletion.

24. The apparatus of any one of claims 21-23, wherein, The obtaining module is specifically configured to: in a case where the data of the initial low-precision map is obtained, verify the initial low-precision map to obtain the low-precision map that passes the verification.

25. A map detection device, characterized by A processor is included, the processor is coupled with a memory, and the memory stores a program, and when program instructions stored in the memory are executed by the processor, the steps of the method in any one of claims 1-7 are implemented.

26. A model training apparatus, comprising: A processor is included, the processor is coupled with a memory, and the memory stores a program, and when program instructions stored in the memory are executed by the processor, the steps of the method in any one of claims 8-12 are implemented.

27. A vehicle characterized by A processor and a memory are included, and the memory stores a program, and when program instructions stored in the memory are executed by the processor, the steps of the method in any one of claims 1-7 or 8-12 are implemented.

28. A computer-readable storage medium, characterized in that, A program is included, and when the program is executed by a processing unit, the steps of the method in any one of claims 1-7 or 8-12 are executed.

29. A computer program product, characterised in that, The computer program product includes software code for executing the steps of the method in any one of claims 1-7 or 8-12.

Citation Information

Patent Citations

  • Map element detection method and device, model training method and device and electronic equipment

    CN117611547A

  • Online semantic vector map construction method based on navigation map and vision

    CN118115963A

  • Trajectory prediction method and device, electronic equipment and storage medium

    CN118533187A

  • Reconciliation of Map Data and Sensor Data

    US20240133697A1

  • Temporal-based perception for autonomous systems and applications

    US20240312219A1

Cited By

  • Methods, devices, and media for road structure reconstruction based on vehicle-mounted panoramic images

    CN122312953A