Map construction method, electronic equipment and storage medium

By using a mixed feature map and a mixed decoder, the point level and element level information of map elements are extracted interactively, and the problem of insufficient map accuracy in the prior art is solved, and a higher precision map construction is achieved.

CN120014184APending Publication Date: 2025-05-16BEIJING SAMSUNG TELECOM R&D CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311527475.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing high-precision map construction method uses point query representation, and it is difficult to fully express the details of map elements, resulting in insufficient map accuracy.

Method used

Using a mixed feature map, including point features and element features, interact with the bird's-eye feature map through a hybrid decoder, determine map information and build a high-precision map.

Benefits of technology

Through mutual refinement and integration of information, the shape integrity and position accuracy of map elements are improved, and the accuracy of maps is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014184A_ABST
    Figure CN120014184A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a map construction method, electronic equipment and a storage medium, and relates to the field of high-precision maps. The method comprises the following steps: acquiring to-be-processed data; extracting a bird's-eye view feature map according to the data; according to the aerial view feature map and the mixed feature map, based on a mixed decoder, map information is determined, and the map information comprises coordinate information of at least one coordinate point in each map element and category information to which each map element belongs; constructing a map corresponding to the data based on the map information; wherein the map comprises a plurality of map elements, each map element comprises an area formed by a plurality of coordinate points in a map area, the mixed feature map comprises a plurality of mixed features, each mixed feature corresponds to one map element, and each mixed feature comprises a point feature and an element feature. Optionally, the method may be performed using an artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of high-precision maps. Specifically, the present application relates to a map construction method, an electronic device and a storage medium. Background Art

[0002] The task of building a high-precision map can be viewed as the problem of predicting a set of vectorized static map elements under a bird's-eye view (BEV). The element categories of map elements include, for example, crosswalks, lane dividers, road boundaries, etc. High-precision maps (HD maps) provide rich and accurate static environmental information of driving scenarios, which is crucial and challenging for downstream tasks such as autonomous driving system planning and HD map automatic annotation systems.

[0003] At present, in order to build high-precision maps, the vectorized map construction algorithm of MapTR and MapTRv2 is usually used. The feature map (BEV features) of the BEV space is obtained through the map encoder (Map Encoder), and then the vectorized map elements are obtained through the map decoder (Map Decoder). The input parameters of its core module Map Decoder are BEV features (BEV feature) and point query representation (point query), that is, each point on the map element is represented by a set of parameters, and the output parameters are the element category (class) and point coordinates (point coordinate) of the map element. However, this method uses point query representation, and it is difficult for limited points to fully express the details of the map elements, which ultimately leads to insufficient accuracy of the constructed map. Summary of the invention

[0004] In order to at least solve the above problems existing in the prior art, the present invention provides a map construction method, an electronic device and a storage medium.

[0005] According to a first aspect of an embodiment of the present application, a high-precision map construction method is provided, comprising: acquiring data to be processed; extracting a bird's-eye view feature map based on the data; determining map information based on the bird's-eye view feature map and the hybrid feature map and based on a hybrid decoder, wherein the map information includes coordinate information of at least one coordinate point in each map element and category information to which each map element belongs; constructing a map corresponding to the data based on the map information; wherein the map includes a plurality of map elements, each map element includes an area formed by a plurality of coordinate points within a map area, the hybrid feature map includes a plurality of hybrid features, each hybrid feature corresponds to a map element, and each hybrid feature includes a point feature and an element feature, the point feature is used to describe information related to each coordinate point in the corresponding map element, and the element feature is used to describe information related to the corresponding map element.

[0006] Optionally, according to the bird's-eye view feature map and the mixed feature map, based on a mixed decoder, the step of determining map information includes: decomposing the mixed feature map into a first point feature map and a first element feature map, the first point feature map including first point features corresponding to each coordinate point in each map element, and the first element feature map including first element features corresponding to each map element; determining a second point feature map and a second element feature map respectively according to the bird's-eye view feature map, the first point feature map, the first element feature map and the current map information; fusing the second point feature map and the second element feature map to update the mixed feature map; updating the map information according to the bird's-eye view feature map and the updated mixed feature map, and returning to the operation of decomposing the mixed feature map into a first point feature map and a first element feature map to perform the next update; based on the map information, the step of constructing a map corresponding to the data includes: constructing a map corresponding to the data based on the final map information.

[0007] Optionally, the map information includes coordinate information of coordinate points, and the steps of respectively determining a second point feature map and a second element feature map according to the bird's-eye feature map, the first point feature map, the first element feature map and the current map information include: for each reference point, determining the second point feature of the reference point according to the bird's-eye feature map, the first point feature of the reference point and the coordinate information of the reference point, wherein the reference point includes the coordinate point corresponding to each first point feature; fusing the second point features of each reference point to obtain the second point feature map; for each map element, determining the second element feature of the map element according to the bird's-eye feature map, the first element feature of the map element and the coordinate information of each reference point in the map element; fusing the second element features of each map element to obtain the second element feature map.

[0008] Optionally, for each reference point, the step of determining the second point feature of the reference point according to the bird's-eye feature map, the first point feature of the reference point and the coordinate information of the reference point includes: for each reference point, determining a number of sampling points associated with the reference point from the map area according to the coordinate information and the first point feature of the reference point; performing fusion processing based on the bird's-eye feature map and the coordinate information and weight of each sampling point associated with the reference point to obtain the third point feature of the reference point; and determining the second point feature of the reference point according to the first point feature and the third point feature of the reference point.

[0009] Optionally, for each reference point, according to the coordinate information and the first point feature of the reference point, the step of determining a number of sampling points associated with the reference point from the map area includes: for each reference point, according to the coordinate information and the first point feature of the reference point, determining a fourth point feature of the reference point, wherein the fourth point feature is used to represent the point feature after considering the influence of the coordinate information; according to the fourth point feature of the reference point, determining a sampling offset and a weight of a number of sampling points associated with the reference point, wherein the sampling offset is used to represent the degree of positional offset of the sampling point relative to the reference point; and determining the coordinate information of each sampling point of the reference point according to the coordinate information of the reference point and the sampling offset of each sampling point.

[0010] Optionally, the step of determining the fourth point feature of the reference point based on the coordinate information and the first point feature of the reference point includes: encoding the coordinate information of the reference point to obtain the position code of the reference point; and determining the fourth point feature of the reference point based on the first point feature and the position code of the reference point.

[0011] Optionally, the step of performing fusion processing based on the bird's-eye view feature map and the coordinate information and weight of each sampling point associated with the reference point to obtain the third point feature of the reference point includes: determining the sampling feature of each sampling point corresponding to the reference point according to the bird's-eye view feature map and the coordinate information of each sampling point associated with the reference point; and performing fusion processing on the sampling features of the reference point corresponding to each sampling point based on the weight of each sampling point associated with the reference point to obtain the third point feature of the reference point.

[0012] Optionally, for each map element, the step of determining the second element feature of the map element according to the bird's-eye feature map, the first element feature of the map element and the coordinate information of each reference point in the map element includes: for each map element, encoding the coordinate information of each reference point in the map element to obtain the position code of each reference point; fusing the position codes of each reference point in the map element to obtain the position code of the map element; determining the second element feature of the map element according to the bird's-eye feature map, the first element feature of the map element and the position code of the map element using the mask attention module in the hybrid decoder, wherein the mask used by the mask attention module is obtained based on the mask information of each pixel, and the mask information is used to indicate the probability that the corresponding pixel belongs to the map element.

[0013] Optionally, the step of fusing the second point feature map and the second element feature map to update the mixed feature map includes: using the self-attention module in the mixed decoder to process the second point feature map and the second element feature map respectively to obtain a fifth point feature map and a fifth element feature map; converting the fifth point feature map to the same dimension as the fifth element feature map, and fusing the fifth element feature map and the converted fifth point feature map to obtain a sixth element feature map; converting the fifth element feature map to the same dimension as the fifth point feature map, and fusing the fifth point feature map and the converted fifth element feature map to obtain a sixth point feature map; fusing the sixth point feature map and the sixth element feature map to obtain the updated mixed feature map.

[0014] Optionally, the loss function used by the hybrid decoder during training includes a point-element consistency loss, which is used to represent the risk level that the point feature map and the element feature map in the updated hybrid feature map are inconsistent with each other.

[0015] Optionally, the value of the point-element consistency loss is determined by the following method: transforming the point feature map and the element feature map in the updated hybrid feature map respectively to obtain point-level information and element-level information; fusing the information of the coordinate points belonging to the same map element in the point-level information to obtain pseudo-element level information; determining the value of the point-element consistency loss based on the pseudo-element level information and the element level information to indicate the risk level of inconsistency between the pseudo-element level information and the element level information.

[0016] Optionally, the loss function used by the hybrid decoder during training also includes at least one of the following: semantic segmentation loss, classification loss, point regression loss, point direction loss, and mask loss.

[0017] According to a second aspect of an embodiment of the present application, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer executable instructions, wherein the computer executable instructions, when executed by the at least one processor, prompt the at least one processor to execute the high-precision map construction method as described above.

[0018] According to a third aspect of an embodiment of the present application, a computer-readable storage medium storing instructions is provided, wherein, when the instructions are executed by at least one processor, the at least one processor is prompted to execute the high-precision map construction method as described above.

[0019] The beneficial effects brought about by the technical solutions provided by the embodiments of the present application will be described below in conjunction with specific optional embodiments, or can be learned from the description of the embodiments, or can be known through the implementation of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly and easily illustrate and understand the technical solutions in the embodiments of the present application, the following is a brief introduction to the drawings required for use in describing the embodiments of the present application.

[0021] Figure 1 It is a schematic diagram showing a high-precision map.

[0022] Figure 2 is a flowchart illustrating a map construction method according to an exemplary embodiment of the present application.

[0023] Figure 3 FIG. 4 is a flowchart illustrating a map construction method according to an exemplary embodiment of the present application.

[0024] Figure 4 is a system block diagram illustrating a map construction method according to an exemplary embodiment of the present application.

[0025] Figure 5 FIG. 4 is a flowchart illustrating a map construction method according to an exemplary embodiment of the present application.

[0026] Figure 6 is a schematic diagram showing a hybrid decoder workflow according to an exemplary embodiment of the present application.

[0027] Figure 7 is a schematic diagram showing a hybrid decoder workflow according to an exemplary embodiment of the present application.

[0028] Figure 8 is a flowchart diagram showing steps of updating a hybrid feature according to an exemplary embodiment of the present application.

[0029] Fig. 9 is a schematic diagram showing a calculation flow of a point-element consistency loss according to an exemplary embodiment of the present application.

[0030] Fig.10 is a schematic diagram showing a calculation flow of a point-element consistency loss according to an exemplary embodiment of the present application.

[0031] Fig.11 It is a schematic diagram showing the difference in technical concept between the map construction method according to the exemplary embodiment of the present application and the related art.

[0032] Fig.12 is a schematic diagram illustrating the accuracy improvement effect of the map construction method according to an exemplary embodiment of the present application.

[0033] Fig.13 is a schematic structural diagram showing an electronic device according to an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0034] The following description with reference to the accompanying drawings is provided to facilitate a comprehensive understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. This description includes various specific details to facilitate understanding but should be considered as exemplary only. Therefore, one of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, for the sake of clarity and conciseness, descriptions of well-known functions and structures may be omitted.

[0035] The terms and expressions used in the following specification and claims are not limited to their dictionary meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present disclosure. Therefore, it should be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustration purposes only and not for the purpose of limiting the present disclosure as defined in the appended claims and their equivalents.

[0036] It should be understood that the singular forms "a", "an", and "the" may also include plural references unless the context clearly indicates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more such surfaces. When we refer to an element as being "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or the one element and the other element may establish a connection relationship through an intermediate element. In addition, "connected" or "coupled" as used herein may include wireless connection or wireless coupling.

[0037] The term "include" or "may include" refers to the presence of the corresponding disclosed functions, operations or components that can be used in various embodiments of the present disclosure, rather than limiting the presence of one or more additional functions, operations or features. In addition, the term "include" or "have" may be interpreted as indicating certain characteristics, numbers, steps, operations, constituent elements, components or combinations thereof, but should not be interpreted as excluding the possibility of the presence of one or more other characteristics, numbers, steps, operations, constituent elements, components or combinations thereof.

[0038] The term "or" used in various embodiments of the present disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple, or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" may be implemented as parameter A including A1 or A2 or A3, or may be implemented as parameter A including at least two of the three items A1, A2, and A3.

[0039] Unless defined differently, all terms (including technical terms or scientific terms) used in the present disclosure have the same meanings as understood by those skilled in the art described in the present disclosure. Common terms as defined in dictionaries are interpreted as having meanings consistent with the context in the relevant technical field, and should not be interpreted ideally or overly formally unless clearly defined in the present disclosure.

[0040] At least some functions of the device or electronic device provided in the embodiments of the present disclosure can be implemented by an AI model, such as at least one module among multiple modules of the device or electronic device can be implemented by an AI model. Functions associated with AI can be performed by non-volatile memory, volatile memory and processor.

[0041] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., or pure graphics processing units, such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-specific processor, such as a neural processing unit (NPU).

[0042] The one or more processors control the processing of input data according to predefined operating rules or artificial intelligence (AI) models stored in non-volatile memory and volatile memory. The predefined operating rules or artificial intelligence models are provided by training or learning.

[0043] Here, providing by learning means obtaining a predefined operating rule or an AI model with desired characteristics by applying a learning algorithm to a plurality of learning data. The learning can be performed in the device or electronic device itself in which the AI ​​according to the embodiment is executed, and / or can be implemented by a separate server / system.

[0044] The AI ​​model may include multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network calculations by calculating between the input data of the layer (such as the calculation results of the previous layer and / or the input data of the AI ​​model) and the multiple weight values ​​of the current layer. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.

[0045] A learning algorithm is a method of using a plurality of learning data to train a predetermined target device (e.g., a robot) to enable, allow, or control the target device to make a determination or prediction. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0046] The method provided in the present disclosure may involve one or more technical fields such as speech, language, image, video or data intelligence.

[0047] Optionally, when it comes to the field of speech or language, in a method performed by an electronic device according to the present disclosure, a speech signal as an analog signal can be received via a speech input device (e.g., a microphone), and the speech portion can be converted into computer-readable text using an automatic speech recognition (ASR) model. The user's speech intention can be obtained by interpreting the converted text using a natural language understanding (NLU) model. The ASR model or the NLU model can be an artificial intelligence model. The artificial intelligence model can be processed by an artificial intelligence dedicated processor designed in a hardware structure specified for artificial intelligence model processing. Language understanding is a technology for recognizing and applying / processing human language / text, including, for example, natural language processing, machine translation, dialogue systems, question answering, or speech recognition / synthesis.

[0048] Optionally, when it comes to the field of images or videos, in the method performed by the electronic device according to the present disclosure, output data can be obtained by using image data as input data of an artificial intelligence model. The method of the present disclosure may be related to the field of visual understanding of artificial intelligence technology, which is a technology for recognizing and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.

[0049] Optionally, when it comes to the field of intelligent data processing, in the method performed by the electronic device according to the present disclosure, in the inference or prediction stage, an artificial intelligence model can be used to perform predictions by using real-time input data. The processor of the electronic device can perform preprocessing operations on the data to convert it into a form suitable for use as an input to the artificial intelligence model. Inference prediction is a technology for logical reasoning and prediction by determining information, including, for example, knowledge-based reasoning, optimization prediction, preference-based planning or recommendation.

[0050] In the present application, an artificial intelligence model can be obtained by training. Here, "obtained by training" means obtaining a predefined operating rule or artificial intelligence model configured to perform a desired feature (or purpose) by training a basic artificial intelligence model with multiple training data through a training algorithm. The artificial intelligence model may include multiple neural network layers. Each of the multiple neural network layers includes multiple weight values, and the neural network calculation is performed by calculating between the calculation result of the previous layer and the multiple weight values.

[0051] The following describes several optional embodiments to illustrate the technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure. It should be noted that the following embodiments can refer to, draw on or combine with each other, and the same terms, similar features and similar implementation steps in different embodiments will not be described repeatedly.

[0052] A high-definition map (HD map) is a high-precision map used for autonomous driving, which contains map elements such as road shapes, road markings, traffic signs and obstacles. Figure 1 Figure 1 shows a part of a high-precision map. Figure 1As shown, the high-precision map includes a plurality of map elements 100. Each map element 100 is usually represented as a set of coordinate points 110 in the map. These coordinate points 110 are connected into multiple line segments or polygons, so as to represent a meaningful map area instance. The coordinate point 110 is a point in a predefined coordinate system, and its position can be represented by the coordinate information (point coordinate) in the coordinate system. In the process of constructing a high-precision map, a number of coordinate points 110 are initially allocated to each map element 100, but the specific positions of these coordinate points 110 are unknown, that is, the coordinate information of these coordinate points 110 is unknown. The coordinate information of these coordinate points 110 can be continuously updated during the update process, so as to determine the coordinate point 110. There are also pixels in the high-precision map. Pixels are fixed points in the image and will not change during the above-mentioned update process. According to the different meanings represented by the map element 100, different element categories can be configured for the map element 100, for example, crosswalks, lane dividers, road boundaries, etc., which are used to mark different map elements 100. With Figure 1 For example, there are 5 map elements 100, which, from the perspective of element categories, specifically include 2 road boundary map elements (green polylines), 2 lane separator map elements (yellow polylines) and 1 crosswalk map element (blue polygon).

[0053] High-precision maps can be divided into local maps and global maps according to the length of time and distance. Local maps are short-distance and are usually constructed from one frame of data. A frame of data can be single-modal data, such as multi-perspective camera images (camera images, usually RGB images in this case) or point cloud data obtained by LiDAR; it can also be multi-modal data, such as camera images and point clouds; it can also include posture data, that is, coordinate transformation information between different modal data to map different modal data to the same coordinate system. Global maps are long-distance and are usually constructed from data of a scene, and a scene is a multi-frame sequence.

[0054] High-precision map construction is based on the original data (such as the camera images and point cloud data mentioned above) to generate Figure 1 The vectorized static map element set shown. Existing high-precision map construction methods only use point query representation, which makes it difficult to express the details of map elements, lacks learning of the overall information of map elements (such as length and direction), and easily causes confusion and entanglement between different map elements, resulting in insufficient accuracy of the constructed map.

[0055] To this end, the present application provides a high-precision map construction method, which uses a hybrid feature map including a point feature map and an element feature map to describe element-level information and point-level information, and performs hybrid decoding on the bird's-eye view feature map and the hybrid feature map. It can realize the interaction of element-level information and point-level information, and realize the mutual refinement and integration of information, so that the final map has a more complete shape and more accurate position, greatly improving the accuracy of the constructed map.

[0056] The following will refer to Figures 2 to 13 The technical solution proposed in this application is described in detail.

[0057] Figure 2 is a flowchart illustrating a map construction method according to an exemplary embodiment of the present application. Figure 3 FIG. 4 is a flowchart illustrating a map construction method according to an exemplary embodiment of the present application. Figure 4 is a system block diagram illustrating a map construction method according to an exemplary embodiment of the present application. Figure 5 FIG. 4 is a flowchart illustrating a map construction method according to an exemplary embodiment of the present application.

[0058] like Figure 2 and Figure 5 As shown, in step S210, data to be processed is obtained.

[0059] The data to be processed is the data used to construct the map as mentioned above.

[0060] In step S220, a bird's-eye view feature map is extracted based on the data.

[0061] This step is used to extract features from the data obtained in step S210, such as Figure 4 and Figure 5 As shown, a bird's-eye view feature extractor 320 may be used to implement extraction by applying methods in related technologies.

[0062] Specifically, the bird's-eye view feature map is the feature map of the bird's-eye view space. If the data used is a multi-view RGB image, we first use the backbone network (backbone network, such as the commonly used existing networks such as Resnet and Swin Transformer) to extract multi-scale 2D features at each view, and then use FPN (Feature Pyramid Network) to fuse features of different scales to obtain a single-scale fused 2D feature map, and finally use the spatial variation module (specifically, the feature transformation from 2D space to bird's-eye view space, which is the prior art) to convert the 2D feature map into a bird's-eye view feature map. If the input visual data is a laser point cloud, voxelized features can also be obtained through a 3D backbone network (such as SECOND) and flattened into a bird's-eye view feature map. If the data used is two-modal data, the bird's-eye view feature maps obtained from different modalities are connected together, and then a convolution operation is performed to obtain a multi-modal fused bird's-eye view feature map. The multimodal fusion method is also a prior art.

[0063] In summary, for unimodal or multimodal data, the outputs of the bird's-eye view feature extractor 320 are all features in the same space (e.g., bird's-eye view space). Figure X , X is represented by a tensor of H*W*C, where H and W represent the height and width of the image represented by the data, respectively, and C represents the number of channels of the feature map. In order to learn a better bird's-eye view feature map, semantic segmentation loss can be used to implement supervised training during the training process of the bird's-eye view feature extractor 320.

[0064] In step S230, map information is determined based on the hybrid decoder according to the bird's-eye view feature map and the hybrid feature map. The map information includes coordinate information of at least one coordinate point in each map element and category information to which each map element belongs. It should be understood that the map information described here includes information of "each" map element, which can be all map elements in the calculation or part of the map elements remaining after all map elements are screened according to certain rules. Here, only the retained part is indicated. Other "each" and "each" in the following text also have this meaning and will not be explained one by one. The map includes a plurality of map elements, each of which includes an area formed by a plurality of coordinate points in the map area, and the hybrid feature map includes a plurality of hybrid features, each of which corresponds to a map element, and each of which includes a point feature and an element feature, and the point feature is used to describe information related to each coordinate point in the corresponding map element, and the element feature is used to describe information related to the corresponding map element.

[0065] The hybrid feature map is essentially a query representation, and is therefore also called a hybrid query representation (HybrIdQuery, HIQuery for short). The hybrid feature map is a set of learnable parameters. , where E represents the maximum number of map elements, which is a predefined parameter and can be set to a sufficiently large number to ensure that the required number of map elements is covered. P represents the maximum number of coordinate points on each map element, 1 represents the element category to which the map element belongs, and C still represents the number of channels in the feature map (the same symbols will not be repeated in the following text). Each element Corresponding to a map element, i∈{1,...,E} is the index of the map element. It consists of two parts: and (i.e., point features and element features), respectively representing the point-level information and element-level information of the i-th map element. Since the hybrid feature map contains integrated information at the point and element levels, it can be used to generate map information for each map element, including coordinate information (e.g., point coordinates), element category information, and mask information (through Figures 4 to 6 The mixed feature map is randomly initialized, that is, it is obtained by randomly giving an initial value to the mixed feature map, and is gradually updated through interaction with the bird's-eye view feature map.

[0066] Optionally, step S230 includes: decomposing the mixed feature map into a first point feature map and a first element feature map, the first point feature map including first point features corresponding to each coordinate point in each map element, and the first element feature map including first element features corresponding to each map element; determining a second point feature map and a second element feature map respectively according to the bird's-eye view feature map, the first point feature map, the first element feature map and the current map information; fusing the second point feature map and the second element feature map to update the mixed feature map; updating the map information according to the bird's-eye view feature map and the updated mixed feature map, and returning to the operation of decomposing the mixed feature map into a first point feature map and a first element feature map to perform the next update. By repeated looping, continuous interaction and feature updating can be achieved to improve the accuracy of the final mixed feature map. As an example, an end condition can be set to end the loop, and the end condition can be that the number of loops reaches a set number of times.

[0067] Taking the lth layer as an example, the mixed features (The superscript l-1 indicates that the mixed feature is the result of the l-1th layer update) is decomposed into two parts: the initial point feature and the initial element feature, using the following formula (1).

[0068] [Q p,l-1 ,Q e,l-1]=Q h,1-1 , (1)

[0069] Here [,] represents the connection relationship, that is, the initial point feature and initial element characteristics Connected to form the initial mixed feature

[0070] By decomposing point features and element features, point features and element features can also interact with each other in subsequent interactions, so as to interactively extract point-level and element-level information of each map element and encode it into new hybrid features. The motivation for the interaction of point-level and element-level information is their complementarity. Point-level information contains detailed local location knowledge, while element-level information provides overall shape and semantic knowledge. Therefore, the two-level information interaction can make full use of local information and global information to achieve mutual refinement and integration of information. Accordingly, if Figure 8 As shown, the point-element hybrid extractor 311 includes a point feature extractor 3111, an element feature extractor 3112 and a point-element fuser 3113.

[0071] As an example, the structure of the hybrid decoder 310 is as follows: Figure 4 and Figure 7 As shown, it includes L layers, each layer includes three modules: a point-element mixed extractor 311, a self-attention module 312 (a universal operation module), and a forward network 313 (a universal operation module). Each layer iteratively updates the mixed features. The updated mixed features can be input into three prediction heads, namely, a point prediction head (implemented with two linear layers), a category prediction head (implemented with two linear layers), and a mask prediction head (first passed through two linear layers, and the output result is then matrix multiplied with the bird's-eye view feature map to ensure that the size of the obtained mask information is consistent with the bird's-eye view feature map). Figure 1 ), generates coordinate information (such as point coordinates), element category information and mask information for each map element (the operations of the three prediction heads are independent of each other).

[0072] Next, the specific processing process of each cycle is introduced.

[0073] Optionally, the map information includes coordinate information of the coordinate point, and the steps of determining the second point feature map and the second element feature map respectively according to the bird's-eye view feature map, the first point feature map, the first element feature map and the current map information include: for each reference point, according to the bird's-eye view feature map, the first point feature of the reference point and the coordinate information of the reference point, determining the second point feature of the reference point, wherein the reference point includes the coordinate point corresponding to each first point feature; fusing the second point features of each reference point to obtain the second point feature map. For the determination of the second point feature map, how to sample the anchor point (i.e., the coordinate point as the target) and make the anchor point close to the map element to which it belongs is very important. By constructing the anchor point as a learnable coordinate point, randomly giving it at the beginning (i.e., as the above-mentioned reference point), and continuously updating it as a learnable parameter in repeated iterations, effective extraction of point features can be achieved. It should be understood that for the first cycle, the reference point is the coordinate point randomly determined for each map element when the hybrid feature map is initialized, and for subsequent cycles, the reference point is the coordinate point in the map element updated in the previous cycle.

[0074] Optionally, for each reference point, the step of determining the second point feature of the reference point according to the bird's-eye feature map, the first point feature of the reference point and the coordinate information of the reference point includes: for each reference point, according to the coordinate information and the first point feature of the reference point, determining a number of sampling points associated with the reference point from the map area; performing fusion processing based on the bird's-eye feature map and the coordinate information and weight of each sampling point associated with the reference point to obtain the third point feature of the reference point; and determining the second point feature of the reference point according to the first point feature and the third point feature of the reference point. By comprehensively calculating the fusion point feature for each reference point in combination with a number of surrounding sampling points, and then superimposing it on the reference point feature, the local output point feature of the corresponding reference point is obtained, and then the overall output point feature is spliced ​​out, which can realize the interaction between the reference point and the surrounding sampling points, and help to obtain reliable output point features.

[0075] Optionally, for each reference point, according to the coordinate information of the reference point and the first point feature, the step of determining a number of sampling points associated with the reference point from the map area includes: for each reference point, according to the coordinate information of the reference point and the first point feature, determining a fourth point feature of the reference point, wherein the fourth point feature is used to represent the point feature after considering the influence of the coordinate information; according to the fourth point feature of the reference point, determining a sampling offset and a weight of a number of sampling points associated with the reference point, wherein the sampling offset is used to represent the degree of positional offset of the sampling point relative to the reference point; according to the coordinate information of the reference point and the sampling offset of each sampling point, determining the coordinate information of each sampling point of the reference point. By combining the coordinate information of the reference point and the first point feature to determine the fourth point feature, and then determining the sampling offset and the weight of the sampling point, the sampling point associated with the reference point can be obtained, and the sampling point can be reliably determined.

[0076] Optionally, the step of determining the fourth point feature of the reference point according to the coordinate information of the reference point and the first point feature includes: encoding the coordinate information of the reference point to obtain a position code of the reference point; and determining the fourth point feature of the reference point according to the first point feature and the position code of the reference point. By superimposing the position code of the reference point on the basis of the first point feature, the coordinate information of the reference point can be further integrated to improve the feature characterization capability.

[0077] Optionally, based on the coordinate information and weight of each sampling point associated with the bird's-eye view feature map and the reference point, a fusion process is performed to obtain the third point feature of the reference point, including: determining the sampling feature of each sampling point corresponding to the reference point according to the coordinate information of each sampling point associated with the bird's-eye view feature map and the reference point; based on the weight of each sampling point associated with the reference point, a fusion process is performed on the sampling feature of each sampling point corresponding to the reference point to obtain the third point feature of the reference point. By firstly allowing the sampling point to interact with the bird's-eye view feature map to obtain the sampling feature, and then performing a fusion process on the sampling features of each sampling point, such as weighted summation, the third point feature of each reference point can be reliably calculated, and a fusion process is implemented, which helps to realize the calculation of the second point feature.

[0078] In summary, the point feature extractor 3111 is used to obtain the second point feature map The specific process can be expressed by the following formula (2,3,4). First, use formula (2) to generate the fourth feature

[0079]

[0080] in is the point coordinate output by the previous layer and is used as a reference point in the current layer. The subscript j represents a reference point, i.e., j∈{1,...,E×P}, so is a two-dimensional point, the point feature output by the previous layer The first feature point of the current layer is a C-dimensional vector. is a learnable parameter of a linear layer. is the position code of the reference point.

[0081] Then each reference point will sample K points, and the sampling offset of all sampling points will be generated using formula (3) and weights

[0082]

[0083] in These are all learnable parameters of the linear layer, and the softmax operation is performed on the dimension of the sampling points.

[0084] Finally, use formula (4) to get the second feature

[0085]

[0086] Where W v ∈R C×C is a learnable parameter of a linear layer, is the transformed bird’s-eye view feature map, is a two-dimensional point representing the sampling offset of a sampling point (k is an index between 1 and K). is a weight ranging from 0 to 1, satisfying The normalization condition of . The sampling feature of each sampling point is expressed as is the feature of the point obtained by fusion of the features of K sampling points of a reference point (j is the index), It is the second point feature corresponding to a reference point. The second point features of all reference points are combined to form the second point feature map of the current layer. During the calculation process, since the point coordinates are floating point values, Bilinear interpolation is used when upsampling the image.

[0087] Optionally, the map information includes coordinate information of coordinate points, and the step of determining the second point feature map and the second element feature map respectively according to the bird's-eye view feature map, the first point feature map, the first element feature map and the current map information further includes: for each map element, determining the second element feature of the map element according to the bird's-eye view feature map, the first element feature of the map element and the coordinate information of each reference point in the map element; fusing the second element features of each map element to obtain the second element feature map. The coordinate information of a map element is directly related to the coordinate information of each coordinate point in the map element. By using the coordinate information of each reference point in the map element in the interactive update of the first element feature, the association between the coordinate point and the map element can be enhanced and utilized, thereby obtaining a more accurate output element feature.

[0088] Optionally, for each map element, according to the bird's-eye view feature map, the first element feature of the map element and the coordinate information of each reference point in the map element, the step of determining the second element feature of the map element includes: for each map element, encoding the coordinate information of each reference point in the map element to obtain the position code of each reference point; fusing the position code of each reference point in the map element to obtain the position code of the map element; according to the bird's-eye view feature map, the first element feature of the map element and the position code of the map element, using the mask attention module in the hybrid decoder to determine the second element feature of the map element, wherein the mask used by the mask attention module is obtained based on the mask information of each pixel, and the mask information is used to indicate the probability that the corresponding pixel belongs to the map element. The present application uses the mask attention module to extract the second element feature. By fusing the position codes of each reference point in the map element as the position code of the map element, the position code of the map element can be associated with the position code of the reference point, thereby enhancing and improving the correlation between the coordinate point and the map element. As mentioned above, the anchor point can be constructed as a learnable coordinate point, that is, as a learnable parameter. Similarly, the anchor mask can also be constructed as a learnable parameter, such as Figure 8 As shown, the initial value of the anchor mask is randomly given at the beginning.

[0089] In summary, the element feature extractor 3112 is used to obtain the second element feature map The specific process is expressed by the following formula (5,6). First, use formula (5) to generate the position-aware element feature and location-aware bird's-eye view feature maps

[0090]

[0091] in is the feature of the i-th map element (i is in the range of 1 to E, used to index a map element), It is the position code generated for the map element (directly using the position code of the reference point obtained previously) The weighted sum of the position codes of all reference points belonging to a map element may be obtained, for example, by calculating the average value). It is the position coding map corresponding to the bird's-eye view feature map, which can be obtained by position coding using existing technology and superimposed on the bird's-eye view feature map. Figure X By calculating the sum of the two, we can get the location-aware bird's-eye view feature map

[0092] Then use formula (6) to generate the second element feature map

[0093]

[0094] Among them, M l-1 ∈{0, 1} HW is a binary mask image, which is obtained by binarizing the mask information output by the l-1th layer (the binarization threshold is 0.5). Represents the extracted element features of a map element (i is the index). Formula (6) It is the local output element feature corresponding to a map element. The second element features of all map elements are combined to form the second element feature map of the current layer.

[0095] As an example, the fusion processing of the second point feature map and the second element feature map may specifically include two fusions of the output point feature and the output element feature, such as Figure 7 As shown, the first fusion is performed by the point-element mixed extractor 311. After the first fused features are obtained, they are input into the self-attention module 312 and the forward network 313. The first fused features are processed twice in succession, and the final fused features are used as the output mixed features of the current cycle.

[0096] Alternatively, if Figure 8As shown, the first fusion step includes: using the self-attention module in the hybrid decoder to process the second point feature map and the second element feature map respectively to obtain a fifth point feature map and a fifth element feature map; converting the fifth point feature map to the same dimension as the fifth element feature map, and fusing the fifth element feature map and the converted fifth point feature map to obtain a sixth element feature map; converting the fifth element feature map to the same dimension as the fifth point feature map, and fusing the fifth point feature map and the converted fifth element feature map to obtain a sixth point feature map; fusing the sixth point feature map and the sixth element feature map to obtain the updated hybrid feature map. By performing intra-level interaction on the second point feature map and the second element feature map respectively using the self-attention module (different from the self-attention module 312) of the point-element hybrid extractor 311, and then performing cross-level interaction through dimensional conversion and merging fusion, and encoding the interaction result into the updated hybrid feature, the second point feature map and the second element feature map can be fully fused to achieve the purpose of updating the hybrid feature map.

[0097] Specifically, the intra-level interaction performed by the self-attention module is implemented by formula (7):

[0098]

[0099] in and They are point-level interaction and element-level interaction, respectively, and are both implemented by the general self-attention module and forward network in the point-element hybrid extractor 311.

[0100] The cross-level interaction is implemented by formula (8):

[0101]

[0102] in It is to copy the information of the fifth element feature map P times, and then connect them together with the dimension of the fifth point feature map Consistency; It is the weighted sum of the fifth point feature maps of P reference points belonging to the same map element, and the dimension of the result and the fifth element feature map Consistent.

[0103] Formula (8) obtains the updated sixth feature map And the updated sixth element characteristic map Combining the two, we get the updated mixed feature map Specifically, as shown in formula (9):

[0104] Q h,l =[Qp,l , Q e,l ], (9)

[0105] return Figure 2 In step S240, a map corresponding to the data is constructed based on the map information. This step can construct a map corresponding to the data based on the final map information of step S230.

[0106] As mentioned above, at the end of each loop, the prediction head is used to obtain the map information corresponding to the updated hybrid feature map. This step can directly use the map information obtained in the last loop. As an example, the category prediction head can output the category information and confidence of each map element, and a confidence threshold can be configured. If the confidence corresponding to a map element is less than the confidence threshold, the map element can be discarded; similarly, the point prediction head can output the coordinate information and confidence of the coordinate point in each map element, and a confidence threshold can also be configured (which can be the same as or different from the confidence threshold of the element category information). If the confidence corresponding to a coordinate point is less than the confidence threshold, the coordinate point can be discarded, that is, it is not used as an anchor point.

[0107] In addition, the loss function used by the hybrid decoder during the training process includes a point-element consistency loss, which is used to represent the risk level that the point feature map and the element feature map in the updated hybrid feature map are inconsistent with each other. It should be understood that the risk level means describing the size of the risk, and the point-element consistency loss can be a level, a probability, or other reasonable value forms, which is not limited in this application. Considering the original difference between point-level features and element-level features, they focus on local and overall information respectively, and the learning of the two-level features may also interfere with each other, which will increase the difficulty of information interaction and reduce the effectiveness of information interaction. By introducing the point-element consistency loss, point-element consistency constraints can be implemented to enhance the consistency between the point-level and element-level information of each map element, and the distinguishability of the map elements can also be enhanced, which helps to reduce the entanglement between different map elements, thereby improving the accuracy of map construction.

[0108] Optionally, the value of the point-element consistency loss is determined by the following method: transforming the point feature map and the element feature map in the updated hybrid feature map respectively to obtain point-level information and element-level information; fusing the information of the coordinate points belonging to the same map element in the point-level information to obtain pseudo-element level information; determining the value of the point-element consistency loss based on the pseudo-element level information and the element-level information to indicate the degree of risk that the pseudo-element level information and the element-level information are inconsistent with each other. By fusing the point-level information according to the map element to which it belongs, pseudo-element level information that is consistent with the dimension of the element-level information and can reflect the point-level information can be obtained, and then by comparing the two, the value of the point-element consistency loss is determined, thereby achieving reliable calculation of the point-element consistency loss. As an example, when determining the pseudo-element level information, all coordinate points of the same map element or part of the coordinate points can be used, and this application does not impose any restrictions on this.

[0109] As an example, the point-element consistency constraint is defined on the intermediate results of the point prediction head and the mask prediction head. The input data includes the point-level features decomposed from the mixed features. and element-level features The process diagram is as follows Fig.10 After extracting the point-level features and element-level features from the mixed features, first use formula (10) to transform the two:

[0110]

[0111] Among them, Wp and Wm are learnable parameters of the linear layer. and It is the transformed point-level information and element-level information.

[0112] Then, we put The point-level information of all coordinate points belonging to the same map element in the weighted summation is obtained to obtain a pseudo element-level representation. Fig. 9 As shown, the element similarity matrix is ​​calculated using formula (11)

[0113]

[0114] A binary cross entropy loss is applied between the calculated similarity matrix and the binary GT (Ground truth) correspondence matrix, where 1 is assigned to the diagonal entries corresponding to the same element and 0 is assigned to the other elements. By promoting a high similarity between pseudo-element-level information and element-level information, the consistency between point-level information and element-level information is enhanced, which increases the consistency between point-level features and element-level features in the output hybrid features.

[0115] Optionally, the loss function used by the hybrid decoder during training also includes at least one of the following: classification loss (for supervising the category prediction head, the Focal loss function can be used), point regression loss (for supervising the point prediction head, using the L1 loss function), point direction loss (for supervising the point prediction head, using the L1 loss function), mask loss (for supervising the mask prediction head, using the binary cross entropy function and the dice function). By configuring the above loss functions, a reference can be provided for the training of the corresponding structure in the hybrid decoder. The weights of each loss function can be configured as needed.

[0116] like Fig.11 and Fig.12 As shown, the existing technology usually lacks information interaction at two levels, which easily leads to incomplete element-level shapes or inaccurate point-level positions. This application obtains more complete shapes and more accurate positions through hybrid representation and interaction of point-level and element-level, generates richer details and more accurate map element shapes, and reduces entanglement between map elements.

[0117] An embodiment of the present disclosure also provides an electronic device, which includes a processor and, optionally, may also include at least one transceiver and / or at least one memory coupled to the at least one processor, wherein the at least one processor is configured to execute the steps of the method provided in any optional embodiment of the present disclosure.

[0118] Fig.13 A schematic diagram of the structure of an electronic device applicable to an embodiment of the present invention is shown in FIG. Fig.13 As shown, Fig.13The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, each of the processor 4001, the memory 4003 and the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure. Optionally, the electronic device may be a first network node, a second network node or a third network node.

[0119] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0120] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.13 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0121] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.

[0122] The memory 4003 is used to store computer programs or executable instructions for executing the embodiments of the present disclosure, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer programs or executable instructions stored in the memory 4003 to implement the steps shown in the above method embodiments.

[0123] Through the above method proposed in the present application, interactive hybrid learning of point-level and element-level information can be performed. Specifically, the present application includes a hybrid feature and a simple and effective hybrid framework. Among them, the hybrid feature is a set of learnable parameters that represent all map elements in the map. It is iteratively updated and refined by interacting with the bird's-eye view feature map. During the iteration process, both the point-level information and the element-level information of the map elements are integrated and encoded into the hybrid feature map. Each hybrid feature in the hybrid feature map corresponds to a separate map element, so it can be directly converted into the coordinate information (such as point coordinates), element category information, and mask information of the corresponding map element. The difference between this method and the existing methods in the core idea is as follows Fig.11 As shown. Through the hybrid representation and interaction of point-level and element-level, the map elements obtained by the method of the present application have more complete shapes and more accurate positions, and the accuracy is much higher than that of existing methods. The present application also helps to achieve consistency between the two levels of information by introducing point-element consistency constraints, reducing the entanglement between map elements.

[0124] An embodiment of the present disclosure provides a computer-readable storage medium having a computer program or instructions stored thereon. When the computer program or instructions are executed by at least one processor, the steps and corresponding contents of the aforementioned method embodiment can be executed or implemented.

[0125] The embodiments of the present disclosure also provide a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiments when executed by a processor.

[0126] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than that shown or described in the text.

[0127] It should be understood that, although the flowchart of the embodiment of the present disclosure indicates each operation step by arrows, the implementation order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiment of the present disclosure, the implementation steps in each flowchart can be executed in other orders according to demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times. In scenarios with different execution times, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present disclosure does not limit this.

[0128] The above text and drawings are provided only as examples to help readers understand the present disclosure. They are not intended and should not be interpreted as limiting the scope of the present disclosure in any way. Although certain embodiments and examples have been provided, based on the contents disclosed herein, it is obvious to those skilled in the art that the embodiments and examples shown can be changed without departing from the scope of the present disclosure, and other similar implementation means based on the technical ideas of the present disclosure are adopted, which also fall within the protection scope of the embodiments of the present disclosure.

Claims

1. A map construction method, comprising: Get the data to be processed; Extracting a bird's-eye view feature map based on the data; Determine map information based on the hybrid decoder according to the bird's-eye view feature map and the hybrid feature map, wherein the map information includes coordinate information of at least one coordinate point in each map element and category information to which each map element belongs; Based on the map information, construct a map corresponding to the data; The map includes a plurality of map elements, each map element includes an area formed by a plurality of coordinate points within the map area, the mixed feature map includes a plurality of mixed features, each mixed feature corresponds to a map element, each mixed feature includes a point feature and an element feature, the point feature is used to describe information related to each coordinate point in the corresponding map element, and the element feature is used to describe information related to the corresponding map element.

2. The method according to claim 1, wherein: According to the bird's-eye view feature map and the hybrid feature map, based on the hybrid decoder, the step of determining map information comprises: Decomposing the mixed feature map into a first point feature map and a first element feature map, wherein the first point feature map includes first point features corresponding to each coordinate point in each map element, and the first element feature map includes first element features corresponding to each map element; Determine a second point feature map and a second element feature map respectively according to the bird's-eye view feature map, the first point feature map, the first element feature map and current map information; Performing a fusion process on the second point feature map and the second element feature map to update a mixed feature map; Update the map information according to the bird's-eye view feature map and the updated mixed feature map, and return to the operation of decomposing the mixed feature map into the first point feature map and the first element feature map to perform the next update; Based on the map information, the step of constructing a map corresponding to the data includes: Based on the final map information, a map corresponding to the data is constructed.

3. The method according to claim 2, wherein: The map information includes coordinate information of the coordinate points, and the steps of respectively determining a second point feature map and a second element feature map according to the bird's-eye view feature map, the first point feature map, the first element feature map and the current map information include: For each reference point, determine the second point feature of the reference point according to the bird's-eye view feature map, the first point feature of the reference point and the coordinate information of the reference point, wherein the reference point includes the coordinate point corresponding to each first point feature; fuse the second point features of each reference point to obtain the second point feature map; For each map element, the second element feature of the map element is determined according to the bird's-eye view feature map, the first element feature of the map element and the coordinate information of each reference point in the map element; the second element features of each map element are fused to obtain the second element feature map.

4. The method of claim 3, wherein: For each reference point, the step of determining the second point feature of the reference point according to the bird's-eye view feature map, the first point feature of the reference point and the coordinate information of the reference point includes: For each reference point, determining a number of sampling points associated with the reference point from the map area according to the coordinate information of the reference point and the first point feature; Based on the bird's-eye view feature map and the coordinate information and weight of each sampling point associated with the reference point, a fusion process is performed to obtain a third point feature of the reference point; A second point feature of the reference point is determined according to the first point feature and the third point feature of the reference point.

5. The method of claim 4, wherein: For each reference point, the step of determining a number of sampling points associated with the reference point from a map area according to the coordinate information of the reference point and the first point feature includes: For each reference point, determine a fourth point feature of the reference point according to the coordinate information of the reference point and the first point feature, wherein the fourth point feature is used to represent the point feature after considering the influence of the coordinate information; Determine, according to the fourth point feature of the reference point, sampling offsets and weights of a plurality of sampling points associated with the reference point, wherein the sampling offsets are used to indicate a degree of positional offset of the sampling points relative to the reference point; The coordinate information of each sampling point of the reference point is determined according to the coordinate information of the reference point and the sampling offset of each sampling point.

6. The method according to claim 5, wherein: The step of determining the fourth point feature of the reference point according to the coordinate information of the reference point and the first point feature comprises: Encoding the coordinate information of the reference point to obtain a position code of the reference point; A fourth point feature of the reference point is determined according to the first point feature and the position code of the reference point.

7. The method according to claim 4, wherein: The step of performing fusion processing based on the bird's-eye view feature map and the coordinate information and weight of each sampling point associated with the reference point to obtain the third point feature of the reference point includes: Determine, according to the bird's-eye view feature map and the coordinate information of each sampling point associated with the reference point, a sampling feature of each sampling point corresponding to the reference point; Based on the weight of each sampling point associated with the reference point, the sampling features of the reference point corresponding to each sampling point are fused to obtain the third point feature of the reference point.

8. The method according to claim 3, wherein: For each map element, the step of determining the second element feature of the map element according to the bird's-eye view feature map, the first element feature of the map element and the coordinate information of each reference point in the map element comprises: For each map element, encoding the coordinate information of each reference point in the map element to obtain a position code of each reference point; fusing the position codes of the reference points in the map element to obtain the position code of the map element; According to the bird's-eye view feature map, the first element feature of the map element and the position encoding of the map element, the second element feature of the map element is determined using the mask attention module in the hybrid decoder, wherein the mask used by the mask attention module is obtained based on the mask information of each pixel, and the mask information is used to represent the probability that the corresponding pixel belongs to the map element.

9. The method according to claim 2, wherein: The step of fusing the second point feature map and the second element feature map to update the mixed feature map includes: Using the self-attention module in the hybrid decoder to process the second point feature map and the second element feature map respectively, to obtain a fifth point feature map and a fifth element feature map; Converting the fifth point feature map to the same dimension as the fifth element feature map, and fusing the fifth element feature map and the converted fifth point feature map to obtain a sixth element feature map; Converting the fifth element feature map to the same dimension as the fifth point feature map, and fusing the fifth point feature map and the converted fifth element feature map to obtain a sixth point feature map; The sixth point feature map and the sixth element feature map are fused to obtain the updated mixed feature map.

10. The method according to any one of claims 1 to 9, wherein: The loss function used by the hybrid decoder during training includes a point-element consistency loss, which is used to represent the risk degree that the point feature map and the element feature map in the updated hybrid feature map are inconsistent with each other.

11. The method according to claim 10, wherein: The value of the point-element consistency loss is determined by the following method: Transform the point feature map and the element feature map in the updated mixed feature map to obtain point level information and element level information respectively; Performing fusion processing on the information of coordinate points belonging to the same map element in the point-level information to obtain pseudo-element-level information; A value of the point-element consistency loss is determined according to the pseudo-element level information and the element level information to indicate a risk level that the pseudo-element level information and the element level information are inconsistent with each other.

12. The method of claim 10, wherein: The loss function used by the hybrid decoder during training also includes at least one of the following: semantic segmentation loss, classification loss, point regression loss, point direction loss, and mask loss.

13. An electronic device, comprising: at least one processor; as well as at least one memory storing computer executable instructions, Wherein, when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to perform the method according to any one of claims 1 to 12.

14. A computer-readable storage medium storing instructions, wherein: When the instructions are executed by at least one processor, the at least one processor is prompted to perform the method according to any one of claims 1 to 12.