Method executed by electronic equipment, electronic equipment and storage medium

By using an AI network based on the Transformer architecture to extract and fuse features from camera images and radar point cloud data, the problem of insufficient accuracy and robustness in multi-data type map construction in existing technologies is solved, realizing the construction of high-precision maps and sensor adaptability in autonomous driving.

CN120997625APending Publication Date: 2025-11-21BEIJING SAMSUNG TELECOM R&D CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410628033.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies have room for improvement in using AI to build maps of the environment around vehicles, especially in terms of accuracy and robustness when dealing with multiple data types.

Method used

An AI network based on an encoder and decoder Transformer architecture is used to construct high-precision map images by extracting and fusing features from different types of data, including camera images and radar point cloud data. Features are enhanced through a mapping network to adapt to the absence or damage of different sensors.

Benefits of technology

It improves prediction accuracy and network robustness under different data types and mixed data scenarios, ensuring that high-precision maps can be built in autonomous driving and adapt to sensor changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997625A_ABST
    Figure CN120997625A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method executed by electronic equipment, the electronic equipment and a storage medium, and relates to the field of artificial intelligence. The method comprises the following steps: acquiring first data, wherein the first data comprises at least one type of data; based on the first data, using a first AI network to obtain a map image corresponding to the first data; under the condition that the first data comprises one type of data, obtaining a first feature of the first data by using an encoder corresponding to the type based on the first data, and obtaining the map image based on the first feature; and under the condition that the first data comprises at least two types of data, for each type of data, using an encoder corresponding to the type to extract a second feature of the type of data, fusing the second features of the types of data to obtain a first feature of the first data, and obtaining the map image. Optionally, the method executed by the electronic equipment can be executed by using an artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and relates to a method executed by an electronic device, an electronic device and a storage medium. BACKGROUND

[0002] In the field of automatic driving, data in the environment around a vehicle can be collected during driving of the vehicle, so as to construct a map of the environment around the vehicle by using the collected data; this process can be implemented by using an AI technology.

[0003] However, in the field, there is still a large room for improvement in how to better construct a map by using an AI technology. SUMMARY

[0004] The present application provides a method executed by an electronic device, an electronic device and a storage medium. The technical solution is as follows:

[0005] In one aspect, the present application provides a method executed by an electronic device, the method comprising:

[0006] obtaining first data, the first data comprising at least one type of data;

[0007] using a first artificial intelligence (AI) network to obtain a map image corresponding to the first data based on the first data;

[0008] In one aspect, the present application provides a method executed by an electronic device, the method comprising:

[0009] In the case where the first data comprises one type of data, using an encoder corresponding to the type to obtain a first feature of the first data based on the first data, and obtaining the map image based on the first feature of the first data;

[0010] In the case where the first data comprises at least two types of data, for each type of data, using an encoder corresponding to the type to extract a second feature of the type of data, and fusing the second features of the types of data to obtain a first feature of the first data, and obtaining the map image based on the first feature of the first data.

[0011] In one possible implementation, using a first AI network to obtain a map image corresponding to the first data based on the first data comprises:

[0012] determining the type of data contained in the first data;

[0013] In a case where the first data comprises one type of data, based on the first data, a first feature of the first data is obtained using an encoder corresponding to the determined type, and based on the first feature of the first data, the map image is obtained.

[0014] In a case where the first data comprises at least two types of data, for each determined type of data, a second feature of the determined type of data is extracted using an encoder corresponding to the determined type, and the second features of the determined types of data are fused to obtain a first feature of the first data, and based on the first feature of the first data, the map image is obtained.

[0015] In a possible implementation, the obtaining, based on the first feature of the first data, of the map image comprises:

[0016] Based on the first feature of the first data, a third feature corresponding to the first data is obtained by using a mapping network in the first AI network to enhance the first feature.

[0017] Based on the third feature corresponding to the first data, a map image corresponding to the first data is obtained by using a decoder in the first AI network.

[0018] In a possible implementation, the obtaining, based on the first feature of the first data, of the third feature corresponding to the first data by using the mapping network in the first AI network to enhance the first feature comprises:

[0019] Based on the first feature of the first data, a third feature corresponding to the first data is obtained by using a first mapping network in the first AI network to enhance the first feature; or

[0020] Based on the first feature of the first data, a third feature corresponding to the first data is obtained by using a second mapping network corresponding to the first feature in the first AI network to enhance the first feature.

[0021] In a possible implementation, the obtaining, based on the third feature corresponding to the first data, of the map image corresponding to the first data by using the decoder in the first AI network comprises:

[0022] If the first mapping network is used, the map image corresponding to the first data is obtained by using the decoder in the first AI network based on the third feature corresponding to the first data and the first feature.

[0023] In a possible implementation, the first data comprises image data collected by a camera and / or point cloud data collected by a radar.

[0024] In a possible implementation, the first feature is a BEV feature, and the second feature is a BEV feature.

[0025] In another aspect, the present application provides a method performed by an electronic device, the method comprising:

[0026] obtaining a training data set, the training data set comprising at least a plurality of first samples and a plurality of second samples associated with the first samples, the first samples and the second samples being different types of data;

[0027] based on the training data set, using a second AI network to obtain, respectively, fourth features of the first samples, fifth features of the second samples, and sixth features of the first samples and the second samples associated with the first samples;

[0028] based on the fourth features of the first samples, the fifth features of the second samples, and the sixth features of the first samples and the second samples associated with the first samples, respectively, using the second AI network to make predictions to obtain prediction results corresponding to each sample in the training data set;

[0029] wherein the prediction results corresponding to each sample include first images corresponding to the first samples, second images corresponding to the second samples, and third images corresponding to the first samples and the associated second samples;

[0030] based on the prediction results corresponding to each sample in each training data set, training the second AI network to obtain a first AI network.

[0031] In a possible implementation, based on the training data set, using a second AI network to obtain, respectively, fourth features of the first samples, fifth features of the second samples, and sixth features of the first samples and the second samples associated with the first samples, comprises:

[0032] based on each first sample, using an encoder corresponding to the type of the first sample to obtain the fourth feature of the first sample;

[0033] based on each second sample, using an encoder corresponding to the type of the second sample to obtain the fifth feature of the second sample;

[0034] for each first sample, performing fusion processing on the fourth feature of the first sample and the fifth feature of the associated second sample to obtain the sixth feature of the first sample and the second sample associated with the first sample.

[0035] In a possible implementation, the prediction using the second AI network based on the fourth feature of each first sample, the fifth feature of each second sample, and the sixth feature of each first sample and the second sample associated with the first sample obtains a prediction result corresponding to each sample in the training data set, including:

[0036] The first mapping network in the second AI network is used to enhance the fourth feature, the fifth feature, and the sixth feature based on the fourth feature of each first sample, the fifth feature of each second sample, and the sixth feature of each first sample and the second sample associated with the first sample, to obtain enhanced fourth feature, fifth feature, and sixth feature.

[0037] The decoder in the second AI network is used to obtain the prediction result corresponding to each sample based on the enhanced fourth feature, fifth feature, and sixth feature.

[0038] In a possible implementation, the training of the second AI network based on the prediction result corresponding to each sample in each training data set obtains the first AI network, including:

[0039] The training loss corresponding to each set of associated samples is determined based on the prediction result and the sample label corresponding to at least one set of associated samples, and the prediction result corresponding to each set of associated samples includes a first image corresponding to a first sample, a second image corresponding to an associated second sample, and a third image corresponding to the first sample and the associated second sample.

[0040] The second AI network is trained based on the training loss corresponding to each set of associated samples.

[0041] In a possible implementation, the first sample is image data collected by a camera, and the second sample is point cloud data collected by a radar.

[0042] In another aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the method described above.

[0043] In another aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method described above.

[0044] The technical scheme provided by the embodiments of the present application has the following beneficial effects:

[0045] The method provided in this application involves acquiring first data, which includes at least one type of data; and using a first AI network to obtain a map image corresponding to the first data. Specifically, when the first data includes one type of data, an encoder corresponding to that type is used to obtain a first feature of the first data, and the map image is obtained based on the first feature. When the first data includes at least two types of data, for each type of data, an encoder corresponding to that type is used to extract a second feature of that type of data, and the second features of each type of data are fused to obtain a first feature of the first data, and the map image is obtained based on the first feature. Therefore, by utilizing the first AI network, both prediction scenarios involving a single type of data and prediction scenarios involving mixed data containing at least two types can be handled with high accuracy, thus improving the robustness of the network. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0047] Figure 1a A flowchart illustrating a method performed by an electronic device, provided as an embodiment of this application;

[0048] Figure 1b A flowchart illustrating a method performed by an electronic device, provided as an embodiment of this application;

[0049] Figure 2 A schematic diagram of a network architecture provided for an embodiment of this application;

[0050] Figure 3 This is another network architecture diagram provided in the embodiments of this application;

[0051] Figure 4 A schematic diagram of a mapping network and decoder provided in an embodiment of this application;

[0052] Figure 5 This is another network architecture diagram provided in the embodiments of this application;

[0053] Figure 6 A schematic diagram illustrating a robustness test of an existing method provided in an embodiment of this application;

[0054] Figure 7 A schematic diagram illustrating a comparative test between the method of this application and a prior art method, provided as an embodiment of this application;

[0055] Figure 8A structural schematic diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0056] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of various embodiments of the present disclosure as defined by the claims and their equivalents. The description includes various specific details to assist in that understanding but these are to be regarded as merely exemplary. Accordingly, those of ordinary skill in the art will recognize that various changes and modifications of the various embodiments described herein can be made without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and constructions can be omitted for clarity and conciseness.

[0057] The terms and words used in the following description and claims are not limited to the bibliographical meanings, but, are merely used to enable a clear and consistent understanding of the present disclosure. Accordingly, it should be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustration purpose only and not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.

[0058] It should be understood that the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a component surface” includes reference to one or more of such surfaces. When we refer to an element being “connected” or “coupled” to another element, we mean that the one element can be directly connected or coupled to the other element or intervening elements can be present. In addition, the use of “connected” or “coupled” herein also includes wireless connection or wireless coupling.

[0059] The terms “include” or “may include” refer to the presence of a corresponding disclosed function, operation, or component in various embodiments of the present disclosure, and are not limited to the presence of one or more additional functions, operations, or features. In addition, the terms “include” or “have” can be interpreted to mean the presence of certain characteristics, numbers, steps, operations, constituent elements, components, or combinations thereof, but should not be interpreted as excluding the possibility of the presence of one or more other characteristics, numbers, steps, operations, constituent elements, components, or combinations thereof.

[0060] The term "or" used in the various embodiments of the disclosure includes any of the listed terms and all combinations of the listed terms. For example, "A or B" can include A, can include B, or can include both A and B. When describing a plurality of (two or more) items, if the relationship between the plurality of items is not explicitly limited, the plurality of items can refer to one, a plurality, or all of the plurality of items, for example, for a description of "a parameter A includes A1, A2, A3", it can be implemented that the parameter A includes A1 or A2 or A3, and it can also be implemented that the parameter A includes at least two of the three items of the parameter A1, A2, A3.

[0061] Unless defined differently, all terms (including technical or scientific terms) used in the disclosure have the same meaning as understood by a person skilled in the art to which the disclosure pertains. Common terms as defined in a dictionary are interpreted to have meanings consistent with those in the context of the relevant technical field, and should not be ideally or excessively formalized, unless explicitly defined in the disclosure.

[0062] At least part of the functions of the apparatus or electronic device provided in the embodiments of the disclosure can be implemented by an AI model, for example, at least one of the plurality of modules of the apparatus or electronic device can be implemented by an AI model. The functions associated with AI can be performed by a non-volatile memory, a volatile memory, and a processor.

[0063] The processor can include one or more processors. At this time, the one or more processors can be a general-purpose processor such as a central processing unit (CPU), an application processor (AP), or the like, or a pure graphics processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI dedicated processor such as a neural processing unit (NPU).

[0064] The one or more processors control the processing of input data according to a predefined operation rule or an artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operation rule or the artificial intelligence model is provided by training or learning.

[0065] Here, the provision by learning means that a predefined operation rule or an AI model having a desired characteristic is obtained by applying a learning algorithm to a plurality of learning data. The learning can be performed in the apparatus or electronic device itself according to the embodiments, and / or can be implemented by a separate server / system.

[0066] An AI model can include multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network computation through a calculation between input data of the layer (e.g., a result of computation of a previous layer and / or input data of the AI model) and the multiple weight values of the current layer. Examples of neural networks include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial network (GAN), and a deep Q-network.

[0067] A learning algorithm is a method of training a predetermined target device (e.g., a robot) using multiple learning data to enable, allow, or control the target device to determine or predict. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0068] The method provided by the disclosure can relate to one or more of the technical fields of voice, language, image, video, or data intelligence.

[0069] Optionally, in relation to the field of voice or language, according to the disclosure, in the method performed by the electronic device, a voice signal can be received as an analog signal via the electronic device (e.g., a microphone), and a voice portion can be converted into computer-readable text using an automatic speech recognition (ASR) model. An utterance intent of a user can be obtained by interpreting the converted text using a natural language understanding (NLU) model. The ASR model or the NLU model can be an artificial intelligence model. The artificial intelligence model can be processed by an artificial intelligence dedicated processor designed in a hardware structure designated for artificial intelligence processing. The artificial intelligence model can be obtained through training. Here, "obtained through training" means that a pre-defined operation rule or an artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training an algorithm with multiple training data to train a basic artificial intelligence model. Language understanding is a technology for recognizing and applying / handling human language / text, including, for example, natural language processing, machine translation, a dialogue system, question answering, or speech recognition / synthesis.

[0070] Optionally, in the field related to images or videos, according to the present disclosure, in the method performed by the electronic device performed in the electronic device, the method can obtain the repaired audio frame by using the audio data as the input data of the artificial intelligence model. The artificial intelligence model can be obtained by training. Here, "obtained by training" means that a basic artificial intelligence model is trained with a plurality of training data by training an algorithm to obtain a predefined operation rule or an artificial intelligence model configured to perform a desired feature (or purpose). The method of the present disclosure can relate to the field of visual understanding of artificial intelligence technology, which is a technology for recognizing and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / positioning, or image enhancement.

[0071] Optionally, in the field related to data intelligent processing, according to the present disclosure, in the method performed by the electronic device performed in the electronic device, the method can be performed by using the audio data using the artificial intelligence model. The processor of the electronic device can perform a preprocessing operation on the data to convert into a form suitable for use as input to the artificial intelligence model. The artificial intelligence model can be obtained by training. Here, "obtained by training" means that a basic artificial intelligence model is trained with a plurality of training data by training an algorithm to obtain a predefined operation rule or an artificial intelligence model configured to perform a desired feature (or purpose). Inference prediction is a technology for logically inferring and predicting by determining information, including, for example, knowledge-based inference, optimization prediction, preference-based planning, or recommendation.

[0072] The technical solutions of the embodiments of the present disclosure and the technical effects generated by the technical solutions of the present disclosure will be described below through the description of several optional embodiments. It should be pointed out that the following embodiments can be mutually referenced, borrowed or combined. For the same terms, similar features and similar implementation steps in different embodiments, they will not be described repeatedly.

[0073] Figure 1a A flowchart of a method performed by an electronic device is provided for the present application, which can be a server, a cloud computing center device, or a terminal, etc. As shown in Figure 1a The method includes steps 101-102.

[0074] Step 101, the electronic device acquires first data, the first data including at least one type of data.

[0075] Exemplarily, the first data can include only one type of data; alternatively, the first data can include at least two types of data.

[0076] The first data includes at least one of a first type of data or a second type of data. The type of data can represent a characteristic of the data, such as a source of the data, a form of the data, and the like. Different types of data correspond to different characteristics; for example, at least one type of data can include, but is not limited to, image data collected by a camera, point cloud data collected by a laser radar, point cloud data collected by a millimeter wave radar, and the like.

[0077] It should be noted that the type of data can also be referred to as the modality of the data, that is, the first data can be single-modality data including only one modality; or the first data can also be mixed data including at least two modalities.

[0078] In a possible implementation, the first data includes image data collected by a camera and / or point cloud data collected by a radar. That is, the first data can include only image data collected by a camera; or can include only point cloud data collected by a radar, or can include both image data collected by a camera and point cloud data collected by a radar.

[0079] In a possible scenario, in an autonomous driving scenario, the first modality of data is an image of a vehicle environment collected by a camera, and the second modality of data is point cloud data of the vehicle environment collected by a radar. For example, the first data includes camera images of 6 directions around the vehicle collected by the camera at the same time, and point cloud data around the vehicle collected by a laser radar.

[0080] In a possible manner, the above only takes two types as an example, and of course, the first data can include at least one type of data of three or more types, which is not limited by the embodiments of the present application.

[0081] For example, if the first data is single-type data, the first data can include any one of a first type of data, a second type of data, and a third type of data. If the first data is mixed multi-type data, the first data can include any two or three of the first type of data, the second type of data, and the third type of data.

[0082] In step 102, based on the first data, a first artificial intelligence (AI) network is used to obtain a map image corresponding to the first data.

[0083] In step 102, based on the first data, a first artificial intelligence (AI) network is used to obtain a map image corresponding to the first data.

[0084] In a case where the first data comprises one type of data, a first feature of the first data is obtained based on the first data using an encoder corresponding to the type, and the map image is obtained based on the first feature of the first data.

[0085] In a case where the first data comprises at least two types of data, for each type of data, a second feature of the type of data is extracted using an encoder corresponding to the type, and the second features of the types of data are fused to obtain the first feature of the first data, and the map image is obtained based on the first feature of the first data.

[0086] For example, the first AI comprises an encoder corresponding to each type.

[0087] In a case where the first data comprises only one type of data, for example, the first data is single-modal data comprising only one modality. If the first data comprises only image data collected by a camera, a corresponding encoder of the camera can be used to extract a first feature of the image data. Alternatively, if the first data comprises only point cloud data collected by a radar, a corresponding encoder of the radar can be used to extract a first feature of the point cloud data.

[0088] In a case where the first data comprises at least two types of data, for example, the first data is hybrid-modal data comprising at least two modalities. First, a second feature of data of each type is extracted using a corresponding encoder of each type, for example, a second feature of image data in the first data is extracted using a corresponding encoder of a camera, and a second feature of point cloud data is extracted using a corresponding encoder of a radar. Then, the second features of the data of each type are fused, for example, a fusion network can be used to fuse the second feature of the image data and the second feature of the point cloud data to obtain the first feature of the first data.

[0089] The second feature of the image data, the second feature of the point cloud data, and the first feature of the first data have the same feature dimension.

[0090] In a possible implementation manner, the implementation of step 102 can comprise the following steps A1-A3:

[0091] Step A1, determining the types of data included in the first data;

[0092] In a possible manner, the electronic device can input the first data into the first AI network, and use the first AI network to determine at least one type corresponding to the first data to obtain the data of each type included in the first data.

[0093] In yet another possible implementation manner, the electronic device can input the first data and type indication information of the first data into the first AI network, and determine each type of data included in the first data based on the type indication information using the first AI network. The type indication information is used to indicate at least one type to which the first data corresponds.

[0094] Step A2, in a case where the first data includes one type of data, based on the first data, a first feature of the first data is obtained using an encoder corresponding to the determined type;

[0095] Step A3, in a case where the first data includes at least two types of data, for each determined type of data, a second feature of the determined type of data is extracted using an encoder corresponding to the determined type, and the second features of the determined types of data are fused to obtain the first feature of the first data.

[0096] In a possible implementation manner, the first feature is a BEV feature, and the second feature is a BEV feature.

[0097] The first AI network can be an AI network based on a MapTR (Map Transformer, a Transformer architecture based on an encoder-decoder) network structure. As shown in Figure 2 The encoder in the first AI network can be a BEV (Birds Eye View, bird's eye view) feature encoder. The encoder corresponding to the camera image in the first AI network supports taking multi-view images as input and converting the features of the camera image to a BEV feature space while retaining geometric and semantic information. The encoder corresponding to the radar point cloud data in the first AI network supports converting LiDAR features to a BEV feature space. Of course, other encoders can also be used, and the present application only takes MapTR as an example for description.

[0098] For camera images, an encoder corresponding to the image can be used, for example Figure 2 In the 2D image encoding module in FIG. 2, a 2D-to-3D transformation is used to encode pixel-level semantic features in image perspectives. Resnet50 (Residual Network) can be used as a backbone to extract multi-view features, and GKT (Graph-based Knowledge Tracing) is used as a 2D-to-BEV feature conversion module to convert the multi-view features to a BEV space.

[0099] For LiDAR point cloud data, initial features can be extracted using an encoder corresponding to the point cloud data, for example Figure 2 The 3D point cloud encoding module in the middle. Specifically, the SECOND model can be followed, and voxelization and sparse LiDAR encoder are used. LiDAR features are projected to BEV feature space using the flatten operation in the BEVFusion model, realizing the conversion from 3D view to BEV view to obtain a unified BEV feature representation, i.e., the initial features of the point cloud data of the radar, that is, the BEV features of the point cloud data.

[0100] For the case where the first data includes at least two types of data, in order to effectively fuse the BEV features from the camera image and the radar point cloud data, a convolution-based fusion method can be used to fuse the BEV features of the image and the BEV features of the point cloud data. Specifically, concatenation can be performed in the feature channel C dimension, and then convolution is used to fuse the BEV features from each modality to obtain the first feature, and the channel numbers of the first feature and the second feature are the same.

[0101] In a possible implementation, in step 102 or step A2-A3, based on the first feature of the first data, the map image is obtained, including the following steps B1 and B2:

[0102] Step B1, based on the first feature of the first data, using the mapping network in the first AI network, enhancing the first feature to obtain the third feature corresponding to the first data;

[0103] In this step, the electronic device can further map the first feature of the first data to keep the mapped features of the first data aligned in the feature space corresponding to each type. The feature dimensions of the first feature and the third feature are the same.

[0104] In a possible implementation, the first AI network includes a first mapping network, and the weight parameters of the first mapping network are shared between each single type or each mixed type including at least two types. Alternatively, the first AI network includes a second mapping network corresponding to each single type or each mixed type. Correspondingly, the implementation of step B1 can include any one of step B11 or step B12:

[0105] Step B11, based on the first feature of the first data, using the first mapping network in the first AI network, enhancing the first feature to obtain the third feature corresponding to the first data; or

[0106] Step B12: Based on the first feature of the first data, use the second mapping network in the first AI network corresponding to the first feature to enhance the first feature and obtain the third feature corresponding to the first data.

[0107] In step B11, the BEV features of the camera image, the BEV features of the point cloud data, and the fused BEV features obtained by fusing the image and point cloud data are associated with a first mapping network, which can also be called a shared mapping network.

[0108] like Figure 3 As shown, the Shared Projector network is a BEV feature mapping module that can be used to map first features corresponding to one type of first data, or first features corresponding to at least two types of first data. Specifically, it can be a multilayer linear perceptron, represented as a projector(·). For example, the electronic device can use the Shared Projector network to map first features from any of three branches to a new shared feature space; for example, the first features of any branch can include, but are not limited to, the following three cases: first features of an image, first features of point cloud data, or first features corresponding to both image and point cloud data. The specific formula is as follows:

[0109]

[0110]

[0111]

[0112] Here, `projector(·)` represents a multi-layer linear perceptron (MLP) function. For example, if the number of feature channels C is 256, the MLP has 256 input feature channels and 256 output feature channels. The MLP may include 128 hidden layer neurons. The BEV features of the camera image, the BEV features of the point cloud data, and the fused BEV features of the image and point cloud data share the network parameters of these 128 hidden layer neurons.

[0113] Based on this, a multilayer linear perceptron can be used to further mine alignment knowledge between different single and mixed types of BEV features during the training phase, linking different single types and mixed types of BEV features, including at least two types. During the inference phase, a more general and universal feature representation between different single and mixed types can be obtained, thereby improving the network's generalization ability.

[0114] In step B11, there are respective second mapping networks corresponding between the BEV features of the camera images, the BEV features of the point cloud data, and the fusion BEV features including the images and the point cloud data. The second mapping networks can also be referred to as independent mapping networks.

[0115] For example, the second mapping networks are completely independent mapping networks, that is, the respective mapping networks corresponding between the BEV features of the camera images, the BEV features of the point cloud data, and the fusion BEV features including the images and the point cloud data are completely independent from each other. Correspondingly, the completely independent mapping networks can be represented as:

[0116]

[0117]

[0118]

[0119] Among them, projector1(·), projector2(·), projector3(·) respectively represent the mapping network corresponding to the BEV features of the camera images, the mapping network corresponding to the BEV features of the point cloud data, and the mapping network corresponding to the fusion BEV features. For example, it can be a multi-layer linear perception function corresponding to each independent mapping network.

[0120] In a possible implementation manner, the obtaining of the map image corresponding to the first data based on the third feature corresponding to the first data and using the decoder in the first AI network comprises:

[0121] If the first mapping network is used, the map image corresponding to the first data is obtained based on the third feature corresponding to the first data and the first feature and using the decoder in the first AI network.

[0122] For example, the mapping network can adopt a shared mapping network based on residual connection, and the first mapping network corresponding to the BEV features of the camera images, the BEV features of the point cloud data, and the fusion BEV features including the images and the point cloud data is shared. The first mapping network can be a shared mapping network based on residual connection (Skip Shared Projector). For example, based on the skip connection (Skip Connection), also known as residual connection (Residual Connection), the input first feature is directly connected to the output to allow the first feature to jump between different layers. In a possible example, the mapping network based on residual connection can be implemented through an addition operation, that is, the input and the output are added. For example, the mapping network based on residual connection can be represented as:

[0123]

[0124]

[0125]

[0126] wherein, the projector(·) can be a two-layer linear perception function, the BEV feature of the camera image, the BEV feature of the point cloud data, and the fusion BEV feature including the image and the point cloud data can share the network parameters of the skip projector module (mapping network module based on residual connection).

[0127] In addition, the second mapping network can also adopt a mapping network based on residual connection, that is, if the second mapping network is used, the decoder in the first AI network can also be used to obtain the map image corresponding to the first data based on the third feature corresponding to the first data and the first feature; this implementation manner is the same as the above-mentioned residual connection-based shared mapping network, that is, the first mapping network, and therefore, details are not repeated here.

[0128] Step B2, using the decoder in the first AI network to obtain the map image corresponding to the first data based on the third feature corresponding to the first data.

[0129] The map image can be a high-precision map output by the decoder. The high-precision map can be a map image corresponding to a vectorized map element, for example, the map element class set can include but is not limited to: road boundary, lane dividing line, and pedestrian crossing, etc. Different elements can be distinguished by different colors in the high-precision map.

[0130] In order to better cope with the situation of missing or damaged sensors, in the embodiments of the present application, a switching modality strategy (SMS) is proposed to seamlessly adapt to any modality input, ensuring the compatibility of input configurations across various modalities, such as Figure 5 As shown in FIG. 6, in the inference stage, the trained first AI network supports the use of inputs of any modality, and accurate predictions can be made. The switching modality strategy can be expressed as:

[0131]

[0132] It should be noted that the above switching strategy simulates the real situation of missing sensors in the inference stage. As shown in FIG. 7, if the point cloud data of the LiDAR sensor is missing, for example, due to radar uninstallation or damage (uninstallation or broken), etc., the point cloud data is unavailable, and the first AI network can still make accurate predictions. Figure 5 Figure 5 ​(a) in (a), camera sensor input only, camera BEV features can be selected as input to the BEV decoder (also referred to as BEV decoding module). Likewise, if camera sensor is missing, such as Figure 5 (b) in (b), LiDAR point cloud input only, LiDAR BEV features can be selected as input to the map decoder. Furthermore, if both camera and LiDAR sensor inputs are available, such as Figure 5 (c) in (c), hybrid BEV features can be selected as input to the BEV decoder. Based on this, the trained first AI network can support inputting data of any modality to accurately construct high-definition maps, thereby improving its practicability in autonomous driving.

[0133] The method provided in the application comprises the following steps: obtaining first data, the first data comprising at least one type of data; based on the first data, using a first AI network to obtain a map image corresponding to the first data; in the case where the first data comprises one type of data, based on the first data, using an encoder corresponding to the type to obtain first features of the first data, and based on the first features of the first data, obtaining the map image; in the case where the first data comprises at least two types of data, for each type of data, using an encoder corresponding to the type to extract second features of the type of data, and fusing the second features of the data of each type to obtain first features of the first data, and based on the first features of the first data, obtaining the map image. Based on this, the first AI network can cope with the prediction scenarios of different types of data and the prediction scenarios of mixed data comprising at least two types of data, and both have high accuracy, thereby improving the robustness of the network.

[0134] Figure 1b A flowchart of a method provided by the application is executed by an electronic device, which can be a server, a cloud computing center device or a terminal, etc. As shown in Figure 1b the method comprises the following steps 201-204.

[0135] Step 201, the electronic device obtains a training data set, the training data set comprising at least a plurality of first samples and a plurality of second samples associated with the first samples, the first samples and the second samples being different types of data.

[0136] The first sample comprises first type sample data, the second sample comprises second type sample data, and the label of a sample is a real map image corresponding to the sample data.

[0137] It should be noted that the training data set can include at least two types of samples. For example, in addition to the first sample and the second sample, the training data set can also include a plurality of third samples, and the third sample can include third type of sample data. For example, in an autonomous driving system, three types of sensors such as a camera, a laser radar, and a millimeter wave radar can be configured, and the three types of sensors can provide three types of data, respectively. Of course, the training data set can also include a plurality of fourth samples of a fourth type. The types of samples included in the training data set can be configured as needed, and the present application does not limit this.

[0138] For example, the type of data can also be referred to as the modality of the data. A first sample or a second sample is single-modality sample data including only one modality.

[0139] In a possible scenario, a high-definition map (HD map) construction method plays a crucial role in providing static environment data required by an autonomous driving system. In this method, data can be collected using at least one of various sensors such as a camera and a laser radar, and a map can be constructed using the collected data. Data from different sensors can be considered as data of different modalities.

[0140] In an embodiment of the present application, by obtaining a training sample set, a first AI network can be trained based on single-modality first samples and second samples, so that the trained first AI network can cope with various single modalities and mixed modalities (that is, various single types or various mixed types including at least two types), and the robustness and generalization ability of the trained first AI network are improved. For example, the trained first AI network not only supports map construction based on camera images and radar point cloud data together, but also supports map construction based on only camera images or only radar point cloud data. That is, the trained first AI network performs well under any modality of data; thereby a unified robust high-definition map construction network (Uni-Map) is trained.

[0141] In a possible implementation manner, the first sample is of a first type of data; for each first sample, there can be a second sample associated with the first sample, that is, each first type of data corresponds to associated second type of data.

[0142] Each first sample and the second sample associated with the first sample have the same label.

[0143] It should be noted that in the embodiments of the present application, only the first modality and the second modality are taken as examples, but the same is also applicable to other cases. For example, the first sample can correspond to the second sample and the third sample associated therewith.

[0144] In a possible implementation manner, each first sample is a vehicle environment image collected by a camera at each driving moment, and each second sample is point cloud data of a vehicle environment collected by a radar at each driving moment; the vehicle environment image and the point cloud data corresponding to the same driving moment are associated; the labels corresponding to the associated first sample and the second sample are the same, and can be a real map image corresponding to the vehicle environment at the corresponding driving moment.

[0145] For example, in an automatic driving scene, data of each modality can be collected at each driving moment in the process of vehicle driving, and the data of different modalities collected at the same driving moment are associated. For example, in the process of vehicle driving, at T1 moment, the image of the environment around the vehicle can be collected by a camera (for example, a camera), and the point cloud data of the environment around the vehicle can be collected by a laser radar.

[0146] Step 202: The electronic device obtains, based on the training data set, fourth features of each first sample, fifth features of each second sample, and sixth features of each first sample and the second sample associated with the first sample by using the second AI network.

[0147] In a possible implementation manner, the step 202 can include:

[0148] obtaining, based on each first sample, the fourth feature of the first sample by using an encoder corresponding to the type of the first sample;

[0149] obtaining, based on each second sample, the fifth feature of the second sample by using an encoder corresponding to the type of the second sample;

[0150] For each first sample, the fourth feature of the first sample and the fifth feature of the associated second sample are fused to obtain the sixth feature of the first sample and the second sample associated with the first sample.

[0151] For example, the second AI network includes an encoder corresponding to each type.

[0152] In this step, the electronic device can determine, based on the type corresponding to each sample, an encoder corresponding to the type of each sample from the encoders, and use the encoder corresponding to the type of each sample to extract the initial feature of the corresponding sample, that is, the fourth feature of each first sample and the fifth feature of each second sample; and fuse the initial features of each first sample and the associated second sample to obtain the sixth feature of each first sample and the second sample associated with the first sample.

[0153] For example, the second AI network can include an encoder corresponding to the first type and an encoder corresponding to the second type. The first type of encoder can be used to extract initial features of the first sample, and the second type of encoder can be used to extract initial features of the second sample. For the associated first sample and second sample, a fusion network can be used to fuse the initial features of the first sample and the initial features of the associated second sample to obtain features corresponding to the associated first sample and second sample, i.e., the sixth features. The feature dimensions of the initial features of the first sample, the initial features of the second sample, and the sixth features are the same.

[0154] In one possible embodiment, sensor data χ of any modality can be used as input, and vectorized map elements in the BEV space can be predicted. The map element class set includes, but is not limited to, road boundaries, lane dividers, and pedestrian crossings, etc. As shown in Figure 2 The input data of each modality can be represented as χ = {Camera, LiDAR}, where Camera represents images from a camera, and LiDAR represents point cloud data from a radar. The images from the camera can include multi-view RGB camera images in a perspective view, and can specifically include a total of 6 images from each direction of the front, back, and sides of the vehicle. For example, the images can include front, back, left, right, top, and bottom images. Where B, N cam , H cam , and W cam represent batch size, number of cameras, image height, and image width, respectively. For example, the number of cameras can be 6, and the batch size can be the number of image samples used in one iteration of training. The point cloud data LiDAR can be represented as LiDAR ∈ R B×P×5 , where B represents the batch size, P represents the number of points, and the data of each point can include 3D coordinates, reflectivity, and ring index of the point. The ring index is optional, and if not included, the last dimension can be changed from 5 to 4.

[0155] The second AI network can be an AI network based on a MapTR (Map Transformer) network structure. As shown in Figure 2As shown, the encoder in the second AI network can be a BEV (Birds Eye View) feature encoder. The encoder corresponding to the camera image in the first AI network supports taking multi-view images as input and converting the features of the camera image to a BEV feature space while retaining geometric and semantic information. The encoder corresponding to the radar point cloud data in the second AI network supports converting LiDAR features to a BEV feature space. Of course, other encoders can also be used, and the present application only takes MapTR as an example for illustration.

[0156] For camera images, an image corresponding encoder can be used, for example Figure 2 The 2D image encoding module uses 2D to 3D transformation to encode pixel-level semantic features in image perspective. Resnet50 (Residual Network) can be used as a backbone to extract multi-view features, and GKT (Graph-based Knowledge Tracing) can be used as a 2D to BEV feature conversion module to convert multi-view features to a BEV space. The generated BEV features can be represented as: B, H, W, and C represent batch size, image height and image width, and feature channel number, respectively.

[0157] In one possible example, a bottom-up BEV feature modeling method of LSS (Lift, Splat, Shoot) can be used to extract features from the collected images by explicitly estimating the depth information of the images, and to convert image features to BEV features according to the estimated discrete depth information. Specifically, 2D convolution can be used to extract perspective features from images in various directions In one possible example, a bottom-up BEV feature modeling method of LSS (Lift, Splat, Shoot) can be used to extract features from the collected images by explicitly estimating the depth information of the images, and to convert image features to BEV features according to the estimated discrete depth information. Specifically, 2D convolution can be used to extract perspective features from images in various directions and predict the depth distribution of D equidistant scatter points related to each pixel in the image, where D can be 512 or 256, and the pixel and the scatter point correspond one-to-one; secondly, the perspective features are assigned to D equidistant scatter points along the camera ray direction to obtain D x H x W pseudo point cloud features Finally, the pseudo point cloud features are flattened to a BEV feature space through a pooling operation to realize the conversion from a 2D perspective view to a BEV view, and to obtain the initial features of each image, i.e., the BEV features of the image C represents the number of feature channels, which can be 512 or 256.

[0158] For LiDAR point cloud data, a point cloud data corresponding encoder can be used to extract initial features, for example Figure 2The 3D point cloud encoding module. Specifically, the voxelization and sparse LiDAR encoder can be used in the SECOND model. The LiDAR features are projected to the BEV feature space using the flatten operation in the BEVFusion model to achieve the conversion from the 3D view to the BEV view to obtain the unified BEV feature representation, i.e., the initial features of the point cloud data of the radar, i.e., the BEV features of the point cloud data

[0159] For the associated camera image and point cloud data, the feature fusion method can be used to obtain the fusion BEV features corresponding to the associated camera image and point cloud data. In order to effectively fuse the BEV features from the camera image and the radar point cloud data, a convolution-based fusion method can be used to fuse the BEV features of the image and the BEV features of the point cloud data. Specifically, the features from different modalities can be concatenated in the feature channel C dimension, and then the convolution is used to fuse the features. and The C-channel feature fusion is performed to obtain the fusion features The channel numbers of the features of the three are the same.

[0160] In step 203, the electronic device uses the second AI network to perform prediction based on the fourth features of each first sample, the fifth features of each second sample, and the sixth features of each first sample and the second sample associated with the first sample, to obtain the prediction results corresponding to each sample in the training data set.

[0161] The prediction results corresponding to each sample include the first image corresponding to each first sample, the second image corresponding to each second sample, and the third image corresponding to each first sample and the associated second sample.

[0162] In one possible implementation, step 203 includes steps 2031-2032.

[0163] In step 2031, the electronic device uses the first mapping network in the second AI network to enhance the fourth features, the fifth features, and the sixth features based on the fourth features of each first sample, the fifth features of each second sample, and the sixth features of each first sample and the second sample associated with the first sample, to obtain the enhanced fourth features, the fifth features, and the sixth features.

[0164] In this step, the electronic device can further map the fourth features of each first sample, the fifth features of each second sample, and the sixth features of each first sample and the second sample associated with the first sample, so that the mapped features of each sample remain aligned in the feature space.

[0165] In one possible implementation, mapping processing can be performed using a mapping network. This mapping network can be a shared mapping network or an independent mapping network. For example, the second AI network includes a first mapping network to be trained, which is shared among individual types or mixed types. Alternatively, the second AI network includes separate second mapping networks to be trained for each individual type or mixed type. Accordingly, step 2031 can be implemented in any of the following ways:

[0166] Based on the fourth feature of each first sample, the fifth feature of each second sample, and the sixth feature of each first sample and the second sample associated with it, the fourth, fifth, and sixth features are enhanced using the first mapping network to be trained, resulting in enhanced fourth, fifth, and sixth features; or

[0167] Based on the fourth feature of each first sample, the fifth feature of each second sample, and the sixth feature of each first sample and the second sample associated with the first sample, the fourth feature, the fifth feature, and the sixth feature are enhanced by using the second mapping network to be trained corresponding to each feature, respectively, to obtain the enhanced fourth feature, the fifth feature, and the sixth feature.

[0168] In one possible example, the first mapping network may be called a shared mapping network, in which each sample corresponds to the same network parameters.

[0169] For example, each training sample can correspond to a shared mapping network, that is, different single-modal or mixed-modal sample data share weight parameters in the shared mapping network.

[0170] For example, such as Figure 3 As shown, the shared projector network can be the BEV feature mapping module in the first AI network used to map the initial features. Specifically, it can be a multilayer linear perceptron, which can be represented as a projector(·). For example, the electronic device can use a learnable shared projector network to map the initial features of the image from three branches, the initial features of the point cloud data, and the initial features of the mixed modality to a new shared feature space. The specific formula is as follows:

[0171]

[0172]

[0173]

[0174] wherein, projector(·) represents a multi-layer linear perceptron (MLP) function. For example, if the number of feature channels C is 256, the multi-layer linear perceptron has an input feature channel number of 256 and an output feature channel number of 256, and the multi-layer linear perceptron can include 128 hidden layer neurons, and the network parameters of the 128 hidden layer neurons are shared between the BEV features of the camera image, the BEV features of the point cloud data, and the BEV features of the mixed modality of the mixed image and point cloud data.

[0175] Based on this, different single modalities and mixed modalities can be associated through the multi-layer linear perceptron, further mining alignment knowledge between different single modalities and mixed modalities from the BEV features, and more widely learning more general and universal feature expressions between different single modalities and mixed modalities, thereby improving the generalization ability of the network.

[0176] It should be noted that through step 202, the BEV features of different single modalities and mixed modalities corresponding to the same BEV feature space are obtained; however, due to the depth inaccuracy in the view transformer and the large inter-modality gap, there is still a certain degree of spatial misalignment between the camera image BEV features, the laser radar BEV features, and the BEV features of the mixed modality. That is, the BEV features of different single modalities and mixed modalities can be located in completely independent regions in the BEV feature space. As shown in FIG. 2, the mixed stack modality (MSM) training scheme proposed in the embodiments of the present application can enhance the semantic consistency between the BEV features of different single modalities and mixed modalities through the BEV feature mapping module, so as to achieve more robust alignment and obtain feature expressions with stronger generalization ability. This enables the subsequent decoder to learn rich information from different modalities and improves the generalization of the network. Figure 3

[0177] In another possible manner, each sample can correspond to a partially shared mapping network (Partially shared Projector), which mainly differs from the shared mapping network in that in the partially shared mapping network, the first linear layer (the first linear layer) is not shared, and the second linear layer (the second linear layer) is shared. That is, the first linear layer of the partially shared mapping network learns the knowledge of different single modalities and mixed modalities independently, while the second linear layer learns the knowledge of multiple modalities, for example, the first modality, the second modality, and the mixed modality each correspond to the network parameters of the respective first linear layer, and the network parameters are shared in the second linear layer. ​

[0178] In yet another possible way, different single modalities, different mixed modalities can correspond to different mapping networks. Illustratively, step 2022 can include: based on the initial features of each training sample and the corresponding modality, using the mapping network in the first AI network corresponding to the modality of each training sample, determining the target features corresponding to each training sample.

[0179] Illustratively, the first sample data of different modalities can correspond to different mapping networks, and the second sample data of different mixed modalities can correspond to different mapping networks, that is, each modality corresponds to an independent mapping network (Independent Projector). The processing process of the independent mapping network can be represented as:

[0180]

[0181]

[0182]

[0183] Wherein, projector1(·), projector2(·), projector3(·) respectively represent the mapping network corresponding to the BEV features of the first sample, the mapping network corresponding to the BEV features of the second sample, and the mapping network corresponding to the fusion BEV features of the first sample and the associated second sample; For example, each independent mapping network can correspond to a multi-layer linear perception function.

[0184] By using the mapping network corresponding to the BEV features of the first sample, the BEV features of the second sample, and the fusion BEV features, feature enhancement is performed respectively. Based on this, through subsequent iterative training, the semantic consistency between the BEV features of different single modalities and mixed modalities can also be enhanced, and feature expression with stronger generalization ability can be obtained.

[0185] In yet another possible way, a shared mapping network based on residual connection (Skip Shared Projector) can also be used between different single modalities and different mixed modalities. For example, based on the skip connection (also known as residual connection), the initial features of each modality are directly connected to the output to allow the initial features to be transmitted between different layers. In a possible example, the shared mapping network based on residual connection can be implemented through an addition operation, that is, the input and the output are added. For example, the shared mapping network based on residual connection can be represented as:

[0186]

[0187]

[0188]

[0189] wherein the projector(·) can be a two-layer linear perceptual function, and different single modalities and mixed modalities can share a skip projector module (a mapping network module based on residual connection), for example, the network parameters of the skip projector module are shared between different training samples.

[0190] In step 2032, the electronic device uses the decoder in the second AI network to obtain a prediction result corresponding to each sample based on the enhanced fourth feature, the fifth feature, and the sixth feature, respectively.

[0191] In this step, each sample corresponds to a shared decoder, that is, the network parameters of the decoder are shared between different single modalities and mixed modalities of sample data.

[0192] In this step, the prediction image can be a high-precision map output by the decoder, as shown in the following formula (1) or (2). Figure 2 or Figure 3 The high-precision map can be a map image corresponding to a vectorized map element, for example, the map element class set can include but is not limited to road boundaries, lane dividers, and pedestrian crossings, etc. Different elements in the high-precision map can be distinguished by different colors.

[0193] Based on this, the decoder can learn the knowledge of different single modalities and mixed modalities, and can have high accuracy when dealing with prediction scenarios of any modality, thereby improving the robustness of the network.

[0194] In step 204, the electronic device trains the second AI network based on the prediction results corresponding to each sample in each training data set to obtain the first AI network.

[0195] In this step, the electronic device can divide each sample into a plurality of groups of associated samples, and iteratively train based on the training loss corresponding to each group of associated samples after division.

[0196] In one possible implementation of step 204, the implementation of step 204 can include the following steps 2041-2042:

[0197] In step 2041, the electronic device determines the training loss corresponding to each group of associated samples based on the prediction result and the sample label corresponding to at least one group of associated samples, wherein the prediction result corresponding to a group of associated samples includes a first image corresponding to a first sample, a second image corresponding to an associated second sample, and a third image corresponding to the first sample and the associated second sample.

[0198] For example, taking the first mode and the second mode as examples, each set of associated samples may include a first sample and a second sample associated with the first sample. For instance, a set of associated samples may include images acquired at the same time (which may include 6 images acquired at the same time) and radar point cloud data.

[0199] For each set of associated samples, the loss function trained using the MapTR model can be used to calculate the loss for each predicted image (e.g., the first image, the second image, or the third image). This function consists of three parts, including classification loss. Point-to-point loss and edge direction loss These loss terms are then combined. Specifically, the loss for each predicted image can be calculated using the following formula:

[0200]

[0201] Wherein, in the formula Let λ1, λ2, and λ3 represent the loss corresponding to a predicted image. λ1, λ2, and λ3 are hyperparameters used to balance these loss terms. For example, during training, λ1 is set to 2, λ2 to 5, and λ3 to 5e. -3 .

[0202] For each set of associated samples, the training loss can be calculated using the following formula:

[0203]

[0204] Wherein, in the formula This represents the training loss corresponding to a set of related samples; in a set of related samples, These represent the losses for the first, second, or third image, respectively, with λ4, λ5, and λ6 being the coefficients of the loss for the corresponding predicted image. For example, λ4, λ5, and λ6 can all be 1, meaning the average loss of each predicted image in a set of associated samples is used as the training loss for that set of associated samples. Based on this, the first AI network can learn information from samples of different single-modal and mixed-modal characteristics, achieving high accuracy in prediction scenarios of any modality within different single-modal and mixed-modal contexts, thereby improving the robustness of the AI ​​network.

[0205] Step 2042: The electronic device trains the second AI network based on the training loss corresponding to each group of associated samples.

[0206] In this step, the electronic device can adjust the parameters of each network in the first AI network based on the training loss corresponding to each set of associated samples, for example, including but not limited to: the encoder corresponding to each modality, the mapping network, the decoder, etc. For example, the parameters of the entire network can be updated by the stochastic gradient descent (SGD) algorithm and the chain rule.

[0207] It should be noted that for the decoder, in the prior art high-definition map construction method, the decoder intelligently learns a mode of BEV features, and therefore can only limit the input configuration of one modality. In the method of the present application, in order to ensure that the first AI network trained has good accuracy in any modality prediction scene, the decoder designed in the embodiments of the present application can be applied to a novel hybrid stacked modality training scheme. As shown in Figure 4 In the training phase, a novel hybrid stacked modality training scheme is designed; by sharing the network parameters of the decoder through different single modalities and hybrid modalities, the trained decoder can learn rich knowledge about different single modalities and hybrid modalities. For example, for the BEV features of camera images, the BEV features of lidar, and the BEV features of hybrid modalities, in the training phase, a hybrid stacked modality form can be used to input the decoder for hybrid stacked learning. For example, the hybrid stacked learning process can be represented as:

[0208]

[0209] wherein the stacked BEV features represent the features learned from the BEV features of camera images, the BEV features of lidar, or the BEV features of hybrid modalities, which can be used in the decoder for high-definition map construction tasks.

[0210] It should be noted that the above stacking process is for the batch dimension, and the feature map shape is still HxWxC, so the subsequent map decoder module can be directly applied to the existing network, such as MapTR; based on this, the method of the present application is plug and play. This hybrid stacking strategy allows the map decoder module to learn rich knowledge from camera, lidar, or hybrid modalities of both, improving the robustness of the AI network. Based on this, the trained AI network can be applied to any modality prediction scene.

[0211] It should be noted that the Uni-Map proposed in this application is a novel unified robust high-definition map building network. After training, this network is an all-in-one model that supports operation under arbitrary modal input configurations. To this end, during the training phase, features from all input configurations, such as different single-modal and mixed-modal models, are provided to the network's decoder, and during the inference phase, a specific feature is processed according to the input configuration of the deployed modality.

[0212] like Figure 2 , Figure 3 As shown, given different perceptual inputs, firstly, for different modalities, corresponding modality-specific encoders are used to extract features. The features from different single-modal and mixed-modal approaches are then converted into unified BEV features, which preserve geometric and semantic information.

[0213] Then, this application proposes a novel Hybrid Stacked Modality (MSM) training scheme, enabling the decoder to acquire rich knowledge from fused features of camera, LiDAR, or a hybrid of both modalities. Furthermore, a mapping network is proposed to align BEV features from different single-modality and hybrid modalities into a shared feature space, thereby enhancing representation learning and overall model performance. Finally, the hybrid stacked BEV features are input into the detector and prediction heads for the high-definition map building task. During inference, this application proposes a modality-switching strategy that enables accurate predictions via Uni-Map when using arbitrary modal inputs.

[0214] The method provided in the application comprises the following steps: obtaining a training data set, the training data set comprising at least a plurality of first samples and a plurality of second samples associated with the first samples, the first samples and the second samples being different types of data; obtaining, based on the training data set, fourth features of each first sample, fifth features of each second sample, and sixth features of each first sample and the second sample associated with the first sample by using a second AI network; obtaining, based on the fourth features of each first sample, the fifth features of each second sample, and the sixth features of each first sample and the second sample associated with the first sample, a prediction result corresponding to each sample in the training data set by using the second AI network; wherein the prediction result corresponding to each sample comprises a first image corresponding to each first sample, a second image corresponding to each second sample, and a third image corresponding to each first sample and the associated second sample; and training the second AI network based on the prediction result corresponding to each sample in each training data set to obtain a first AI network. In this way, the first AI network can learn the knowledge of different single modal and mixed modal sample data, and has high accuracy in dealing with any modal prediction scene, thereby improving the robustness of the network.

[0215] It should be noted that in the related art, cameras and lidars are the main sensors used, and the input configurations are different (only camera image input, only lidar input, or camera image-lidar fusion input) according to cost performance considerations. Among them, the method based on camera image-lidar fusion performs best. However, the existing method has two major technical defects: 1. The existing method needs to train and deploy a separate model for each input configuration, resulting in a large amount of development, maintenance and deployment overhead. 2. The existing method usually needs to assume that the model is always designed to access complete sensor information. The model has high cost and low robustness in the case of missing or damaged sensors, that is, when the sensor is missing or damaged, its performance may decrease significantly or even crash. Among them, sensor loss means that the sensor is unavailable, for example, camera loss only has radar point cloud data as input. Sensor damage means that the sensor provides damaged data, for example, one of the six direction cameras of the camera is damaged, and there are five images and radar point cloud data as input. As shown in Figure 6 , the impact of multi-sensor damage on the camera image-lidar fusion model, as shown in Figure 6 , each case and the evaluation index corresponding to each case, for example, the evaluation index can use mAP (Mean Average Precision, average precision mean); in the prior art, when the sensor is damaged or missing, whether it is from one sensor damage or two sensor damage, the performance of the camera image-lidar fusion model decreases significantly.

[0216] The unified robust HD map construction network (Uni-Map) proposed in the embodiments of the present application is a single model that performs well in all input configurations. A new hybrid stack mode (MSM) training scheme is specifically proposed to enable the map decoder to effectively learn information from features of camera, lidar, and mixed modalities. In addition, a mapping module is proposed to align BEV features from different single modalities and mixed modalities into a shared feature space, thereby enhancing feature expression and improving overall network performance. In the inference stage, a switching modal strategy (SMS) is proposed to seamlessly adapt to any modality input, ensuring compatibility across various input configurations. The Uni-Map can achieve higher performance under different input configurations while reducing the training and deployment costs of the model.

[0217] Figure 6 The performance comparison results of the method (Uni-Map) of the present application and the prior art are also shown in the Figure 6 As shown in the Figure 6 As shown in the

[0218] The structure of the first AI network can be implemented using the deep learning framework Pytorch and tested on the nuScenes dataset. As shown in the Figure 7 Figure 7 The high-precision maps output by the method of the present application and the method of the prior art under normal conditions, loss of point cloud data, and loss of camera images, respectively, are shown in the Figure 7 As shown in the

[0219] ​The electronic device provided in the embodiments of the present disclosure includes a processor, and optionally, a transceiver and / or a memory coupled to the processor, and the processor is configured to perform the steps of the method provided in any of the optional embodiments of the present disclosure.

[0220] Figure 8 The structure of an electronic device to which the embodiments of the present disclosure are applied is shown in FIG. 4A. Figure 8 As shown in FIG. 4A, Figure 8 The electronic device 4000 shown in FIG. 4A includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, through a bus 4002. Optionally, the electronic device 4000 can further include a transceiver 4004, which can be used for data interaction, such as data transmission and / or data reception, between the electronic device and other electronic devices. It should be noted that the transceiver 4004 is not limited to one in actual application, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure. Optionally, the electronic device can be a first network node, a second network node or a third network node.

[0221] The processor 4001 can be a CPU (Central Processing Unit, central processing unit), a general-purpose processor, a DSP (Digital Signal Processor, digital signal processor), an ASIC (Application Specific Integrated Circuit, application specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure. The processor 4001 can also be a combination of computing functions, such as one or more microprocessor combinations, combinations of DSP and microprocessor, etc.

[0222] The bus 4002 can include a path for transmitting information between the above-mentioned components. The bus 4002 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 8 In the embodiments of the present disclosure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0223] The memory 4003 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing computer instructions and capable of being read by a computer, without limitation.

[0224] The memory 4003 is configured to store a computer program for implementing the embodiments of the present disclosure, and the processor 4001 is configured to control the execution of the computer program stored in the memory 4003. The processor 4001 is configured to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0225] The embodiments of the present disclosure provide a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0226] The embodiments of the present disclosure also provide a computer program product, and the computer program product includes a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0227] The terms "first", "second", "third", "fourth", "1", "2", and the like (if any) in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that shown or described.

[0228] It should be understood that although the various operation steps in the flowcharts of the embodiments of the present disclosure are indicated by arrows, the implementation order of the steps is not limited to the order indicated by the arrows. Unless otherwise specified herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be executed in other orders as required. In addition, part or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on the actual implementation scenario. Part or all of these sub-steps or stages can be executed at the same time, and each of these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured as required, and the embodiments of the present disclosure do not limit this.

[0229] The above text and drawings are provided only as examples to help the reader understand the present disclosure. They are not intended to and should not be interpreted to limit the scope of the present disclosure in any way. Although certain embodiments and examples have been provided, it will be apparent to those skilled in the art based on the disclosure herein that modifications can be made to the embodiments and examples shown, and other similar implementations based on the technical ideas of the present disclosure, and these also belong to the protection scope of the embodiments of the present disclosure.

Claims

1. A method performed by an electronic device, characterized in that, The method includes: Acquire first data, wherein the first data includes at least one type of data; Based on the first data, a map image corresponding to the first data is obtained using a first artificial intelligence (AI) network. Specifically, based on the first data, a first AI network is used to obtain the map image corresponding to the first data, including: If the first data includes a type of data, based on the first data, an encoder corresponding to the type is used to obtain a first feature of the first data, and based on the first feature of the first data, the map image is obtained; When the first data includes at least two types of data, for each type of data, an encoder corresponding to the type is used to extract the second feature of the data of that type, and the second features of each type of data are fused to obtain the first feature of the first data. Based on the first feature of the first data, the map image is obtained.

2. The method according to claim 1, characterized in that, Based on the first data, using the first AI network, a map image corresponding to the first data is obtained, including: Determine the type of data contained in the first data; If the first data includes a type of data, based on the first data, an encoder corresponding to the determined type is used to obtain the first feature of the first data; When the first data includes at least two types of data, for each determined type of data, an encoder corresponding to the determined type is used to extract the second feature of the determined type of data, and the second features of each determined type of data are fused to obtain the first feature of the first data.

3. The method according to claim 1 or 2, characterized in that, The process of obtaining the map image based on the first feature of the first data includes: Based on the first feature of the first data, the mapping network in the first AI network is used to enhance the first feature to obtain the third feature corresponding to the first data; Based on the third feature corresponding to the first data, the map image corresponding to the first data is obtained using the decoder in the first AI network.

4. The method according to claim 3, characterized in that, The step of enhancing the first feature based on the first data using the mapping network in the first AI network to obtain the third feature corresponding to the first data includes: Based on the first feature of the first data, the first mapping network in the first AI network is used to enhance the first feature to obtain the third feature corresponding to the first data; or Based on the first feature of the first data, the first feature is enhanced using the second mapping network in the first AI network that corresponds to the first feature, thereby obtaining the third feature corresponding to the first data.

5. The method according to claim 4, characterized in that, The step of obtaining the map image corresponding to the first data based on the third feature corresponding to the first data using the decoder in the first AI network includes: If the first mapping network is used, based on the third feature and the first feature corresponding to the first data, the decoder in the first AI network is used to obtain the map image corresponding to the first data.

6. The method according to any one of claims 1-5, characterized in that, The first data includes image data acquired by the camera and / or point cloud data acquired by the radar.

7. The method according to any one of claims 1-5, characterized in that, The first feature is a BEV feature, and the second feature is a BEV feature.

8. A method performed by an electronic device, characterized in that, The method includes: Obtain a training dataset, which includes at least a plurality of first samples and a plurality of second samples associated with the first samples, wherein the first samples and the second samples are data of different types; Based on the training dataset, the second AI network is used to obtain the fourth feature of each first sample, the fifth feature of each second sample, and the sixth feature of each first sample and the second sample associated with the first sample. Based on the fourth feature of each first sample, the fifth feature of each second sample, and the sixth feature of each first sample and the second sample associated with the first sample, the second AI network is used to make predictions to obtain the prediction results corresponding to each sample in the training dataset. The prediction results for each sample include a first image corresponding to each first sample, a second image corresponding to each second sample, and a third image corresponding to each first sample and its associated second sample. The second AI network is trained based on the prediction results of each sample in each training dataset to obtain the first AI network.

9. The method according to claim 8, characterized in that, Based on the training dataset, the second AI network is used to obtain the fourth feature of each first sample, the fifth feature of each second sample, and the sixth feature of each first sample and the second sample associated with the first sample, including: Based on each first sample, the fourth feature of each first sample is obtained using an encoder corresponding to the type of the first sample; Based on each second sample, the fifth feature of each second sample is obtained using an encoder corresponding to the type of the second sample; For each first sample, the fourth feature of the first sample and the fifth feature of the associated second sample are fused to obtain the sixth feature of the first sample and the second sample associated with the first sample.

10. The method according to claim 8, characterized in that, The second AI network is used to make predictions based on the fourth feature of each first sample, the fifth feature of each second sample, and the sixth feature of each first sample and the second sample associated with the first sample, respectively, to obtain the prediction results corresponding to each sample in the training dataset, including: Based on the fourth feature of each first sample, the fifth feature of each second sample, and the sixth feature of each first sample and the second sample associated with the first sample, the fourth feature, the fifth feature, and the sixth feature are enhanced using the first mapping network in the second AI network to obtain the enhanced fourth feature, the fifth feature, and the sixth feature. Based on the enhanced fourth, fifth, and sixth features, respectively, the decoder in the second AI network is used to obtain the prediction results corresponding to each sample.

11. The method according to any one of claims 8-10, characterized in that, The process of training the second AI network based on the prediction results corresponding to each sample in each training dataset to obtain the first AI network includes: Based on the prediction results and sample labels corresponding to at least one set of associated samples, the training loss corresponding to each set of associated samples is determined. The prediction results corresponding to a set of associated samples include a first image corresponding to a first sample, a second image corresponding to an associated second sample, and a third image corresponding to the first sample and the associated second sample. The second AI network is trained based on the training loss corresponding to each group of associated samples.

12. The method according to any one of claims 8-11, characterized in that, The first sample is image data acquired by a camera, and the second sample is point cloud data acquired by radar.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 12.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.