Map generation method and device, equipment and storage medium
By extracting features from image data from multiple cameras and using a map generation model for spatial transformation, the problem of limited field of view of vehicle-mounted cameras and high calibration costs of roadside cameras is solved, generating efficient and low-cost vectorized maps suitable for autonomous driving.
Patent Information
- Application Number
- CN202511101324.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
In complex road sections such as intersections, the limited field of view of vehicle-mounted cameras results in poor map quality. Existing roadside cameras have high calibration costs and poor adaptability to dynamic environments.
By extracting features from image data from multiple cameras and using geometric transformation parameters in the map generation model to perform spatial transformation, a vectorized map is generated, replacing traditional on-site calibration equipment and processes.
It improves the efficiency of map generation and updating, reduces calibration costs and calibration errors in dynamic environments, and the generated vectorized maps are suitable for autonomous driving systems.
Smart Images

Figure CN120997335A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic map technology, and in particular to a map generation method, apparatus, device and storage medium. Background Technology
[0002] In electronic map creation scenarios such as high-precision maps, vehicle-mounted cameras can be used to collect map data and generate maps. However, in complex road sections such as intersections, the limited field of view of vehicle-mounted cameras makes it difficult to fully capture the entire intersection, resulting in poor map quality.
[0003] In related technologies, fixed roadside cameras are used to supplement the limited field of view of vehicle-mounted cameras. These roadside cameras can provide a wider field of view and capture more intersection details. However, because the parameters of each camera need to be precisely calibrated to ensure that the captured image data can be accurately used for map generation, this usually requires setting up specialized calibration equipment and scenes at the intersection or roadside, involving a large amount of manpower and material resources, resulting in high costs. Furthermore, changes in external factors (such as slight movements of camera positions) may cause changes in camera parameters, which may require multiple calibrations at irregular intervals, further increasing costs.
[0004] Therefore, how to reduce the calibration cost of roadside cameras and improve their adaptability in dynamic environments has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a map generation method, apparatus, device, and storage medium to at least solve one of the above-mentioned technical problems.
[0006] According to one aspect of this application, a map generation method is provided, comprising:
[0007] After acquiring image data from multiple cameras, a first image feature is extracted from the image data; wherein, the image data includes road images from multiple perspectives, and the first image feature includes first image features from all perspectives;
[0008] According to the preset map generation model, the first image features are spatially transformed, and a vectorized map is generated based on the spatially transformed second image features.
[0009] The map generation model includes geometric transformation parameters, which are used to spatially transform the first image features from each viewpoint to the second image features from the target viewpoint.
[0010] In one implementation, the map generation model includes a spatial transformation network for determining the geometric transformation parameters;
[0011] The spatial transformation processing of the first image features includes:
[0012] For the first image feature under each viewpoint, the spatial transformation network determines the corresponding geometric transformation parameters based on the first image feature to obtain the geometric transformation parameters of the first image feature under each viewpoint.
[0013] Based on the geometric transformation parameters of the first image features under each viewpoint, spatial transformation is performed on the corresponding first image features to obtain the spatially transformed second image features.
[0014] In one embodiment, the spatial transformation network includes a localization network, a mesh generator, and a sampler; wherein the geometric transformation parameters are determined based on the localization network after locating the first image features;
[0015] The step of performing spatial transformation on the corresponding first image features according to the geometric transformation parameters of the first image features under each viewpoint to obtain the spatially transformed second image features includes:
[0016] For each first image feature under each viewpoint, a sampling grid for spatial mapping of the first image feature is generated by the grid generator based on the geometric transformation parameters. The sampling grid defines the mapping relationship between the pixels of the first image feature and the second image feature.
[0017] The sampler performs sampling processing on the first image features according to the sampling grid to spatially transform the first image features and obtain the second image features; wherein, the sampling processing includes extracting pixel values from the first image features and mapping them to the corresponding positions of the second image features.
[0018] In one implementation, the training method of the spatial transformation network includes:
[0019] Extract the first historical image features from the historical image data;
[0020] The first historical image features are input into the initial spatial transformation network to obtain the second historical image features after spatial transformation of the first historical image features;
[0021] Based on the predicted values of vector map elements obtained from the second historical image features and the loss calculation results between them and the true values of the map elements, the initial spatial transformation network is trained to obtain the spatial transformation network.
[0022] In one implementation, the map generation model includes a feature fusion network; the second image feature includes image features corresponding to the first image features from all viewpoints after spatial transformation.
[0023] The generation of a vectorized map based on the spatially transformed second image features includes:
[0024] According to the feature fusion network, the image features corresponding to the first image features under each viewpoint after spatial transformation are fused to obtain fused image features; wherein, the feature fusion network is used to perform feature fusion on the first image features under each viewpoint in the feature layer in the feature channel and / or feature space.
[0025] Based on the fusion features, vectorized map elements in the image data are identified, and a vectorized map is constructed based on the vectorized map elements.
[0026] In one implementation, the map generation model further includes a map output head, which comprises a classification branch and a point regression branch;
[0027] The step of identifying vectorized map features in the image data based on the fused image features includes:
[0028] The fused image features are input into the map output header;
[0029] Based on the classification branch, the fused image features are classified and predicted to obtain the feature categories of map elements;
[0030] Based on the point regression branch, the location of the fused image features is predicted to obtain the location coordinates of the map elements;
[0031] Based on the feature category and the location coordinates, identify the vectorized map features in the image data.
[0032] In one embodiment, the map generation model further includes a feature extraction network, which includes a backbone network and a feature pyramid network.
[0033] The step of extracting the first image feature from the image data includes:
[0034] Multi-scale features of the image data are extracted based on the backbone network, and the multi-scale features are input into the feature pyramid network;
[0035] The first image feature is obtained by fusing the multi-scale features according to the feature pyramid network.
[0036] In one embodiment, the method is applied to an edge computing device, on which the map generation model is deployed; wherein the deployment method of the map generation model includes:
[0037] The map generation model is converted into a map generation model corresponding to the target format using a pre-set conversion tool; wherein the target format is compatible with the inference engine of the edge device.
[0038] The converted map generation model is deployed in the edge computing device, which is an embedded device that can be embedded in multiple cameras.
[0039] According to a second aspect of this application, a map generation apparatus is provided, comprising:
[0040] The first processing module is configured to extract a first image feature from the image data after acquiring image data from multiple cameras; wherein the image data includes road images from multiple perspectives, and the first image feature includes first image features from all perspectives.
[0041] The second processing module is configured to perform spatial transformation processing on the first image features according to a preset map generation model, and generate a vectorized map based on the spatially transformed second image features.
[0042] The map generation model includes geometric transformation parameters, which are used to spatially transform the first image features from each viewpoint to the second image features from the target viewpoint.
[0043] In one implementation, the map generation model includes a spatial transformation network for determining the geometric transformation parameters;
[0044] The second processing module includes:
[0045] The determining unit is configured to determine the corresponding geometric transformation parameters for the first image features under each viewpoint through the spatial transformation network, thereby obtaining the geometric transformation parameters of the first image features under each viewpoint.
[0046] The spatial transformation unit is configured to perform spatial transformation on the corresponding first image features according to the geometric transformation parameters of the first image features under each viewpoint, so as to obtain the spatially transformed second image features.
[0047] In one embodiment, the spatial transformation network includes a localization network, a mesh generator, and a sampler; wherein the geometric transformation parameters are determined based on the localization network after locating the first image features;
[0048] The spatial transformation unit is specifically configured as follows: for a first image feature under each viewpoint, a sampling grid for spatial mapping of the first image feature is generated by the grid generator based on the geometric transformation parameters. The sampling grid defines the mapping relationship between the pixels of the first image feature and the second image feature. The first image feature is sampled according to the sampling grid by the sampler to perform spatial transformation on the first image feature to obtain the second image feature. The sampling process includes extracting pixel values from the first image feature and mapping them to the corresponding positions of the second image feature.
[0049] In one embodiment, a training module for training the spatial transformation network is further included, the training module comprising:
[0050] The first extraction unit is configured to extract the first historical image feature from the historical image data;
[0051] The acquisition unit is configured to input the first historical image features into an initial spatial transformation network to obtain the second historical image features after spatial transformation of the first historical image features;
[0052] An optimization unit is configured to train the initial spatial transformation network using the loss calculation result between the predicted values of vector map elements obtained from the second historical image features and the true values of the map elements, in order to obtain the spatial transformation network.
[0053] In one implementation, the map generation model includes a feature fusion network; the second image feature includes image features corresponding to the first image features from all viewpoints after spatial transformation.
[0054] The second processing module includes:
[0055] The first fusion unit is configured to fuse the image features corresponding to the first image features under each viewpoint after spatial transformation according to the feature fusion network to obtain fused image features; wherein, the feature fusion network is used to perform feature fusion on the first image features under each viewpoint at the feature layer in the feature channel and / or feature space.
[0056] The identification unit is configured to identify vectorized map elements in the image data based on the fusion features, and construct a vectorized map based on the vectorized map elements.
[0057] In one implementation, the map generation model further includes a map output head, which includes a classification branch and a point regression branch;
[0058] The identification unit is specifically configured to: input the fused image features into the map output head; perform classification prediction on the fused image features according to the classification branch to obtain the feature category of the map element; perform position prediction on the fused image features according to the point regression branch to obtain the position coordinates of the map element; and identify the vectorized map element in the image data according to the feature category and the position coordinates.
[0059] In one embodiment, the map generation model further includes a feature extraction network, which includes a backbone network and a feature pyramid network.
[0060] The first processing module includes:
[0061] The second extraction unit is configured to extract multi-scale features of the image data based on the backbone network and input the multi-scale features into the feature pyramid network;
[0062] The second fusion unit is configured to perform feature fusion on the multi-scale features according to the feature pyramid network to obtain the first image features.
[0063] In one embodiment, the method is applied to an edge computing device, on which the map generation model is deployed; wherein the deployment method of the map generation model includes:
[0064] The map generation model is converted into a map generation model corresponding to the target format using a pre-set conversion tool; wherein the target format is compatible with the inference engine of the edge device.
[0065] The converted map generation model is deployed in the edge computing device, which is an embedded device that can be embedded in multiple cameras.
[0066] According to a third aspect of this application, an electronic device is provided, comprising: a memory and a processor;
[0067] The memory stores computer-executed instructions;
[0068] The processor executes computer execution instructions stored in the memory, causing the electronic device to perform the map generation method provided by any of the first aspects above.
[0069] According to a fourth aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the map generation method provided in any of the first aspects above.
[0070] According to a fifth aspect of this application, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the map generation method provided in any of the first aspects above.
[0071] The map generation method, apparatus, device, and storage medium provided in this application, after acquiring image data from multiple cameras, extracts first image features from the image data. The image data includes road images from multiple perspectives, and the first image features include first image features from all perspectives. Based on a preset map generation model, the first image features undergo spatial transformation processing, and a vectorized map is generated based on the spatially transformed second image features. The map generation model includes geometric transformation parameters used to spatially transform the first image features from each perspective to second image features from the target perspective. During this process, when the map generation model generates the vectorized map, the geometric transformation parameters in the model are used to spatially transform the first image features from different perspectives. The spatially transformed second image features are represented in a unified coordinate system, eliminating geometric differences caused by different camera perspectives. This eliminates the need to deploy calibration equipment at intersections or roadsides, effectively improving the efficiency of map building or updating, and reducing calibration costs and calibration errors in dynamic environments. Attached Figure Description
[0072] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0073] Figure 1 This is a schematic diagram of a possible scenario provided for an embodiment of this application;
[0074] Figure 2 This is an example diagram of image data collected by multiple cameras in an embodiment of this application;
[0075] Figure 3 A schematic flowchart illustrating the map generation method provided in this application embodiment;
[0076] Figure 4 This is an example diagram of the feature extraction network in the embodiments of this application;
[0077] Figure 5 This is an example diagram of the map generation model in the embodiments of this application;
[0078] Figure 6 This is an example diagram of the vectorized map output by the map generation model in the embodiments of this application;
[0079] Figure 7 This is a schematic diagram of the structure of the map generation device provided in the embodiments of this application;
[0080] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0081] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0082] Currently, the creation and updating of electronic maps, such as high-precision maps, are mainly achieved through offline technologies, such as professional surveying vehicles collecting map data and generating maps, or online methods, such as using vehicle-mounted cameras to collect map data in real time and generate maps.
[0083] Offline methods rely on specialized surveying vehicles equipped with cameras and LiDAR sensors to collect LiDAR point cloud and image data. This data is then manually annotated and processed offline to generate high-precision vectorized maps. While this method offers high accuracy, the high cost of data acquisition and processing, coupled with long update cycles, makes it difficult to meet the real-time and high-precision requirements of autonomous driving systems. Online methods, on the other hand, utilize onboard cameras and LiDAR sensors to generate maps in real time. Examples include High Definition Map Networks (HDMapNet), Vector Map Networks (VectorMapNet), and Map Transformers (MapTR). These methods achieve real-time map generation through multimodal data fusion. However, limitations imposed by the field of view and resolution of onboard cameras and LiDAR sensors can lead to occlusion and limited view at complex intersections, resulting in incomplete map generation and compromising the safety of autonomous driving systems.
[0084] In related technologies, to address the performance limitations of online methods at complex intersections, image data collected by roadside cameras is used to supplement intersection details, thereby improving map completeness. However, the aforementioned schemes using roadside cameras to collect image data require calibration for each camera. Calibration of camera parameters necessitates the deployment of calibration equipment and scenarios at the intersection or roadside, resulting in high costs. Furthermore, changes in external factors may cause variations in camera parameters, potentially requiring multiple calibrations periodically, further increasing costs.
[0085] In view of this, embodiments of this application provide a map generation method, apparatus, device, and storage medium. After acquiring image data from multiple cameras, a first image feature is extracted from the image data. The image data includes road images from multiple perspectives, and the first image feature includes first image features from all perspectives. Based on a preset map generation model, the first image feature undergoes spatial transformation processing, and a vectorized map is generated based on the spatially transformed second image feature. The map generation model includes geometric transformation parameters, which are used to spatially transform the first image feature from each perspective to the second image feature from the target perspective. In this process, when the map generation model generates a vectorized map, the geometric transformation parameters in the model are used to spatially transform the image features from different perspectives, replacing existing calibration methods. This eliminates the need to deploy calibration equipment at intersections or roadsides, effectively improving the efficiency of map building or updating, and reducing calibration costs and calibration errors in dynamic environments.
[0086] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0087] The embodiments of this application will be explained below in conjunction with application scenarios. The map generation method provided in the embodiments of this application can be applied to the application scenario of intelligent driving, and more specifically, it can be applied to the application scenario of autonomous driving based on vehicle cloud computing. For example, the execution subject of the method provided in the embodiments of this application can be a server, and more specifically, for example, the server of the high-precision map producer. The following will introduce the method provided in the embodiments of this application as the execution subject of the server.
[0088] Figure 1 This is a schematic diagram of a map generation method provided in an embodiment of this application, such as... Figure 1 As shown, server 110 establishes network connections with multiple roadside cameras 120 (hereinafter referred to as multiple roadside cameras) and intelligent vehicle 130. Server 110 generates high-precision map data (in this embodiment, "generating" includes "updating" the high-precision map data) and transmits the high-precision map data to intelligent vehicle 130. Intelligent vehicle 130 can use the high-precision map data to assist in autonomous driving. Server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, and cloud computing.
[0089] Optionally, during the generation of high-precision map data, server 110 receives image data from multiple roadside cameras 120, extracts image features from the image data, and generates a high-precision map using a map generation model deployed on the server. These multiple roadside cameras 120 can be cameras located at different positions at the intersection, used to collect image data containing objects in the high-precision map (such as lane lines, poles, signs, guardrails, etc.). For example, the image data collected by multiple roadside cameras at an intersection may be as follows: Figure 2 As shown.
[0090] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that these specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0091] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0092] Figure 3 This is a flowchart illustrating a map generation method provided in an embodiment of this application, as shown below. Figure 3 As shown, the method includes steps S301 and S302:
[0093] Step S301: After acquiring image data from multiple cameras, extract the first image feature from the image data; wherein the image data includes road images from multiple perspectives, and the first image feature includes first image features from all perspectives.
[0094] In this embodiment, the multiple cameras can be multiple roadside cameras already deployed in the traffic environment. Utilizing multiple roadside cameras provides multi-view images, enhancing the completeness and accuracy of the map. For example, the collected image data can be image data of intersections, such as... Figure 2 As shown. Each camera can have a resolution of 2 megapixels to provide high-quality images.
[0095] like Figure 2As shown, the image data from multiple cameras are typically road images from different perspectives. In related technologies, to ensure consistency in the perspectives of multiple cameras, calibration equipment is used for physical calibration to unify the perspectives of road images from different angles. However, this method has high calibration costs and is prone to requiring recalibration due to external factors such as changes in camera positions.
[0096] Instead of physical calibration, this embodiment extracts features from image data from uncalibrated cameras and uses geometric transformation parameters in the map generation model to spatially transform the image features, directly generating a vectorized map to improve map generation efficiency.
[0097] Optionally, the image feature extraction process can be performed within the map generation model or independently of it. One approach is to directly input the image data from multiple cameras into the map generation model, and after a series of processing steps (including image feature extraction), output a vectorized map. Another approach is to extract image features using a separate feature extraction network after acquiring the image data from multiple cameras. This embodiment does not impose any particular limitation on this approach.
[0098] It is understood that the first image feature and the second image feature in this embodiment are used to describe similar objects, wherein the first image feature is an image feature after feature extraction, and the second image feature is an image feature after spatial transformation processing of the first image feature.
[0099] Next, we will further introduce the first method mentioned above: the map generation model can include a feature extraction network, which includes a backbone network and a feature pyramid network.
[0100] In this embodiment, the feature extraction network is used to extract multi-scale, multi-level feature information from the input image. It includes a backbone network and a feature pyramid network (FPN). The backbone network can employ a convolutional neural network (CNN) architecture. The feature pyramid network is an architecture used to further process and integrate the features extracted by the backbone network to generate feature maps with multi-scale, multi-level features. In other words, the main goal of the FPN is to generate multi-scale feature maps with rich semantic information to better handle scale variation problems in tasks such as object detection and semantic segmentation.
[0101] The extraction of the first image feature from the image data in step S301 above can be achieved in the following manner:
[0102] Multi-scale features of the image data are extracted based on the backbone network, and the multi-scale features are input into the feature pyramid network;
[0103] The first image feature is obtained by fusing the multi-scale features according to the feature pyramid network.
[0104] For example, such as Figure 4 As shown, a convolutional neural network (CNN) can be used to extract features from a set of images. Specifically, ResNet50 can be used as the backbone to extract image features, and then the multi-scale features are input into a feature pyramid network for fusion. To improve the detection capability of small and occluded targets, a path aggregation network (PANet) can be used as the feature pyramid network to enhance the fusion of low-level details and high-level semantic information.
[0105] Thus, by adopting an end-to-end map generation model, vectorized maps can be generated directly from input image data, simplifying the multi-stage processing flow in related technologies and reducing computational complexity.
[0106] In some embodiments, as many cameras as possible can be added, such as up to eight cameras, to increase the number of input images, cover more viewpoints, reduce occlusion and view limitation issues, thereby further improving the integrity and accuracy of the map.
[0107] Step S302: According to the preset map generation model, the first image features are spatially transformed, and a vectorized map is generated based on the spatially transformed second image features; wherein, the map generation model includes geometric transformation parameters, which are parameters used to spatially transform the first image features under each viewpoint to the second image features under the target viewpoint.
[0108] In related technologies, there are also methods for generating vectorized maps using machine learning models. However, to improve map generation accuracy, the image data used to generate vectorized maps is calibrated. In this embodiment, during the image feature processing of the map generation model, spatial transformation of image features can be achieved using the geometric transformation parameters in the map generation model. These geometric transformation parameters can be implicit parameters in the model, serving as intermediate variables in the generation of the bird's-eye view feature map, used for spatial transformation during feature mapping. Alternatively, they can act as parameters in the large convolutional kernel used in the feature fusion module to achieve spatial transformation. In this process, after image data is input into the model, spatial transformation can be directly performed using the geometric transformation parameters in the model, ensuring spatial consistency of the image. This eliminates the need for additional calibration equipment or an additional calibration environment to perform other calibration processes, effectively reducing calibration costs and avoiding the need for recalibration due to invalid calibration caused by external factors, thus simplifying map generation efficiency.
[0109] Specifically, this embodiment utilizes the geometric transformation parameters included in the map generation model to perform spatial transformation on the first image features, converting image features from each viewpoint to a unified viewpoint. This spatial transformation process can correct geometric distortions caused by different camera installation positions and angles, enabling image data acquired from different cameras to be compared and analyzed in the same coordinate system. Thus, during the map generation process of the map generation model, the physical calibration effect of multiple cameras can be achieved without the need for additional calibration methods.
[0110] In this embodiment, the second image feature can be a feature that is closer to the bird's-eye view (BEV) feature space, hereinafter referred to as BEV feature.
[0111] In one implementation, the map generation model may include a spatial transformation network for determining the geometric transformation parameters. The spatial transformation processing of the first image features in step S302 above can be performed in the following manner:
[0112] For the first image feature under each viewpoint, the spatial transformation network determines the corresponding geometric transformation parameters based on the first image feature to obtain the geometric transformation parameters of the first image feature under each viewpoint.
[0113] Based on the geometric transformation parameters of the first image features under each viewpoint, spatial transformation is performed on the corresponding first image features to obtain the spatially transformed second image features.
[0114] In this embodiment, the map generation model includes a Spatial Transformer Network (STN). The STN performs geometric transformations of image features, so that the network input is only the camera image and does not depend on the camera parameters. This reduces calibration costs and the impact of parameter errors on the generated map, and enhances the network's adaptability to different viewpoints.
[0115] Specifically, a spatial transformation network is introduced. This network, part of the map generation model, can be used to automatically determine geometric transformation parameters. For the first image features at each viewpoint, the spatial transformation network can calculate the corresponding geometric transformation parameters through a learning algorithm. These parameters are used to transform the image features from different viewpoints to the target viewpoint, generating spatially transformed second image features. Finally, these features are integrated to generate a vectorized map. This process of generating a vectorized map is more flexible and adaptable, allowing for dynamic adjustment of geometric transformation parameters to accommodate different camera configurations and environmental changes. In some embodiments, in addition to setting the spatial transformation network within the map generation model to determine geometric transformation parameters, the parameters can also be predefined (those skilled in the art can adaptively determine these parameters based on practical applications or empirical values). For example, when the intersection scene is known, the camera's viewpoint and position are fixed, and the required parameters can be predefined as the model's geometric transformation parameters. In other embodiments, the map generation model can also be trained using supervised learning by using labeled training data, so that the map generation model can predict geometric transformation parameters. The training data may include the input image and its corresponding target transformation (such as the target image or transformation parameters), etc. This embodiment does not particularly limit the specific determination process of the geometric transformation parameters.
[0116] In a further example of the above implementation, the spatial transformation network may include a localization network, a mesh generator, and a sampler; wherein the geometric transformation parameters are determined based on the localization network after locating the first image features;
[0117] The step of performing spatial transformation on the corresponding first image features according to the geometric transformation parameters of the first image features under each viewpoint to obtain the spatially transformed second image features includes:
[0118] For each first image feature under each viewpoint, a sampling grid for spatial mapping of the first image feature is generated by the grid generator based on the geometric transformation parameters. The sampling grid defines the mapping relationship between the pixels of the first image feature and the second image feature.
[0119] The sampler performs sampling processing on the first image features according to the sampling grid to spatially transform the first image features and obtain the second image features. The sampling processing includes extracting pixel values from the first image features and mapping them to the corresponding positions of the second image features. A localization network can be used to predict the geometric transformation parameters required for the input image; the localization network can employ a small convolutional neural network (CNN) structure. A grid generator can generate a sampling grid for the target image based on the geometric transformation parameters output by the localization network. This grid defines how pixels in the input image are mapped to pixel positions in the output image, and its output is a two-dimensional coordinate grid representing the position of each pixel in the input image in the output image. The sampler extracts pixel values from the input image based on the coordinate grid provided by the grid generator to generate the transformed output image. The sampler can use interpolation methods (such as bilinear interpolation) to calculate pixel values at non-integer coordinate positions to address the issue that the coordinates output by the grid generator are usually not integers. Based on the above structure, the output is a spatially transformed image.
[0120] Specifically, in this spatial transformation process, the localization network first analyzes the first image features at each viewpoint to determine geometric transformation parameters. These parameters describe how to transform the image features from the original viewpoint to the target viewpoint. The localization network can utilize a deep learning model to identify the spatial layout of the features and output a set of geometric transformation parameters, such as rotation, translation, and scaling parameters.
[0121] Next, the mesh generator generates a sampling mesh based on these geometric transformation parameters. This mesh defines how each pixel in the first image feature maps to the second image feature. The mesh generator determines the new position of each pixel in the target viewpoint by calculating the geometric transformation. Then, the sampler performs a spatial transformation on the first image feature according to the generated sampling mesh. The sampler extracts pixel values from the first image feature and maps these values to the corresponding positions in the second image feature. Furthermore, to ensure image smoothness during the transformation process, interpolation techniques, such as bilinear interpolation, can be used.
[0122] Ultimately, the transformed second image features are represented in a unified coordinate system, eliminating geometric differences caused by different camera perspectives. By integrating the second image features from all camera perspectives, a complete vectorized map is generated. This map provides a consistent environmental representation, suitable for autonomous driving and advanced driver assistance systems, significantly improving the accuracy and reliability of environmental perception.
[0123] In some embodiments, besides combining a localization network, a mesh generator, and a sampler to perform spatial transformation on an image, other methods can also be used for spatial transformation of the image, and this embodiment does not particularly limit this. For example, matrix operations can be performed directly in the spatial transformation network, that is, geometric transformation parameters (such as affine transformation matrices) can be directly used to transform image features. For example, this can be implemented using a linear algebra library (such as NumPy), which maps features to a new coordinate system through matrix multiplication, thereby achieving spatial transformation of the image.
[0124] In a further example of the above implementation, the training method of the spatial transformation network can be as follows: extracting a first historical image feature from historical image data; inputting the first image feature into an initial spatial transformation network to obtain a second historical image feature after spatial transformation of the first historical image feature; training the initial spatial transformation network based on the loss calculation result between the predicted value of the vector map element obtained by predicting the second historical image feature and the true value of the map element, so as to obtain the spatial transformation network.
[0125] In this embodiment, first historical image features are extracted from historical image data. These features may include road image features captured in the same or similar environments in the past, such as lane lines, traffic signs, and other road elements. The feature extraction method is the same as that used in the first image feature extraction method in the previous embodiment to reduce discrepancies. These features will be used as input to the training data.
[0126] Next, an initial spatial transformation network is used, which can be an untrained model or a model pre-trained based on prior knowledge. By inputting the extracted first historical image features into the initial spatial transformation network, these features are transformed into second historical image features. This process simulates the transformation of image features from the original viewpoint to the target viewpoint in real-world applications. Based on the second historical image features, the network predicts vector map features, which may include road geometry, lane line positions, and other key map elements. The predicted vector map features are compared with known ground truth map features, and the loss between the predicted and ground truth values is calculated (using various loss functions, such as mean squared error (MSE) or cross-entropy loss). The ground truth values are typically derived from high-precision map data or manually labeled datasets. In this embodiment, the network weights and parameters are updated using the loss calculation results via backpropagation. This process continuously optimizes the network, enabling it to perform spatial transformations more accurately. The network can be iterated multiple times by repeating the above process, using different historical image data and features for training. The network gradually learns more accurate geometric transformation parameters, improving its performance in practical applications, and outputting the final spatial transformation network.
[0127] This training method uses the ground truth values of map features and the predicted values of vector map features to calculate the loss, gradually converging the BEV features to the ground feature features. This guides the training of geometric transformation parameters, completing the transformation from image features to the BEV feature space. The spatial transformation network trained in this way can provide reliable spatial transformation performance under different environments and conditions, supporting the high-precision map generation requirements of autonomous driving and advanced driver assistance systems.
[0128] In some embodiments, the map generation model may further include a feature fusion network, wherein the second image feature includes the image feature corresponding to the first image feature under all viewpoints after spatial transformation;
[0129] In the above step S302, generating a vectorized map based on the spatially transformed second image features includes:
[0130] According to the feature fusion network, the image features corresponding to the first image features under each viewpoint after spatial transformation are fused to obtain fused image features; wherein, the feature fusion network is used to perform feature fusion on the first image features under each viewpoint in the feature layer in the feature channel and / or feature space.
[0131] Based on the fusion features, vectorized map elements in the image data are identified, and a vectorized map is constructed based on the vectorized map elements.
[0132] In this embodiment, the map generation model includes a Feature Fusion Network (FFN) to fuse image features from different viewpoints. Compared to related technologies, this embodiment does not use camera parameters for BEV feature space mapping, but relies on geometric transformations learned by the STN. Therefore, it can use a two-layer convolutional network for feature fusion, expanding the receptive field and improving fusion accuracy.
[0133] For example, in the case of multiple cameras at an intersection, the input consists of images from multiple perspectives. After processing by the STN network, each image yields a set of spatially transformed image feature data. The feature fusion network can fuse image features from different perspectives. Unlike image stitching, this fusion is performed at the feature layer, fusing feature channels and feature space. For instance, through convolutional layers, each feature point can be fused to the features of all channels within the convolutional kernel range around that point (when the convolutional kernel is 5, the range is a square area with a side length of 5 centered at that point). This ensures that the feature exists at the corresponding location when predicting vector map elements.
[0134] As a further example, the feature fusion network is responsible for integrating the spatially transformed second image features. This feature fusion network can employ convolutional neural networks or other deep learning architectures to perform fusion at the feature layer. The feature fusion process can include fusion at the feature channel and feature space levels. Fusion at the feature channel level combines feature information from different perspectives to enhance the expressive power of the features; fusion at the feature space level ensures that spatial information from different perspectives is effectively integrated, forming a more complete understanding of the environment, thereby generating fused features. Through the feature fusion network, the spatially transformed image features from each perspective are fused to generate fused image features. These features contain information from multiple perspectives, providing a more comprehensive description of the environment. In other words, this embodiment improves the completeness and accuracy of the map through multi-view image fusion.
[0135] In some embodiments, the feature fusion network may employ an attention mechanism to enhance the effect of feature fusion, thereby further improving the robustness and accuracy of the model.
[0136] In one implementation, the map generation model may also include a map output head, which includes a classification branch and a point regression branch.
[0137] The steps described above for identifying vectorized map elements in the image data based on the fused image features specifically involve: inputting the fused image features into the map output header; performing classification prediction on the fused image features according to the classification branch to obtain the element category of the map element; performing location prediction on the fused image features according to the point regression branch to obtain the location coordinates of the map element; and identifying the vectorized map elements in the image data based on the element category and the location coordinates.
[0138] For example, the Map Output Head is designed with a classification branch and a point regression branch. The classification branch outputs an N*C dimensional vector representing the category scores of the N output features; the point regression branch outputs an N*K*2 dimensional vector representing the normalized BEV plane coordinates of the K points of the N output features. Here, N represents the number of output features, C is the number of feature categories supported by the model, and K represents the number of nodes describing the features.
[0139] Specifically, the main task of the classification branch is to classify and predict the fused image features to determine the category of each map element. This branch outputs an N*C dimensional vector, where N represents the number of output features and C is the number of feature categories supported by the model. That is, each feature has a corresponding category score vector, representing its probability of belonging to each category. The classification branch can be used to identify different types of map features, such as roads, arrows, and lane lines, thus providing foundational data for subsequent map creation and navigation.
[0140] The point regression branch is responsible for predicting the location of fused image features to determine the coordinates of map elements. This branch outputs an N*K*2 dimensional vector, where K represents the number of nodes describing the element, and each node has two coordinate values representing its position on the normalized BEV plane. Through the point regression branch, the shape and location of map elements can be accurately plotted.
[0141] By combining the outputs of the classification and point regression branches, the model can effectively identify vectorized map features in image data. This identification is not limited to static features but can also be extended to dynamic features such as traffic flow and pedestrian activity. The map generation model, which includes an STN network, a feature fusion network, and a map output head, is as follows: Figure 5 As shown, the output vector map (rendered image) of the intersection scene is as follows: Figure 6 As shown, the technical solution provided in this embodiment effectively solves the problem that in complex intersection scenarios, due to the limited field of view of the vehicle-mounted camera, it is difficult to fully capture the entire view of the intersection.
[0142] In some embodiments, the method is applied to an edge computing device, on which the map generation model is deployed; wherein the deployment method of the map generation model includes:
[0143] The map generation model is converted into a map generation model corresponding to the target format using a pre-set conversion tool; wherein the target format is compatible with the inference engine of the edge device.
[0144] The converted map generation model is deployed in the edge computing device, which is an embedded device that can be embedded in multiple cameras.
[0145] In this embodiment, by deploying the model to an edge computing device, which can be an embedded device that can be embedded in multiple cameras, a vectorized map can be generated (or updated) for the road segment (such as an intersection) corresponding to the multiple cameras when image data is collected from them. This generated vectorized map is then transmitted to a map generation server and combined with map data from other road segments to generate (or update) a comprehensive map. By deploying this map generation model on an edge computing device, data processing and storage can be performed locally, reducing computational complexity. Since data does not need to be transmitted over a network, the risk of network attacks and data tampering is also effectively reduced.
[0146] For example, the model can be deployed on an edge computing device. The deployment process can begin by converting the model from the PyTorch deep learning framework format to the Open Neural Network Exchange (ONNX) format, and then using a pre-built conversion tool (such as trtexec) to convert the ONNX model to an edge computing device-compatible inference engine format. Furthermore, to optimize inference efficiency, the model can be converted to 16-bit floating-point precision (FP16). Ultimately, in this embodiment, the map generation method can utilize approximately 1.6GB of GPU memory on the edge computing device, achieve an inference speed of 20 frames per second (FPS), and have an accuracy reduction of less than 1%.
[0147] In some embodiments, model compression and pruning techniques can be used to compress and prune the map generation model, further reducing the computational complexity of the model, improving inference speed, and making it more suitable for deployment on edge devices.
[0148] Figure 7 This application provides a map generation apparatus, such as... Figure 7 As shown, the device 700 includes a first processing module 701 and a second processing module 702, wherein,
[0149] The first processing module 701 is configured to extract a first image feature from the image data after acquiring image data from multiple cameras; wherein the image data includes road images from multiple perspectives, and the first image feature includes first image features from all perspectives.
[0150] The second processing module 702 is configured to perform spatial transformation processing on the first image features according to a preset map generation model, and generate a vectorized map based on the spatially transformed second image features.
[0151] The map generation model includes geometric transformation parameters, which are used to spatially transform the first image features from each viewpoint to the second image features from the target viewpoint.
[0152] In one embodiment, the map generation model includes a spatial transformation network for determining the geometric transformation parameters; the second processing module 702 includes:
[0153] The determining unit is configured to determine the corresponding geometric transformation parameters for the first image features under each viewpoint through the spatial transformation network, thereby obtaining the geometric transformation parameters of the first image features under each viewpoint.
[0154] The spatial transformation unit is configured to perform spatial transformation on the corresponding first image features according to the geometric transformation parameters of the first image features under each viewpoint, so as to obtain the spatially transformed second image features.
[0155] In one embodiment, the spatial transformation network includes a localization network, a mesh generator, and a sampler; wherein the geometric transformation parameters are determined based on the localization network after locating the first image features;
[0156] The spatial transformation unit is specifically configured as follows: for a first image feature under each viewpoint, a sampling grid for spatial mapping of the first image feature is generated by the grid generator based on the geometric transformation parameters. The sampling grid defines the mapping relationship between the pixels of the first image feature and the second image feature. The first image feature is sampled according to the sampling grid by the sampler to perform spatial transformation on the first image feature to obtain the second image feature. The sampling process includes extracting pixel values from the first image feature and mapping them to the corresponding positions of the second image feature.
[0157] In one embodiment, a training module for training the spatial transformation network is further included, the training module comprising:
[0158] The first extraction unit is configured to extract the first historical image feature from the historical image data;
[0159] The acquisition unit is configured to input the first image features into an initial spatial transformation network to obtain the second historical image features after spatial transformation of the first historical image features;
[0160] An optimization unit is configured to train the initial spatial transformation network using the loss calculation result between the predicted values of vector map elements obtained from the second historical image features and the true values of the map elements, in order to obtain the spatial transformation network.
[0161] In one embodiment, the map generation model includes a feature fusion network; the second image features include image features corresponding to the first image features from all viewpoints after spatial transformation; the second processing module 702 includes:
[0162] The first fusion unit is configured to fuse the image features corresponding to the first image features under each viewpoint after spatial transformation according to the feature fusion network to obtain fused image features; wherein, the feature fusion network is used to perform feature fusion on the first image features under each viewpoint at the feature layer in the feature channel and / or feature space.
[0163] The identification unit is configured to identify vectorized map elements in the image data based on the fusion features, and construct a vectorized map based on the vectorized map elements.
[0164] In one implementation, the map generation model further includes a map output head, which includes a classification branch and a point regression branch;
[0165] The identification unit is specifically configured to: input the fused image features into the map output head; perform classification prediction on the fused image features according to the classification branch to obtain the feature category of the map element; perform position prediction on the fused image features according to the point regression branch to obtain the position coordinates of the map element; and identify the vectorized map element in the image data according to the feature category and the position coordinates.
[0166] In one embodiment, the map generation model further includes a feature extraction network, which includes a backbone network and a feature pyramid network; the first processing module 701 includes:
[0167] The second extraction unit is configured to extract multi-scale features of the image data based on the backbone network and input the multi-scale features into the feature pyramid network;
[0168] The second fusion unit is configured to perform feature fusion on the multi-scale features according to the feature pyramid network to obtain the first image features.
[0169] In one embodiment, the method is applied to an edge computing device, on which the map generation model is deployed; wherein the deployment method of the map generation model includes:
[0170] The map generation model is converted into a map generation model corresponding to the target format using a pre-set conversion tool; wherein the target format is compatible with the inference engine of the edge device.
[0171] The converted map generation model is deployed in the edge computing device, which is an embedded device that can be embedded in multiple cameras.
[0172] The relevant explanations can be understood by referring to the relevant descriptions and effects corresponding to the steps in the method embodiments of this application, and will not be elaborated further here.
[0173] Figure 8 This is an electronic device provided in the embodiments of this application, such as... Figure 8 As shown, the electronic device 801 includes: a processor 801, and a memory 802 communicatively connected to the processor 801;
[0174] The memory 802 stores computer-executed instructions;
[0175] The processor 801 executes computer execution instructions stored in the memory 802 to implement the high-precision map generation method, wherein the memory 802 and the processor 801 are connected via a bus 803.
[0176] The relevant explanations can be understood by referring to the relevant descriptions and effects corresponding to the steps in the method embodiments of this application, and will not be elaborated further here.
[0177] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the map generation method corresponding to the above method embodiments.
[0178] The computer-readable storage medium can be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0179] The relevant explanations can be understood by referring to the relevant descriptions and effects corresponding to the steps in the method embodiments of this application, and will not be elaborated further here.
[0180] This application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the map generation method corresponding to the above method embodiments.
[0181] The relevant explanations can be understood by referring to the relevant descriptions and effects corresponding to the steps in the method embodiments of this application, and will not be elaborated further here.
[0182] This application also provides a chip, including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory to execute the map generation method corresponding to the above method embodiment.
[0183] The relevant explanations can be understood by referring to the relevant descriptions and effects corresponding to the steps in the method embodiments of this application, and will not be elaborated further here.
[0184] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0185] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0186] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A map generation method, characterized in that, include: After acquiring image data from multiple cameras, a first image feature is extracted from the image data; wherein, the image data includes road images from multiple perspectives, and the first image feature includes first image features from all perspectives; According to the preset map generation model, the first image features are spatially transformed, and a vectorized map is generated based on the spatially transformed second image features. The map generation model includes geometric transformation parameters, which are used to spatially transform the first image features from each viewpoint to the second image features from the target viewpoint.
2. The method according to claim 1, characterized in that, The map generation model includes a spatial transformation network for determining the geometric transformation parameters; The spatial transformation processing of the first image features includes: For the first image feature under each viewpoint, the spatial transformation network determines the corresponding geometric transformation parameters based on the first image feature to obtain the geometric transformation parameters of the first image feature under each viewpoint. Based on the geometric transformation parameters of the first image features under each viewpoint, spatial transformation is performed on the corresponding first image features to obtain the spatially transformed second image features.
3. The method according to claim 2, characterized in that, The spatial transformation network includes a localization network, a mesh generator, and a sampler; wherein, the geometric transformation parameters are determined based on the localization network after locating the first image features; The step of performing spatial transformation on the corresponding first image features according to the geometric transformation parameters of the first image features under each viewpoint to obtain the spatially transformed second image features includes: For each first image feature under each viewpoint, a sampling grid for spatial mapping of the first image feature is generated by the grid generator based on the geometric transformation parameters. The sampling grid defines the mapping relationship between the pixels of the first image feature and the second image feature. The sampler performs sampling processing on the first image features according to the sampling grid to spatially transform the first image features and obtain the second image features; wherein, the sampling processing includes extracting pixel values from the first image features and mapping them to the corresponding positions of the second image features.
4. The method according to claim 2 or 3, characterized in that, The training methods for the spatial transformation network include: Extract the first historical image features from the historical image data; The first historical image features are input into the initial spatial transformation network to obtain the second historical image features after spatial transformation of the first historical image features; Based on the predicted values of vector map elements obtained from the second historical image features and the loss calculation results between them and the true values of the map elements, the initial spatial transformation network is trained to obtain the spatial transformation network.
5. The method according to any one of claims 1-4, characterized in that, The map generation model includes a feature fusion network; the second image feature includes the image feature corresponding to the first image feature under all viewpoints after spatial transformation. The generation of a vectorized map based on the spatially transformed second image features includes: According to the feature fusion network, the image features corresponding to the first image features under each viewpoint after spatial transformation are fused to obtain fused image features; wherein, the feature fusion network is used to perform feature fusion on the first image features under each viewpoint in the feature layer in the feature channel and / or feature space. Based on the fusion features, vectorized map elements in the image data are identified, and a vectorized map is constructed based on the vectorized map elements.
6. The method according to claim 5, characterized in that, The map generation model also includes a map output head, which includes a classification branch and a point regression branch. The step of identifying vectorized map features in the image data based on the fused image features includes: The fused image features are input into the map output header; Based on the classification branch, the fused image features are classified and predicted to obtain the feature categories of map elements; Based on the point regression branch, the location of the fused image features is predicted to obtain the location coordinates of the map elements; Based on the feature category and the location coordinates, identify the vectorized map features in the image data.
7. The method according to any one of claims 1-4, characterized in that, The map generation model also includes a feature extraction network, which includes a backbone network and a feature pyramid network. The step of extracting the first image feature from the image data includes: Multi-scale features of the image data are extracted based on the backbone network, and the multi-scale features are input into the feature pyramid network; The first image feature is obtained by fusing the multi-scale features according to the feature pyramid network.
8. The method according to any one of claims 1-4, characterized in that, The method is applied to an edge computing device, on which the map generation model is deployed; wherein, the deployment method of the map generation model includes: The map generation model is converted into a map generation model corresponding to the target format using a pre-set conversion tool; wherein the target format is compatible with the inference engine of the edge device. The converted map generation model is deployed in the edge computing device, which is an embedded device that can be embedded in multiple cameras.
9. A map generation device, characterized in that, include: The first processing module is configured to extract a first image feature from the image data after acquiring image data from multiple cameras; wherein the image data includes road images from multiple perspectives, and the first image feature includes first image features from all perspectives. The second processing module is configured to perform spatial transformation processing on the first image features according to a preset map generation model, and generate a vectorized map based on the spatially transformed second image features. The map generation model includes geometric transformation parameters, which are used to spatially transform the first image features from each viewpoint to the second image features from the target viewpoint.
10. An electronic device / computer-readable storage medium / computer program product, characterized in that, The electronic device includes a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the electronic device to perform the map generation method according to any one of claims 1 to 8; And / or, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the map generation method as described in any one of claims 1-8; and / or, The computer program product includes a computer program that, when executed by a processor, implements the map generation method as described in any one of claims 1-8.