Model training method, heat map generation method and device

CN115273011BActive Publication Date: 2026-09-08BEIJING HORIZON ROBOTICS TECH RES & DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210923417.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-02
Publication Date
2026-09-08
Estimated Expiration
2042-08-02

AI Technical Summary

Technical Problem

[0003]为了解决采用上述方式热力图生成效率低的问题,提出了本公开

Benefits of technology

[0023] According to another aspect of the present disclosure, an electronic device is provided, comprising:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273011B_ABST
    Figure CN115273011B_ABST
Patent Text Reader

Abstract

A model training method, a heat map generation method and device are disclosed. The model training method comprises: obtaining multiple training images, the multiple training images being multiple images of different perspectives collected by multiple image acquisition devices arranged at different positions of a first movable device at the same time for an environment around the first movable device; generating, based on the multiple training images, a first heat map of a road target of a predetermined category in the environment via a neural network model; generating, based on position information of the first movable device at the time of collecting the multiple training images and a high-precision map, a second heat map of the road target of the predetermined category in the environment; determining a model loss value of the neural network model by comparing the first heat map and the second heat map; and training the neural network model based on the model loss value. The disclosed embodiments can improve the generation efficiency while ensuring the reliability of the generated heat map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to driving technology, and in particular to a model training method, a heatmap generation method, and an apparatus. Background Technology

[0002] Environmental perception is a key research area in fields such as autonomous driving and scene understanding. When performing environmental perception, there is a need to generate heat maps of specific road targets in some cases. To meet this need, a common approach is to provide images collected by cameras on vehicles and other mobile devices to a neural network model. The neural network model then generates the detection results of the specific road target based on these images. The detection results from multiple cameras are then fused together to generate the required heat map. Summary of the Invention

[0003] To address the issue of low efficiency in heatmap generation using the aforementioned methods, this disclosure is proposed. Embodiments of this disclosure provide a model training method, a heatmap generation method, and an apparatus.

[0004] According to one aspect of the present disclosure, a model training method is provided, comprising:

[0005] Multiple training images are acquired, which are multiple images from different perspectives of the environment around the first mobile device acquired at the same time by multiple image acquisition devices set at different locations of the first mobile device.

[0006] Based on multiple training images, a first heatmap of road surface targets of a predetermined category in the environment is generated via a neural network model;

[0007] Based on the location information of the first mobile device at the acquisition time of multiple training images and the high-precision map, a second heat map of road targets of a predetermined category in the environment is generated.

[0008] The model loss value of the neural network model is determined by comparing the first heatmap and the second heatmap.

[0009] The neural network model is trained based on the model loss value.

[0010] According to another aspect of the present disclosure, a method for generating a heatmap is provided, comprising:

[0011] Multiple environmental images are acquired simultaneously by multiple image acquisition devices located at different positions on a second mobile device, resulting in multiple environmental images, each corresponding to a different viewpoint.

[0012] Based on multiple environmental images, a heat map of road targets of a predetermined category in the environment surrounding the second mobile device is generated via a neural network model.

[0013] According to another aspect of the present disclosure, a model training apparatus is provided, comprising:

[0014] The first acquisition module is used to acquire multiple training images. The multiple training images are multiple images from different perspectives captured at the same time by multiple image acquisition devices set at different positions of the first mobile device, targeting the environment around the first mobile device.

[0015] The first generation module is used to generate a first heat map of road surface targets of a predetermined category in the environment based on multiple training images acquired by the first acquisition module via a neural network model.

[0016] The second generation module is used to generate a second heat map of road targets of a predetermined category in the environment based on the location information of the first mobile device at the acquisition time of multiple training images acquired by the first acquisition module, and a high-precision map.

[0017] The determination module is used to determine the model loss value of the neural network model by comparing the first heat map generated by the first generation module and the second heat map generated by the second generation module.

[0018] The training module is used to train the neural network model based on the model loss value determined by the determining module.

[0019] According to another aspect of the present disclosure, a heatmap generation apparatus is provided, comprising:

[0020] The second acquisition module is used to acquire environmental images acquired at the same time by multiple image acquisition devices set at different positions of the second mobile device, so as to obtain multiple environmental images, each of which corresponds to a different viewpoint.

[0021] The third generation module is used to generate a heat map of road surface targets of a predetermined category in the environment around the second mobile device based on multiple environmental images obtained by the second acquisition module via a neural network model.

[0022] According to another aspect of the present disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program for performing the above-described model training method or heatmap generation method.

[0023] According to another aspect of the present disclosure, an electronic device is provided, comprising:

[0024] processor;

[0025] Memory used to store the processor's executable instructions;

[0026] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the above-described model training method or heatmap generation method.

[0027] Based on the model training method, heatmap generation method, apparatus, computer-readable storage medium, and electronic device provided in the above embodiments of this disclosure, a first heatmap of road targets of a predetermined category in the environment surrounding a first mobile device can be generated based on multiple training images via a neural network model. A second heatmap of road targets of a predetermined category in the environment surrounding the first mobile device can be generated based on the location information of the first mobile device at the time of acquisition of the multiple training images and a high-precision map. The first heatmap can be considered as the predicted data of the neural network model, and the second heatmap can be considered as the real data. By comparing the first and second heatmaps, the model loss value of the neural network model can be determined, and the model loss value can be used for training the neural network model. Since the model training utilizes multiple images from different perspectives and a high-precision map, when the trained neural network model is used, the trained neural network model can directly and efficiently generate the corresponding heatmap based on images from multiple perspectives, thereby improving generation efficiency while ensuring the reliability of the generated heatmap.

[0028] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0029] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0030] Figure 1 This is a schematic flowchart of a model training method provided in an exemplary embodiment of this disclosure.

[0031] Figure 2 This is a flowchart illustrating a model training method provided in another exemplary embodiment of this disclosure.

[0032] Figure 3 This is a flowchart illustrating a model training method provided in yet another exemplary embodiment of this disclosure.

[0033] Figure 4This is a schematic diagram illustrating the generation principle of the first heat map in an exemplary embodiment of this disclosure.

[0034] Figure 5 This is a schematic flowchart of a heatmap generation method provided in an exemplary embodiment of this disclosure.

[0035] Figure 6 This is a schematic diagram of the structure of a model training apparatus provided in an exemplary embodiment of the present disclosure.

[0036] Figure 7 This is a schematic diagram of the structure of a model training apparatus provided in another exemplary embodiment of this disclosure.

[0037] Figure 8 This is a schematic diagram of the structure of a heat map generation apparatus provided in an exemplary embodiment of this disclosure.

[0038] Figure 9 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0039] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. It is obvious that the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0040] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0041] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.

[0042] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.

[0043] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.

[0044] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.

[0045] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.

[0046] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.

[0047] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.

[0048] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0049] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0050] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.

[0051] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.

[0052] Application Overview

[0053] When performing environmental perception, there is a need to generate heat maps of specific road targets in some cases, such as heat maps of intersections, zebra crossings, and road directional arrows.

[0054] Taking the need to generate intersection heatmaps as an example, a common approach to meet this need is to provide images captured by cameras installed on vehicles and other mobile devices to a neural network model. The neural network model then generates intersection detection results based on these images, fuses the intersection detection results from multiple cameras, and generates an intersection heatmap based on the fused results.

[0055] It is easy to see that generating intersection heatmaps using the above method requires processing images collected by different cameras separately and fusing the processing results from different cameras. Therefore, the above method suffers from low heatmap generation efficiency.

[0056] Exemplary methods

[0057] Figure 1 This is a schematic flowchart of a model training method provided in an exemplary embodiment of this disclosure. Figure 1 The method shown includes steps 110, 120, 130, 140 and 150, which are explained below.

[0058] Step 110: Acquire multiple training images. The multiple training images are multiple images from different perspectives of the environment around the first mobile device, which are acquired at the same time by multiple image acquisition devices set at different locations of the first mobile device.

[0059] Optionally, the first mobile device can be a vehicle; the image acquisition device can be a camera; the number of image acquisition devices set on the first mobile device can be four, and the four image acquisition devices can be set on the left front, right front, left rear, and right rear of the first mobile device respectively; there can be a one-to-one correspondence between multiple training images and multiple image acquisition devices.

[0060] Step 120: Based on multiple training images, generate a first heat map of road surface targets of a predetermined category in the environment via a neural network model.

[0061] Optionally, the neural network model involved in the embodiments of this disclosure can be an end-to-end neural network model; the predetermined category of road surface targets includes, but is not limited to, intersections, zebra crossings, road directional arrows, etc.

[0062] In step 120, multiple training images can be provided as input to the neural network model, which can then perform calculations to generate a first heat map of road targets of a predetermined category in the environment surrounding the first mobile device. Each pixel in the first heat map has a heat value, and a larger heat value indicates that the pixel is closer to the road target of the predetermined category. The maximum value of the heat value can be 1, and the minimum value can be 0.

[0063] Step 130: Based on the location information of the first mobile device at the acquisition time of multiple training images and the high-precision map, generate a second heat map of road targets of a predetermined category in the environment.

[0064] Optionally, the first mobile device may be equipped with a positioning device, such as a Global Positioning System (GPS). The location information of the first mobile device at the time of acquisition of multiple training images can be obtained by calling the positioning device.

[0065] Alternatively, high-precision maps can be created based on data collected by traditional data collection vehicles equipped with professional surveying equipment such as lidar, Global Navigation Satellite System (GNSS), and Inertial Measurement Unit (IMU). Compared to ordinary maps, high-precision maps have advantages such as higher accuracy, more data dimensions, and more precise positioning.

[0066] It should be noted that the first heatmap and the second heatmap can have the same image size. The pixels in the first heatmap can correspond one-to-one with the pixels in the second heatmap. Furthermore, similar to the first heatmap, each pixel in the second heatmap also has a heat value. The larger the heat value, the closer the pixel is to the road surface target of the predetermined category. The maximum value of the heat value can be 1, and the minimum value can be 0.

[0067] Step 140: Determine the model loss value of the neural network model by comparing the first heatmap and the second heatmap.

[0068] Since the first heatmap is generated by a neural network model, it can be considered as the predicted data of the neural network model. Since the second heatmap is generated based on the location information of the first mobile device at the time of acquisition of multiple training images and a high-precision map, and since both the location information and the high-precision map are objective and accurate data, the second heatmap can be considered as real data (or labeled data). In step 140, the model loss value of the neural network model can be determined by comparing the predicted data and the real data.

[0069] In one example, a first pixel exists in the first heatmap, and a second pixel corresponding to the first pixel exists in the second heatmap. Based on the heatmap values ​​of the first and second pixels, and a preset loss function, the loss value of the pixel pair consisting of the first and second pixels can be calculated. Since both the first and second heatmaps contain multiple pixels, multiple pixel pairs can be formed. Following the above method, multiple loss values ​​corresponding one-to-one with multiple pixel pairs can be obtained. By calculating the average of these multiple loss values, the model loss value of the neural network model can be obtained.

[0070] Optionally, the preset loss function can be the L1 loss function, the L2 loss function, or other loss functions; among them, the L1 loss function can also be called the Mean Abs Error (MAE) loss function, and the L2 loss function can also be called the Mean Square Error (MSE) loss function.

[0071] Step 150: Train the neural network model based on the model loss value.

[0072] In step 150, the model parameters of the neural network model can be adjusted based on the model loss value using the stochastic gradient descent method until the neural network model converges, thereby completing the training of the neural network model.

[0073] Based on the model training method provided in the above embodiments of this disclosure, a first heatmap of road targets of a predetermined category in the environment surrounding a first mobile device can be generated by a neural network model based on multiple training images. A second heatmap of road targets of a predetermined category in the environment surrounding the first mobile device can be generated based on the location information of the first mobile device at the time of acquisition of the multiple training images and a high-precision map. The first heatmap can be considered as the predicted data of the neural network model, and the second heatmap can be considered as the real data. By comparing the first and second heatmaps, the model loss value of the neural network model can be determined, so that the model loss value can be used for training the neural network model. Since the model training utilizes multiple images from different perspectives and a high-precision map, when the trained neural network model is used, the trained neural network model can directly and efficiently generate the corresponding heatmap based on images from multiple perspectives, thereby improving generation efficiency while ensuring the reliability of the generated heatmap.

[0074] exist Figure 1 Based on the illustrated embodiments, as Figure 2 As shown, step 130 includes steps 1302, 1304 and 1306.

[0075] Step 1302: Based on the location information of the first mobile device at the acquisition time of multiple training images and the size of the first preset area, determine the local map in the high-precision map.

[0076] Optionally, the first preset area size may include the first preset area length H1 and the first preset area width W1. H1 and W1 can both be 80 meters, 100 meters, 120 meters, etc., and will not be listed here. H1 and W1 can be the same or different.

[0077] In step 1302, given the location information of the first mobile device at the acquisition time of multiple training images, as well as the known H1 and W1, a rectangular area with a length of H1 and a width of W1 centered on the location represented by the location information can be determined on the high-precision map. The determined rectangular area can then be used as a local map in the high-precision map.

[0078] It is understandable that the category of each point in the local map is known, and these categories include, but are not limited to, intersection category, road category, zebra crossing category, road directional arrow category, lane line category, pedestrian category, curb category, etc.

[0079] Step 1304: Convert the local map into a target feature map containing semantic segmentation information.

[0080] It should be noted that the size of the heat map can be preset, for example, the length H2 and the width W2 of the heat map can be preset.

[0081] In step 1304, a feature map with length H2 and width W2 can be initialized, and then all points in the local map are mapped one by one to the feature map. The pixel value of each pixel in the feature map represents the category of the corresponding point in the local map, thereby forming a target feature map containing semantic segmentation information. The size of the target feature map can be represented as (H2, W2), and the semantic segmentation information includes the category of each pixel in the target feature map.

[0082] Step 1306: Based on the target feature map, generate a second heat map of road surface targets of a predetermined category in the environment.

[0083] Optionally, the second heatmap and the target feature map can have the same size.

[0084] In one specific implementation, a second heatmap of road surface targets of a predetermined category in the environment is generated based on the target feature map, including:

[0085] Based on semantic segmentation information, determine all pixels in the target feature map that belong to a predetermined category;

[0086] All pixels in the target feature map belonging to a predetermined category are divided into at least one set of target pixels, and each set of target pixels in the at least one set of target pixels corresponds to a road surface target of a predetermined category;

[0087] Based on the bounding boxes corresponding to each set of at least one set of target pixels determined on the target feature map, a second heatmap of road surface targets of a predetermined category in the environment is generated.

[0088] Optionally, the bounding box corresponding to any set of pixels can be the smallest rectangle that can enclose all pixels in the set.

[0089] Since semantic segmentation information includes the category of each pixel in the target feature map, referring to this information allows for convenient and quick filtering of all pixels in the target feature map that belong to a predetermined category. Next, for all pixels in the target feature map belonging to the predetermined category, a target pixel set can be partitioned. Each partitioned target pixel set corresponds to a road surface target of a predetermined category. For example, each partitioned target pixel set can represent a real intersection in the environment. Then, based on the bounding boxes corresponding to each target pixel set determined on the target feature map, a second heatmap can be generated.

[0090] In one example, a heatmap matrix of size (H2, W2) with all zeros can be initialized. For any set of target pixels on the target feature map, corresponding to the bounding box (let's assume it's the first bounding box), a second bounding box corresponding to the first bounding box can be determined on the heatmap matrix. Furthermore, pixels closer to the center of the second bounding box have larger heatmap values, and pixels farther from the center have smaller heatmap values. This forms a two-dimensional Gaussian heatmap on the heatmap matrix, which can then be used as the second heatmap. It's easy to see that the size of the second heatmap can be represented as (H2, W2), and correspondingly, the size of the first heatmap also needs to be (H2, W2).

[0091] In this implementation, by referring to semantic segmentation information, all pixels in the target feature map that belong to a predetermined category can be selected. All pixels in the target feature map that belong to a predetermined category can be considered as pixels that are truly helpful in generating the second heatmap. Then, by dividing these pixels, at least one set of target pixels can be determined, and the second heatmap can be generated based on the bounding box corresponding to each set of target pixels. This ensures the efficiency of generating the second heatmap.

[0092] In the embodiments of this disclosure, a local map can be extracted from a high-precision map by referring to the location information and the size of a first preset region. Then, the local map is converted into a target feature map containing semantic segmentation information so that the target feature map can be used to generate a second heatmap. This can effectively utilize the advantages of the high-precision map to generate an accurate and reliable second heatmap. In this way, when the second heatmap is used for model training, the accuracy and reliability of the training results can be well guaranteed.

[0093] In an optional example, all pixels in the target feature map belonging to a predetermined category are divided into at least one set of target pixels, including:

[0094] Determine a target bounding box on the target feature map that encloses all pixels in the target feature map that belong to a predetermined category;

[0095] In response to the actual region size corresponding to the target bounding box being less than or equal to the second preset region size, all pixels in the target feature map belonging to the predetermined category are assigned to the same target pixel set;

[0096] In response to the fact that the size of the real region corresponding to the target bounding box is greater than the second preset region size, based on the preset clustering algorithm, all pixels in the target feature map belonging to the predetermined category are divided into at least two target pixel sets.

[0097] Optionally, the target bounding box can be the smallest rectangle that can enclose all pixels of a predetermined category in the target feature map.

[0098] Optionally, the second preset area size may include a second preset area length H3 and a second preset area width W3. H3 and W3 can be determined by statistically analyzing the maximum size of road surface targets of a predetermined category in the real world. H3 and W3 may be the same or different.

[0099] Optionally, the preset clustering algorithm can be the K-means clustering algorithm; wherein, the K-means clustering algorithm is an iterative clustering analysis algorithm.

[0100] It should be noted that there can be a conversion relationship between the dimensions in the target feature map and the dimensions in the real world. After determining the target bounding box on the target feature map, the size of the real region corresponding to the target bounding box can be determined based on this conversion relationship, and it can be determined whether the size of the real region corresponding to the target bounding box is greater than the second preset region size.

[0101] The size of the real-world region corresponding to the target bounding box can be represented as (H4, W4). If H4 is less than or equal to H3, and W4 is less than or equal to W3, then it is reasonable to use all pixels enclosed by the target bounding box to represent a real road surface target of a predetermined category in the environment. Therefore, all pixels in the target feature map belonging to the predetermined category can be grouped into the same target pixel set. If H4 is greater than H3, and / or W4 is greater than H4, then it is unreasonable to use all pixels enclosed by the target bounding box to represent a real road surface target of a predetermined category in the environment. Therefore, based on a pre-defined clustering algorithm, all pixels in the target feature map belonging to the predetermined category can be divided into at least two target pixel sets.

[0102] In one specific implementation, based on a preset clustering algorithm, all pixels in the target feature map belonging to a predetermined category are divided into at least two sets of target pixels, including:

[0103] Using a pre-defined clustering algorithm, all pixels in the target feature map belonging to a predetermined category are divided into i sets of reference pixels; where i is greater than or equal to 2.

[0104] Determine the bounding boxes corresponding to each of the i reference pixel sets on the target feature map to obtain i bounding boxes;

[0105] In response to the fact that the actual region size corresponding to each of the i bounding boxes is less than or equal to the second preset region size, the set of i reference pixels is determined as the set of i target pixels;

[0106] In response to the fact that the actual region size corresponding to at least some of the bounding boxes in the i bounding boxes is greater than the second preset region size, the target feature map is divided into i+1 sets of pixels with the category of the predetermined category by a preset clustering algorithm, and at least two sets of target pixels are determined based on the i+1 sets of pixels.

[0107] It should be noted that the goal of the K-means clustering algorithm is to divide a sample set containing multiple samples into K non-overlapping clusters through iteration. The general steps of the K-means clustering algorithm are as follows: randomly select K samples from the sample set, and use each of the K selected samples as an initial cluster center to determine the K cluster centers; for any sample in the sample set, calculate the distance between it and each of the K cluster centers, and assign the sample to the cluster center closest to it. The cluster center and the sample assigned to it represent a cluster. Each time a sample is assigned, the cluster center is recalculated based on the existing samples in the cluster. This process is repeated until a certain termination condition is met. The termination condition can be that no (or a minimum number) samples are reassigned to different cluster centers, no (or a minimum number) cluster centers are changed, and the sum of squared errors is locally minimized.

[0108] In this implementation, we can first set i=2, which is equivalent to setting K in the K-means clustering algorithm to 2. Then, through the execution of the relevant steps of the K-means clustering algorithm, all pixels in the target feature map that belong to the predetermined category are divided into two reference pixel sets. Next, we can determine the bounding boxes corresponding to the two reference pixel sets on the target feature map to obtain two bounding boxes.

[0109] If the actual region size corresponding to each of the two bounding boxes is less than or equal to the second preset region size, it can be considered that all the pixels enclosed by the target bounding box are divided into two reference pixel sets, and each of the two reference pixel sets represents a real road surface target of a predetermined category in the environment. Therefore, each of the two reference pixel sets can be used as a target reference pixel set, thereby determining the two target reference pixel sets.

[0110] If the size of the real region corresponding to at least one of the two bounding boxes is larger than the second preset region size, it can be considered unreasonable to divide all the pixels enclosed by the target bounding box into two reference pixel sets, and each of the two reference pixel sets represents a real road surface target of a predetermined category in the environment. In this case, i=3, which is equivalent to updating K in the K-means clustering algorithm from 2 to 3. Next, by executing the relevant steps of the K-means clustering algorithm, all pixels in the target feature map belonging to the predetermined category are divided into three reference pixel sets. Then, the bounding boxes corresponding to the three reference pixel sets can be determined on the target feature map to obtain three bounding boxes. The subsequent process is similar to the process after obtaining two bounding boxes, and will not be repeated here.

[0111] In this implementation, we can start by dividing all pixels in the target feature map that belong to a predetermined category into two sets of reference pixels. If the attempt is successful, each set of reference pixels is used as a target pixel set. If the attempt is unsuccessful, we divide all pixels in the target feature map that belong to the predetermined category into a larger number of sets of reference pixels. In this way, we can quickly and reasonably divide all pixels in the target feature map that belong to the predetermined category into at least two sets of target pixels.

[0112] Of course, in the actual implementation, it is not necessary to try K in the order of 2, 3, 4, 5... Instead, first try whether K is 2. If K is 2, then divide all pixels in the target feature map that belong to the predetermined category into two target pixel sets. If K is not 2, then try whether K is 4. If K is not 4, then try whether K is 6, and so on. This will not be elaborated here.

[0113] In the embodiments of this disclosure, it can first be determined whether it is reasonable to divide all pixels in the target feature map of the predetermined category into a target pixel set based on the relationship between the actual size of the target bounding box and the size of the second preset region. If it is reasonable, all pixels in the target feature map of the predetermined category are directly divided into the same target pixel set. If it is not reasonable, all pixels in the target feature map of the predetermined category are efficiently and reasonably divided into at least two target pixel sets by using a preset clustering algorithm.

[0114] exist Figure 1 Based on the illustrated embodiments, as Figure 3 As shown, step 120 includes steps 1202, 1204 and 1206.

[0115] Step 1202: Extract features from multiple training images using a neural network model to obtain the first set of extracted features.

[0116] Optionally, the neural network model may include a feature extraction network. In step 1202, the feature extraction network can be used to extract features from multiple training images to obtain the extracted features corresponding to each of the multiple training images. These extracted features can form a first set of extracted features.

[0117] Step 1204: Convert the first set of extracted features into the second set of extracted features within the bird's-eye view space.

[0118] Alternatively, the bird's-eye view space can also be referred to as the BEV (Bird's Eye View) space.

[0119] In one specific implementation, converting the first set of extracted features into a second set of extracted features within the bird's-eye view space includes:

[0120] Based on the intrinsic and extrinsic parameters of the first image acquisition device, a first transformation matrix between the image coordinate system and the bird's-eye view coordinate system is determined; wherein, the first image acquisition device is any one of a plurality of image acquisition devices set at different locations of the first mobile device, and the training image corresponding to the first image acquisition device is the first training image.

[0121] The first extraction feature corresponding to the first training image in the first group of extracted features is transformed into the second extraction feature in the bird's-eye view space using the first transformation matrix.

[0122] Based on the second extracted feature, a second set of extracted features is determined.

[0123] It should be noted that the intrinsic parameters of the first image acquisition device are used to realize the transformation between the camera coordinate system and the pixel coordinate system, and the extrinsic parameters of the first image acquisition device are used to realize the transformation between the world coordinate system and the camera coordinate system. Furthermore, there is a transformation relationship between the pixel coordinate system and the image coordinate system, and there is also a transformation relationship between the world coordinate system and the bird's-eye view coordinate system. In this way, based on the intrinsic and extrinsic parameters of the first image acquisition device, the first transformation matrix between the image coordinate system and the bird's-eye view coordinate system can be determined efficiently and quickly.

[0124] After obtaining the first transformation matrix, the first extracted feature can be multiplied by the first transformation matrix to obtain the second extracted feature. Similarly, each extracted feature in the first group of extracted features can be transformed to the bird's-eye view space, thereby obtaining multiple extracted features in the bird's-eye view space. These extracted features can be combined to form the second group of extracted features.

[0125] By employing this implementation method, and referencing the intrinsic and extrinsic parameters of the image acquisition device, the conversion of extracted features into a bird's-eye view can be achieved efficiently and reliably.

[0126] Step 1206: Based on the second set of extracted features, generate a first heat map of road surface targets of a predetermined category in the environment via a neural network model.

[0127] Optionally, the neural network model may also include a feature fusion network and a prediction network. In step 1202, the feature fusion network can fuse all the extracted features in the second set of extracted features to obtain fused features, and the prediction network can make predictions based on the fused features to generate the first heatmap.

[0128] In the embodiments of this disclosure, by using a neural network model to extract features from multiple training images and transforming the extracted features into a bird's-eye view space, the neural network model can efficiently and reliably generate a first heatmap from a bird's-eye view based on the transformed second set of extracted features, so that the first heatmap can be used for training the neural network model.

[0129] In an optional example, you can first obtain v panoramic images from different perspectives (equivalent to multiple training images mentioned above), namely I1, I2, I3, ... I v A panoramic image with v perspectives can be represented as Figure 4 The form is v*3*h*w.

[0130] Based on I1, I2, I3, ... I v The process of generating the first heatmap may include:

[0131] (1) Using the feature extraction network in the neural network model, I1, I2, I3, ... I v Feature extraction is performed to obtain I1, I2, I3, ..., I v Each corresponding feature F is extracted, and I1, I2, I3, ..., I are used. v The corresponding first transformation matrix H BEV→img (Right now Figure 4 (homography matrix in the image), with I1, I2, I3, ... I v All features are transformed to the BEV space, resulting in transformed features BEV_F. The set of these BEV_F features constitutes the second set of extracted features mentioned above. The second set of extracted features can be represented as follows: Figure 4 The form of v*c′*h′*w′;

[0132] (2) I1, I2, I3, ... I v The second set of extracted features, composed of their respective corresponding bev_F values, is provided to the feature fusion network in the neural network model. The resulting fused features can be represented as follows: Figure 4 The form of 1*c″*h″*w″;

[0133] (3) The fused features are provided to the prediction network in the neural network model, and the prediction network can generate a predicted heatmap based on the fused features. pred (Equivalent to the first heatmap mentioned above), the predicted heatmap pred It can be represented as Figure 4 The form is 1*1*h″*w″.

[0134] Based on the first mobile device at I1, I2, I3, ... I vThe process for generating the second heat map based on the position information at the acquisition time and the high-definition map may include:

[0135] (1) Assuming that the target feature map obtained based on position information and a high-definition map is represented as BEVSeg, and the predetermined category is a sidewalk category, all pixel points belonging to the sidewalk category in BEVSeg can be extracted to obtain a point set P={(i,j)|0≤i<h, 0≤j<w} (equivalent to all pixel points of the predetermined category in the target feature map described above);

[0136] (2) First attempt to use point set P as a target pixel point set. If the attempt fails (equivalent to the case where the real area size corresponding to the target bounding box described above is larger than the second predetermined area size), the K-means clustering algorithm is performed on point set P. Since it is unknown how many intersections exist in the current scene, it may first be assumed that there are only n intersections, and new point sets P1, P2...Pn are formed after clustering (each new point set is equivalent to a reference pixel point set described above), and each new point set can determine the range of one intersection. If the range of a certain intersection is larger than the predetermined range (equivalent to the real area size corresponding to a certain bounding box among the i bounding boxes described above is larger than the second predetermined area size), it indicates that the assumption is unreasonable, and an attempt to cluster into n+1 clusters is performed, where n starts from 2, until the range of each intersection obtained by clustering is within the predetermined maximum range (equivalent to the real area sizes respectively corresponding to the i bounding boxes are all less than or equal to the second predetermined area size);

[0137] (3) Assuming that n intersections are obtained by clustering in the previous step, and new point sets P1, P2...Pn are formed, then the bounding boxes B1, B2...Bn of the n intersections can be determined;

[0138] (4) Initialize an all-zero heat map matrix with a shape of (H2, W2), and form a 2D Gaussian heat map Heatmap on the matrix for all the bounding boxes of intersections determined in (3) gt (equivalent to the second heat map described above).

[0139] So far, Heatmap as the first heat map pred and Heatmap as the second heat map gt have both been successfully obtained, then Heatmap pred and Heatmap gt can be subjected to L1 loss, and the formula used for L1 loss is:

[0140] Loss = |Heatmap pred -Heatmap g |

[0141] Based on L1 loss, the parameters of the neural network model can be updated, thereby completing the entire training process.

[0142] After the neural network model is trained, it can generate intersection heatmaps from a bird's-eye view by simply providing images from multiple perspectives as input, thereby improving the efficiency of heatmap generation.

[0143] Figure 5 This is a schematic flowchart of a heatmap generation method provided in an exemplary embodiment of this disclosure. Figure 5 The method shown includes steps 510 and 520, which are explained below.

[0144] Step 510: Obtain environmental images simultaneously captured by multiple image acquisition devices located at different positions on the second mobile device, resulting in multiple environmental images, each with a different viewing angle.

[0145] Optionally, the second mobile device can be a vehicle; the image acquisition device can be a camera; the number of image acquisition devices installed on the second mobile device can be four, and the four image acquisition devices can be respectively installed at the left front, right front, left rear, and right rear of the first mobile device; there can be a one-to-one correspondence between multiple environmental images and multiple image acquisition devices.

[0146] Step 520: Based on multiple environmental images, generate a heat map of road surface targets of a predetermined category in the environment surrounding the second mobile device via a neural network model.

[0147] It should be noted that the neural network model involved in step 520 can be a neural network model trained using the model training method described above.

[0148] In step 520, multiple environmental images can be provided as input to a neural network model, which can then perform calculations to generate a heat map of road surface targets of a predetermined category in the environment surrounding the second mobile device.

[0149] Based on the heat map generation method provided in the above embodiments of this disclosure, by simply providing images from multiple perspectives as input to the neural network model, the neural network model can efficiently and reliably generate heat maps from a bird's-eye view, thereby improving generation efficiency while ensuring the reliability of the generated heat maps.

[0150] Any model training method provided in this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any model training method provided in this disclosure can be executed by a processor, such as by a processor executing any model training method mentioned in this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.

[0151] Any of the heatmap generation methods provided in this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any of the heatmap generation methods provided in this disclosure can be executed by a processor, such as by a processor executing any of the heatmap generation methods mentioned in this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.

[0152] Exemplary device

[0153] Figure 6 This is a schematic diagram of the structure of a model training apparatus provided in an exemplary embodiment of the present disclosure. Figure 6 The apparatus shown includes a first acquisition module 610, a first generation module 620, a second generation module 630, a determination module 640, and a training module 650.

[0154] The first acquisition module 610 is used to acquire multiple training images. The multiple training images are multiple images from different perspectives captured at the same time by multiple image acquisition devices set at different positions of the first mobile device, targeting the environment around the first mobile device.

[0155] The first generation module 620 is used to generate a first heat map of road surface targets of a predetermined category in the environment based on multiple training images acquired by the first acquisition module 610 via a neural network model.

[0156] The second generation module 630 is used to generate a second heat map of road targets of a predetermined category in the environment based on the location information of the first mobile device at the acquisition time of multiple training images acquired by the first acquisition module 610 and the high-precision map.

[0157] The determination module 640 is used to determine the model loss value of the neural network model by comparing the first heat map generated by the first generation module 620 and the second heat map generated by the second generation module 630.

[0158] The training module 650 is used to train the neural network model based on the model loss value determined by the determination module 640.

[0159] exist Figure 6 Based on the illustrated embodiments, as Figure 7As shown, the second generation module 630 includes:

[0160] The determination submodule 6302 is used to determine a local map in the high-precision map based on the location information of the first mobile device at the acquisition time of multiple training images and the size of the first preset area.

[0161] The transformation submodule 6304 is used to transform the local map determined by the determining submodule 6302 into a target feature map containing semantic segmentation information;

[0162] The first generation submodule 6306 is used to generate a second heat map of road surface targets of a predetermined category in the environment based on the target feature map obtained by the transformation submodule 6304.

[0163] In one optional example, the first generated submodule 6306 includes:

[0164] The first determining unit is used to determine all pixels in the target feature map that belong to a predetermined category based on the semantic segmentation information carried by the target feature map transformed by the transformation submodule 6304.

[0165] The partitioning unit is used to divide all pixels of a predetermined category in the target feature map determined by the first determining unit into at least one set of target pixels, and each set of target pixels in the at least one set of target pixels corresponds to a road surface target of a predetermined category;

[0166] The generation unit is used to generate a second heatmap of road targets of a predetermined category in the environment based on the bounding boxes corresponding to each set of at least one set of target pixels obtained by the partitioning unit on the target feature map.

[0167] In one optional example, the division of units includes:

[0168] A sub-unit is defined to determine a target bounding box on the target feature map that surrounds all pixels of a predetermined category in the target feature map defined by the first defining unit.

[0169] The first dividing subunit is used to divide all pixels in the target feature map that belong to the predetermined category into the same target pixel set in response to the determination subunit determining that the actual region size corresponding to the target bounding box is less than or equal to the second preset region size.

[0170] The second partitioning subunit is used to divide all pixels in the target feature map that belong to a predetermined category into at least two sets of target pixels based on a preset clustering algorithm in response to the determination subunit determining that the size of the real region corresponding to the target bounding box is greater than the second preset region size.

[0171] In one optional example, the second dividing subunit is specifically used for:

[0172] Using a pre-defined clustering algorithm, all pixels in the target feature map belonging to a predetermined category are divided into i reference pixel sets, where i is greater than or equal to 2. Bounding boxes are determined for each of the i reference pixel sets in the target feature map, resulting in i bounding boxes. In response to the fact that the actual region size corresponding to each of the i bounding boxes is less than or equal to a second pre-defined region size, the i reference pixel sets are determined as i target pixel sets. In response to the fact that the actual region size corresponding to at least some of the i bounding boxes is greater than the second pre-defined region size, using the pre-defined clustering algorithm, all pixels in the target feature map belonging to the predetermined category are divided into i+1 pixel sets, and based on the i+1 pixel sets, at least two target pixel sets are determined.

[0173] exist Figure 6 Based on the illustrated embodiments, as Figure 7 As shown, the first generation module 620 includes:

[0174] The acquisition submodule 6202 is used to extract features from multiple training images acquired by the first acquisition module 610 through a neural network model to obtain a first set of extracted features;

[0175] The conversion submodule 6204 is used to convert the first set of extracted features obtained by the acquisition submodule 6202 into a second set of extracted features in the bird's-eye view space.

[0176] The second generation module 6206 is used to generate a first heat map of road surface targets of a predetermined category in the environment based on the second set of extracted features obtained by the transformation submodule 6204 via a neural network model.

[0177] In one optional example, conversion module 6204 includes:

[0178] The second determining unit is used to determine a first transformation matrix between the image coordinate system and the bird's-eye view coordinate system based on the intrinsic and extrinsic parameters of the first image acquisition device; wherein, the first image acquisition device is any one of a plurality of image acquisition devices set at different locations of the first mobile device, and the training image corresponding to the first image acquisition device is the first training image.

[0179] The transformation unit is used to convert the first extracted features corresponding to the first training image in the first set of extracted features obtained by the acquisition submodule 6202 into the second extracted features in the bird's-eye view space through the first transformation matrix determined by the second determining unit.

[0180] The third determining unit is used to determine the second set of extracted features based on the second extracted features obtained by the transformation unit.

[0181] Figure 8 This is a schematic diagram of the structure of a heat map generation apparatus provided in an exemplary embodiment of this disclosure. Figure 8 The apparatus shown includes a second acquisition module 810 and a third generation module 820.

[0182] The second acquisition module 810 is used to acquire environmental images acquired at the same time by multiple image acquisition devices set at different positions of the second mobile device, so as to obtain multiple environmental images, each of which has a different viewing angle.

[0183] The third generation module 820 is used to generate a heat map of road surface targets of a predetermined category in the environment around the second mobile device based on multiple environmental images obtained by the second acquisition module 810 via a neural network model.

[0184] Exemplary electronic devices

[0185] Below, for reference Figure 9 This describes an electronic device according to embodiments of the present disclosure. The electronic device may be either or both of a first device and a second device, or a standalone device independent of them, which may communicate with the first device and the second device to receive acquired input signals from them.

[0186] Figure 9 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0187] like Figure 9 As shown, the electronic device 900 includes one or more processors 910 and memory 920.

[0188] The processor 910 may be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device 900 to perform desired functions.

[0189] The memory 920 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium. The processor 910 may execute the program instructions to implement the model training method or heatmap generation method of the various embodiments of this disclosure described above. Additionally, the processor 910 may also execute the program instructions to implement other desired functions.

[0190] In one example, the electronic device 900 may also include an input device 930 and an output device 940, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0191] For example, when the electronic device 900 is a first device or a second device, the input device 930 may be a microphone or a microphone array. When the electronic device 900 is a standalone device, the input device 930 may be a communication network connector for receiving acquired input signals from the first device and the second device.

[0192] In addition, the input device 930 may also include, for example, a keyboard, a mouse, etc. The output device 940 can output various information to the outside. The output device 940 may include, for example, a monitor, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0193] Of course, for the sake of simplicity, Figure 9 Only some of the components of the electronic device 900 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 900 may include any other suitable components depending on the specific application.

[0194] Exemplary computer program products and computer-readable storage media

[0195] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products, including computer program instructions that, when executed by a processor, cause the processor to perform the steps in the model training methods or heatmap generation methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.

[0196] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0197] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps of the model training method or heatmap generation method according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.

[0198] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0199] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0200] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0201] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0202] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.

[0203] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.

[0204] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0205] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A model training method, comprising: Multiple training images are acquired, which are multiple images from different perspectives of the environment around the first mobile device acquired at the same time by multiple image acquisition devices set at different locations of the first mobile device. Based on multiple training images, a first heatmap of road surface targets of a predetermined category in the environment is generated via a neural network model; Based on the location information of the first mobile device at the acquisition time of multiple training images and the high-precision map, a second heat map of road targets of a predetermined category in the environment is generated. The model loss value of the neural network model is determined by comparing the first heatmap and the second heatmap. The neural network model is trained based on the model loss value; The step of generating a second heatmap of road surface targets of a predetermined category in the environment based on the location information of the first mobile device at the acquisition time of multiple training images and a high-precision map includes: Based on the location information of the first mobile device at the acquisition time of multiple training images and the size of the first preset area, a local map in the high-precision map is determined; The local map is converted into a target feature map containing semantic segmentation information; Based on the target feature map, a second heat map of road surface targets of a predetermined category in the environment is generated.

2. The method according to claim 1, wherein, The step of generating a second heatmap of road surface targets of a predetermined category in the environment based on the target feature map includes: Based on the semantic segmentation information, all pixels in the target feature map that belong to a predetermined category are determined; The target feature map is divided into at least one set of target pixels, and each set of target pixels corresponds to a road surface target of the predetermined category. Based on the bounding boxes corresponding to each set of at least one set of target pixels determined on the target feature map, a second heat map of the road surface target of the predetermined category in the environment is generated.

3. The method according to claim 2, wherein, The step of dividing all pixels in the target feature map belonging to the predetermined category into at least one set of target pixels includes: Determine a target bounding box on the target feature map that surrounds all pixels in the target feature map that belong to the predetermined category; In response to the actual region size corresponding to the target bounding box being less than or equal to the second preset region size, all pixels in the target feature map belonging to the predetermined category are assigned to the same target pixel set; In response to the fact that the actual region size corresponding to the target bounding box is greater than the second preset region size, based on a preset clustering algorithm, all pixels in the target feature map belonging to the predetermined category are divided into at least two target pixel sets.

4. The method according to claim 3, wherein, The method based on a preset clustering algorithm divides all pixels in the target feature map belonging to the predetermined category into at least two sets of target pixels, including: The preset clustering algorithm is used to divide all pixels in the target feature map that belong to the predetermined category into i sets of reference pixels; where i is greater than or equal to 2. On the target feature map, determine the bounding boxes corresponding to each of the i sets of reference pixels to obtain i bounding boxes; In response to the fact that the actual region size corresponding to each of the i bounding boxes is less than or equal to the second preset region size, the i reference pixel set is determined as the i target pixel set; In response to the fact that the actual region size corresponding to at least some of the bounding boxes in the i bounding boxes is greater than the second preset region size, the preset clustering algorithm is used to divide all pixels in the target feature map belonging to the predetermined category into i+1 pixel sets, and based on the i+1 pixel sets, at least two target pixel sets are determined.

5. The method according to claim 1, wherein, The step of generating a first heatmap of road surface targets of a predetermined category in the environment based on multiple training images via a neural network model includes: The neural network model is used to extract features from multiple training images to obtain a first set of extracted features; The first set of extracted features is converted into a second set of extracted features within the bird's-eye view space; Based on the second set of extracted features, a first heat map of road surface targets of a predetermined category in the environment is generated via the neural network model.

6. The method according to claim 5, wherein, The step of converting the first set of extracted features into a second set of extracted features within the bird's-eye view space includes: Based on the intrinsic and extrinsic parameters of the first image acquisition device, a first transformation matrix between the image coordinate system and the bird's-eye view coordinate system is determined; wherein, the first image acquisition device is any one of a plurality of image acquisition devices set at different locations of the first mobile device, and the training image corresponding to the first image acquisition device is the first training image. The first extraction feature corresponding to the first training image in the first group of extracted features is converted into the second extraction feature in the bird's-eye view space using the first transformation matrix. Based on the second extracted features, a second set of extracted features is determined.

7. A method for generating a heatmap, comprising: Multiple environmental images are acquired simultaneously by multiple image acquisition devices located at different positions on a second mobile device, resulting in multiple environmental images, each corresponding to a different viewpoint. Based on multiple environmental images, a heat map of road targets of a predetermined category in the environment surrounding the second mobile device is generated via a neural network model; The neural network model is obtained by training using any one of the model training methods described in claims 1-6.

8. A model training device, comprising: The first acquisition module is used to acquire multiple training images. The multiple training images are multiple images from different perspectives captured at the same time by multiple image acquisition devices set at different positions of the first mobile device, targeting the environment around the first mobile device. The first generation module is used to generate a first heat map of road surface targets of a predetermined category in the environment based on multiple training images acquired by the first acquisition module via a neural network model. The second generation module is used to generate a second heat map of road targets of a predetermined category in the environment based on the location information of the first mobile device at the acquisition time of multiple training images acquired by the first acquisition module, and a high-precision map. The second generation module includes: a determination submodule, used to determine a local map in the high-precision map based on the location information of the first mobile device at the acquisition time of multiple training images and a first preset area size; a conversion submodule, used to convert the local map into a target feature map containing semantic segmentation information; and a first generation submodule, used to generate a second heat map of road surface targets of a predetermined category in the environment based on the target feature map. The determination module is used to determine the model loss value of the neural network model by comparing the first heat map generated by the first generation module and the second heat map generated by the second generation module. The training module is used to train the neural network model based on the model loss value determined by the determining module.

9. A heat map generating apparatus, comprising: The second acquisition module is used to acquire environmental images acquired at the same time by multiple image acquisition devices set at different positions of the second mobile device, so as to obtain multiple environmental images, each of which corresponds to a different viewpoint. The third generation module is used to generate a heat map of road surface targets of a predetermined category in the environment around the second mobile device based on multiple environmental images obtained by the second acquisition module via a neural network model. The neural network model is obtained by training using any one of the model training methods described in claims 1-6.

10. A computer-readable storage medium storing a computer program for performing the model training method according to any one of claims 1-6 or the heatmap generation method according to claim 7.

11. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the model training method according to any one of claims 1-6 or to execute the heatmap generation method according to claim 7.

Citation Information

Patent Citations

  • Target detection network training method and vehicle detection method

    CN113947188A

  • Map updating method and map updating device

    CN114372068A

  • Target identification network training method, device and system

    CN114821242A