Map generation method and device, vehicle and storage medium
By acquiring and integrating the feature information of the local navigation map and bird's aerial view of the autonomous driving vehicle, a more accurate local map is generated, which solves the problem of insufficient sensor perception capabilities and improves the road environment perception and driving safety of autonomous driving vehicles.
Patent Information
- Application Number
- CN202410102354.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-24
- Publication Date
- 2025-07-25
AI Technical Summary
Under complex operating conditions, the sensor perception model of autonomous driving vehicles is easily blocked by obstacles, resulting in the inability to accurately sense the road environment.
Obtain the local navigation map and bird's-eye view around the vehicle, obtain the respective feature maps through the feature extraction network, and determine the pixel point correspondence relationship, fuse the feature values to generate the fusion feature map, and conduct road element detection to generate the target local map.
Through the integration of local navigation maps and sensor information, a more accurate target local map is generated, which improves the vehicle's understanding and judgment of road conditions and improves driving safety.
Smart Images

Figure CN120368954A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of autonomous driving technology, and in particular, to a method and apparatus for generating a map, a vehicle, and a storage medium. Background Art
[0002] The local map perception model from a bird's-eye view can sense the road information around the autonomous vehicle in real time through sensors such as panoramic cameras and lidar, including lane lines, crosswalks, and the topological structure of the road. This perception model can reduce the dependence of the autonomous vehicle on high-precision maps.
[0003] However, in complex working conditions, such as when there are many surrounding obstacles, the sensor information is easily affected by factors such as obstacle occlusion, resulting in the model being unable to perceive a reasonable road environment. Summary of the Invention
[0004] The present disclosure aims to solve at least one of the technical problems in the related art to some extent.
[0005] A first aspect embodiment of the present disclosure provides a method for generating a map, including:
[0006] Obtain a local navigation map and a bird's-eye view around the vehicle, where the bird's-eye view is generated based on the vehicle perimeter information collected by the sensors on the vehicle;
[0007] Extract features from the bird's-eye view and the local navigation map respectively to obtain a first feature map corresponding to the bird's-eye view and a second feature map corresponding to the local navigation map;
[0008] Determine multiple second pixel points corresponding to each first pixel point in the first feature map from the second feature map;
[0009] Fuse the first feature value corresponding to each first pixel point in the first feature map with the second feature values of the corresponding multiple second pixel points in the second feature map to obtain a fused feature map;
[0010] Perform road element detection on the fused feature map to obtain a road element detection result;
[0011] Generate a target local map based on the road element detection result.
[0012] A second aspect embodiment of the present disclosure provides a map generation apparatus, including:
[0013] A first acquisition module, configured to obtain a local navigation map and a bird's-eye view around the vehicle, where the bird's-eye view is generated based on the vehicle perimeter information collected by the sensors on the vehicle;
[0014] A second acquisition module, configured to respectively perform feature extraction on the bird's-eye view and the local navigation map to obtain a first feature map corresponding to the bird's-eye view and a second feature map corresponding to the local navigation map;
[0015] A determination module, configured to determine, from the second feature map, a plurality of second pixel points corresponding to each first pixel point in the first feature map;
[0016] A fusion module, configured to fuse a first feature value corresponding to each first pixel point in the first feature map with second feature values of the corresponding plurality of second pixel points in the second feature map to obtain a fused feature map;
[0017] A detection module, configured to perform road element detection on the fused feature map to obtain a road element detection result;
[0018] A generation module, configured to generate a target local map based on the road element detection result.
[0019] An embodiment of the third aspect of the present disclosure provides a vehicle, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to: perform steps of the map generation method as proposed in the embodiment of the first aspect of the present disclosure.
[0020] An embodiment of the fourth aspect of the present disclosure provides a computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enabling the mobile terminal to execute a map generation method, the method including:
[0021] Obtain a local navigation map and a bird's-eye view around the vehicle, wherein the bird's-eye view is generated based on vehicle perimeter information collected by sensors on the vehicle;
[0022] Respectively perform feature extraction on the bird's-eye view and the local navigation map to obtain a first feature map corresponding to the bird's-eye view and a second feature map corresponding to the local navigation map;
[0023] Determine, from the second feature map, a plurality of second pixel points corresponding to each first pixel point in the first feature map;
[0024] Fuse a first feature value corresponding to each first pixel point in the first feature map with second feature values of the corresponding plurality of second pixel points in the second feature map to obtain a fused feature map;
[0025] Perform road element detection on the fused feature map to obtain a road element detection result;
[0026] Generate a target local map based on the road element detection result.
[0027] The map generation method, device, vehicle, and storage medium provided by the present disclosure have the following beneficial effects:
[0028] In the embodiments of the present disclosure, first, a local navigation map and an aerial view around the vehicle are obtained, and feature extraction is respectively performed on the aerial view and the local navigation map to obtain a first feature map corresponding to the aerial view and a second feature map corresponding to the local navigation map. Then, from the second feature map, a plurality of second pixel points corresponding to each first pixel point in the first feature map are determined, and the first feature value corresponding to each first pixel point in the first feature map is fused with the second feature values of the corresponding plurality of second pixel points in the second feature map to obtain a fused feature map. Finally, road element detection is performed on the fused feature map to obtain a road element detection result, and based on the road element detection result, a target local map is generated. Thus, the information provided by the local navigation map can be fused with the information collected by the vehicle sensors, and the global information provided by the local navigation map can effectively make up for the problem of insufficient sensing ability of the sensors on the vehicle, so that a more accurate target local map can be generated, improving the vehicle's understanding and judgment ability of the road conditions and enhancing the safety of vehicle driving.
[0029] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The above and / or additional aspects and advantages of the present disclosure will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0031] Figure 1 is a flowchart of a map generation method provided by an embodiment of the present disclosure;
[0032] Figure 2 is a flowchart of a map generation method provided by another embodiment of the present disclosure;
[0033] Figure 3 is a flowchart of a method for obtaining a fused feature map provided by an embodiment of the present disclosure;
[0034] Figure 4 is a structural diagram of a map generation device provided by another embodiment of the present disclosure;
[0035] Figure 5 shows a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as a limitation of the present disclosure.
[0037] The map generation method, device, vehicle, and storage medium according to embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0038] Figure 1 It is a schematic flowchart of a map generation method provided by an embodiment of the present disclosure.
[0039] In the embodiments of the present disclosure, the map generation method is configured in a map generation device as an example. The map generation device can be applied to any vehicle so that the vehicle can perform the map generation function.
[0040] As Figure 1 shown, the map generation method may include the following steps:
[0041] Step 101: Obtain a local navigation map and a bird's-eye view around the vehicle, where the bird's-eye view is generated based on the vehicle perimeter information collected by sensors on the vehicle.
[0042] The vehicle may be an autonomous vehicle.
[0043] The local navigation map may include map elements such as lane areas, crosswalk areas, traffic lights, etc. In some possible implementation manners, the map elements may be rendered within a local range centered on the host vehicle to obtain an image corresponding to the rasterized local navigation map.
[0044] In some possible implementation manners, the current position of the vehicle may be determined by a high-precision positioning system, and then a request is sent to the navigation map for the local navigation map around the vehicle.
[0045] The sensors on the vehicle may include panoramic cameras, lidar, etc., and the present disclosure does not limit this.
[0046] In some possible implementation manners, multiple panoramic cameras with different perspectives may be used to collect environmental images around the vehicle. Each panoramic camera is used to collect a panoramic image of one perspective, and the multiple panoramic cameras cover the environmental range around the vehicle. Each panoramic camera defines its own camera perspective coordinate system, and forms its own camera perspective space through its respective camera perspective coordinate system. The panoramic image collected by each panoramic camera is an image in the corresponding camera perspective space.
[0047] In one embodiment, multiple panoramic cameras can collect multiple panoramic images from different perspectives in real time, such as Image 1, 2... N, and send the collected panoramic images to the electronic device in real time. In this way, the panoramic images obtained by the electronic device can represent the real situation of the vehicle's surrounding environment at the current moment, and then multiple panoramic images can be converted into a bird's-eye view based on the BEV perception algorithm.
[0048] In one embodiment, the vehicle may include 6 panoramic cameras. The 6 panoramic cameras are respectively arranged at the front, left front, right front, rear, left rear, and right rear of the carrier. The present disclosure does not limit this.
[0049] Step 102: Extract features from the bird's-eye view and the local navigation map respectively to obtain a first feature map corresponding to the bird's-eye view and a second feature map corresponding to the local navigation map.
[0050] In some possible implementation manners, a feature extraction network (such as a convolutional neural network, a Transformer network, etc.) may be used to extract features from the local navigation map and the bird's-eye view.
[0051] It should be noted that the same feature extraction network may be used to extract features from the bird's-eye view and the local navigation map, or different feature extraction networks may be used to extract features from the bird's-eye view and the local navigation map respectively.
[0052] In some possible implementation manners, the BEV perception algorithm may also be used to directly obtain the first feature map.
[0053] Step 103: Determine a plurality of second pixel points corresponding to each first pixel point in the first feature map from the second feature map.
[0054] Optionally, based on the position of the first pixel point in the first feature map, a reference pixel point with the same position as the first pixel point in the second feature map may be determined, and then the reference pixel point and a preset number of pixel points around the reference pixel point are determined as the plurality of second pixel points corresponding to the first pixel point.
[0055] Optionally, based on the Deformable Attention (DAN) mechanism, a plurality of second pixel points corresponding to each first pixel point in the first feature map may be determined from the second feature map. Thereby, it can be avoided that the same pixel point in the first feature map and the second feature map corresponds to different information around the vehicle, and the plurality of second pixel points corresponding to the first pixel point can be accurately determined.
[0056] Optionally, before determining multiple second pixel points corresponding to each first pixel point in the first feature map from the second feature map, the second feature map may also be scaled based on the scale of the first feature map so that the scaled second feature map has the same scale as the first feature map, thereby accurately determining multiple second pixel points corresponding to the first pixel point, and then the first feature map and the second feature map can be better fused.
[0057] Step 104: Fuse the first feature values corresponding to each first pixel point in the first feature map with the second feature values of the corresponding multiple second pixel points in the second feature map to obtain a fused feature map.
[0058] In some embodiments, the second feature values of multiple second pixel points in the second feature map may be fused first, and the fused feature values may be fused with the first feature values to obtain the target feature value corresponding to the first pixel point.
[0059] In some embodiments, the weight value corresponding to the first pixel point and the weight value corresponding to each second pixel point may also be determined respectively, and then based on the weight value corresponding to each pixel point, the corresponding feature value of each pixel point is weighted to obtain the target feature value corresponding to the first pixel point.
[0060] For example, the first pixel point corresponds to four second pixel points, and the weight values corresponding to the first pixel point and each second pixel point may be the same, both being 0.2. Alternatively, the weight of the first pixel point is greater than the weight of each second pixel point. For example, the weight of the first pixel point is 0.5, and the weight corresponding to each second pixel point is 0.125.
[0061] It should be noted that the present disclosure does not limit the weights corresponding to the first pixel point and each second pixel point. The weight values corresponding to each second pixel point may be the same or different.
[0062] Step 105: Perform road element detection on the fused feature map to obtain a road element detection result.
[0063] Optionally, road elements may include elements such as lane lines, arrows, zebra crossings, road edges, drivable areas, etc.
[0064] Optionally, the road element detection result may include information such as the type, size, and position of the road element.
[0065] In some possible implementation manners, a pre-trained multi-head road element detection model may be used to perform road element detection on the fused feature map. That is, the fused feature map may be input into each road element detection head to obtain the road element detection results output by each detection head, where each detection head may be used to detect different types of road elements.
[0066] Step 106: Generate a target local map based on the road element detection result.
[0067] In the embodiments of the present disclosure, after obtaining the road element detection result, each road element in the road element detection result can be fitted to obtain a target local map around the vehicle.
[0068] In the embodiments of the present disclosure, first, obtain a local navigation map and a bird's-eye view around the vehicle, and respectively perform feature extraction on the bird's-eye view and the local navigation map to obtain a first feature map corresponding to the bird's-eye view and a second feature map corresponding to the local navigation map. Then, determine multiple second pixel points corresponding to each first pixel point in the first feature map from the second feature map, and fuse the first feature value corresponding to each first pixel point in the first feature map with the second feature value of the corresponding multiple second pixel points in the second feature map to obtain a fused feature map. Finally, perform road element detection on the fused feature map to obtain a road element detection result, and generate a target local map based on the road element detection result. Thus, the information provided by the local navigation map can be fused with the information collected by the vehicle sensors. Through the global information provided by the local navigation map, the problem of insufficient perception ability of the sensors on the vehicle can be effectively compensated, so that a more accurate target local map can be generated, the vehicle's understanding and judgment ability of the road conditions can be improved, and the safety of vehicle driving can be improved.
[0069] Figure 2 It is a schematic flowchart of a map generation method provided by an embodiment of the present disclosure. As Figure 2 shown, the map generation method may include the following steps:
[0070] Step 201: Obtain a local navigation map and a bird's-eye view around the vehicle, where the bird's-eye view is generated based on the vehicle information collected by the sensors on the vehicle.
[0071] Step 202: Respectively perform feature extraction on the bird's-eye view and the local navigation map to obtain a first feature map corresponding to the bird's-eye view and a second feature map corresponding to the local navigation map.
[0072] For the specific implementation forms of step 201 and step 202, reference may be made to the detailed descriptions in other embodiments of the present disclosure, and details are not described herein again.
[0073] Step 203: Perform a linear transformation on the first feature value corresponding to the first pixel point in the first feature map to obtain a second feature value.
[0074] Optionally, a linear transformation network can be used to perform a linear transformation on the first feature value.
[0075] In some embodiments, the linear transformation network may be a recurrent neural network (RNN), a long short-term memory network (LSTM), etc., and the present disclosure does not limit this.
[0076] Step 204: Input the third eigenvalue into the offset prediction network to obtain multiple offsets.
[0077] Among them, the number of offsets can be set in advance in the offset prediction network and the weight prediction network in the form of hyperparameters.
[0078] In some embodiments, the offset prediction network may be a simple convolutional neural network (CNN), or a more complex network structure, such as a self-attention network or a recurrent neural network (RNN), etc. The present disclosure does not limit this.
[0079] Step 205: Based on the multiple offsets and the position of the first pixel point in the first feature map, obtain multiple second pixel points corresponding to the first pixel point from the second feature map.
[0080] In some embodiments, the reference pixel point with the same position as the first pixel point in the second feature map can be determined based on the position of the first pixel point in the first feature map, and then the reference pixel point is offset based on each offset to obtain the second pixel point corresponding to each offset.
[0081] Step 206: Input the third eigenvalue into the weight prediction network to obtain the attention weight corresponding to each offset.
[0082] Among them, the attention weight prediction network can be used to predict the importance of the second pixel point corresponding to each offset.
[0083] In some embodiments, the attention weight prediction network may be a simple fully connected neural network (FCNN), or a more complex network structure, such as a convolutional neural network (CNN), a self-attention network, or a recurrent neural network (RNN), etc.
[0084] Step 207: Based on the attention weights, fuse the second feature values corresponding to multiple second pixel points to obtain the fourth feature value corresponding to the first pixel point.
[0085] Specifically, based on the attention weights corresponding to each offset, perform weighted summation on the second feature values of the second pixel points corresponding to each offset to obtain the fourth feature value.
[0086] In the embodiments of the present disclosure, the offset and attention weights corresponding to the first pixel point can be predicted. Then, based on the offset, determine multiple second pixel points corresponding to the first pixel point, and fuse the second feature values corresponding to the multiple second pixel points based on the attention weights, so as to avoid different information around the vehicle corresponding to the same pixel point in the first feature map and the second feature map, and further improve the accuracy of the obtained fourth feature value.
[0087] Step 208: Fuse the first feature value corresponding to each first pixel point in the first feature map with the corresponding fourth feature value to obtain a fused feature map.
[0088] In some possible implementation manners, based on the importance degrees of the local navigation map and the bird's-eye view map for the target local map, determine the first weight corresponding to the bird's-eye view map and the second weight corresponding to the local navigation map. Then, based on the first weight and the second weight, fuse the first feature value corresponding to each first pixel point with the fourth feature value to obtain the feature value corresponding to each pixel point in the fused feature map.
[0089] In some possible implementation manners, the first feature value corresponding to each first pixel point in the first feature map can also be added to the fourth feature value to obtain a fused feature map.
[0090] Figure 3 This is a schematic flowchart of a process for obtaining a fused feature map provided by an embodiment of the present disclosure. As Figure 3 shown, input the local navigation map M into the feature extraction network W v to obtain a second feature map Input the first feature value corresponding to the first pixel point in the first feature map B into the linear network W v so that the linear network W v performs a linear transformation on to obtain a second feature value q i and input the second feature value q i into the offset prediction network W Δ and the weight prediction network W a respectively to obtain the offset Δ Δ output by the offset prediction network W ri and the weight prediction network W aThe output attention weight a i , then from the second feature map Determine the reference pixel point r corresponding to the first pixel point i , and combined with the offset Δ ri , determine the second pixel point, and extract the second eigenvalue corresponding to the second pixel point from the second feature map Then based on the attention weight a i The second eigenvalue Fusion is performed to obtain the fourth eigenvalue, and then the fourth eigenvalue is combined with the first eigenvalue Add together to get the fusion feature map The target feature value corresponding to the second pixel point with the same position as the first pixel point Thus, the target pixel value corresponding to each pixel point in the fused feature map is determined in turn to obtain the fused feature map.
[0091] Step 209: Perform road element detection on the fused feature map to obtain a road element detection result.
[0092] Step 210: Generate a target local map based on the road element detection results.
[0093] The specific implementation forms of step 209 and step 210 may refer to the detailed descriptions in other embodiments of the present disclosure, and will not be described in detail here.
[0094] In the disclosed embodiment, a local navigation map and a bird's-eye view around the vehicle are first obtained, and feature extraction is performed on the bird's-eye view and the local navigation map respectively to obtain a first feature map corresponding to the bird's-eye view and a second feature map corresponding to the local navigation map, and then based on the offset prediction network and the weight prediction network, the offset corresponding to each first pixel in the first feature map and the attention weight corresponding to each offset are predicted, and then based on the attention weight corresponding to each offset, the second eigenvalue of the second pixel corresponding to each offset is fused to obtain the fourth eigenvalue corresponding to the first pixel, and the first eigenvalue corresponding to each first pixel in the first feature map is fused with the corresponding fourth eigenvalue to obtain a fused feature map, and finally the fused feature map is subjected to road element detection to obtain a road element detection result, and a target local map is generated based on the road element detection result. Thus, the attention mechanism can be used to accurately determine the multiple second pixels corresponding to the first pixel, and the global information provided by the local navigation map can be better fused with the information collected by the sensor on the vehicle, thereby further improving the accuracy of the generated target local map, further improving the vehicle's understanding and judgment of road conditions, and improving the safety of vehicle driving.
[0095] In order to implement the above embodiments, the present disclosure also proposes a map generating device.
[0096] Figure 4 The structural schematic diagram of the map generation device provided by the embodiments of the present disclosure.
[0097] As Figure 4 shown, the map generation device 400 may include:
[0098] A first acquisition module 401, configured to acquire a local navigation map and an aerial view around the vehicle, where the aerial view is generated based on the vehicle perimeter information collected by sensors on the vehicle;
[0099] A second acquisition module 402, configured to perform feature extraction on the aerial view and the local navigation map respectively to obtain a first feature map corresponding to the aerial view and a second feature map corresponding to the local navigation map;
[0100] A determination module 403, configured to determine a plurality of second pixel points corresponding to each first pixel point in the first feature map from the second feature map;
[0101] A fusion module 404, configured to fuse the first feature value corresponding to each first pixel point in the first feature map with the second feature values of the corresponding plurality of second pixel points in the second feature map to obtain a fused feature map;
[0102] A detection module 405, configured to perform road element detection on the fused feature map to obtain a road element detection result;
[0103] A generation module 406, configured to generate a target local map based on the road element detection result.
[0104] In some possible implementation manners, the determination module 403 is configured to:
[0105] Perform a linear transformation on the first feature value corresponding to the first pixel point to obtain a third feature value;
[0106] Input the third feature value into an offset prediction network to obtain a plurality of offsets;
[0107] Based on the plurality of offsets and the position of the first pixel point in the first feature map, obtain a plurality of second pixel points corresponding to the first pixel point from the second feature map.
[0108] In some possible implementation manners, the determination module 403 is configured to:
[0109] Based on the position of the first pixel point in the first feature map, determine a reference pixel point in the second feature map with the same position as the first pixel point;
[0110] Based on each offset, offset the reference pixel point to obtain a second pixel point corresponding to each offset.
[0111] In some possible implementations, the fusion module 404 is configured to:
[0112] Input the third eigenvalue into the weight prediction network to obtain the attention weight corresponding to each offset;
[0113] Based on the attention weights, fuse the second eigenvalues corresponding to multiple second pixel points to obtain the fourth eigenvalue corresponding to the first pixel point;
[0114] Fuse the first eigenvalue corresponding to each first pixel point in the first feature map with the corresponding fourth eigenvalue to obtain a fused feature map.
[0115] In some possible implementations, the fusion module 404 is configured to:
[0116] Add the first eigenvalue corresponding to each first pixel point in the first feature map to the fourth eigenvalue to obtain a fused feature map.
[0117] In some possible implementations, it further includes a processing module configured to:
[0118] Based on the scale of the first feature map, perform a scaling process on the second feature map so that the scaled second feature map has the same scale as the first feature map.
[0119] For the functions and specific implementation principles of the above-mentioned modules in the embodiments of the present disclosure, reference may be made to the above-mentioned method embodiments, and details are not described herein again.
[0120] The map generation device according to the embodiments of the present disclosure first obtains a local navigation map and a bird's-eye view around the vehicle, and respectively performs feature extraction on the bird's-eye view and the local navigation map to obtain a first feature map corresponding to the bird's-eye view and a second feature map corresponding to the local navigation map. Then, from the second feature map, multiple second pixel points corresponding to each first pixel point in the first feature map are determined. The first eigenvalue corresponding to each first pixel point in the first feature map is fused with the second eigenvalue of the corresponding multiple second pixel points in the second feature map to obtain a fused feature map. Finally, road element detection is performed on the fused feature map to obtain a road element detection result, and based on the road element detection result, a target local map is generated. Thus, the information provided by the local navigation map can be fused with the information collected by the vehicle sensors. Through the global information provided by the local navigation map, the problem of insufficient perception ability of the sensors on the vehicle can be effectively compensated, so that a more accurate target local map can be generated, improving the vehicle's understanding and judgment ability of the road conditions and enhancing the safety of vehicle driving.
[0121] Figure 5FIG. is a schematic functional block diagram of a vehicle shown in an exemplary embodiment. For example, vehicle 500 can be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. Vehicle 500 can be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0122] Referring Figure 5 , vehicle 500 can include various subsystems. For example, the infotainment system 5410, the perception system 520, the decision control system 530, the drive system 540, and the computing platform 550. Among them, vehicle 500 can also include more or fewer subsystems, and each subsystem can include multiple components. In addition, each subsystem and each component of vehicle 500 can be interconnected by wired or wireless means.
[0123] In some embodiments, the infotainment system 510 can include a communication system, an entertainment system, and a navigation system, etc. The perception system 520 can include several sensors for sensing information about the environment around vehicle 500. For example, the perception system 520 can include a global positioning system (the global positioning system can be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter wave radar, ultrasonic radar, and a camera device.
[0124] The decision control system 530 can include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system. The drive system 540 can include components that provide power movement for vehicle 500. In one embodiment, the drive system 540 can include an engine, an energy source, a transmission system, and wheels. The engine can be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine can convert the energy provided by the energy source into mechanical energy.
[0125] Some or all of the functions of vehicle 500 are controlled by the computing platform 550. The computing platform 550 can include at least one processor 551 and a memory 552. The processor 551 can execute instructions 553 stored in the memory 552.
[0126] The processor 551 can be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.
[0127] The memory 552 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0128] In addition to the instructions 553, the memory 552 can also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 552 can be used by the computing platform 550. In the embodiments of the present disclosure, the processor 551 can execute the instructions 553 to complete all or part of the steps of the above-described map generation method.
[0129] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the map generation method provided by the present disclosure are implemented.
[0130] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0131] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0132] Any process or method description represented in a flowchart or described otherwise herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be executed not in the order shown or discussed, including substantially concurrently or in reverse order according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure pertain.
[0133] The logic and / or steps represented in a flowchart or described otherwise herein, for example, may be considered as a sequenced list of executable instructions for implementing a logical function and may be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0134] It should be understood that various parts of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0135] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0136] In addition, in each of the embodiments of the present disclosure, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0137] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A method for generating a map, characterized in that, Including: Obtain a local navigation map and a bird's-eye view around the vehicle, where the bird's-eye view is generated based on the vehicle perimeter information collected by sensors on the vehicle; Extract features from the bird's-eye view and the local navigation map respectively to obtain a first feature map corresponding to the bird's-eye view and a second feature map corresponding to the local navigation map; Determine a plurality of second pixel points corresponding to each first pixel point in the first feature map from the second feature map; Fuse the first feature value corresponding to each first pixel point in the first feature map with the second feature values of the corresponding plurality of second pixel points in the second feature map to obtain a fused feature map; Perform road element detection on the fused feature map to obtain a road element detection result; Generate a target local map based on the road element detection result.
2. The method according to claim 1, wherein The determining a plurality of second pixel points corresponding to each first pixel point in the first feature map from the second feature map includes: Perform a linear transformation on the first feature value corresponding to the first pixel point to obtain a third feature value; Input the third feature value into an offset prediction network to obtain a plurality of offsets; Based on the plurality of offsets and the position of the first pixel point in the first feature map, obtain a plurality of second pixel points corresponding to the first pixel point from the second feature map.
3. The method according to claim 2, characterized in that, The obtaining a plurality of second pixel points corresponding to the first pixel point from the second feature map based on the plurality of offsets and the position of the first pixel point in the first feature map includes: Based on the position of the first pixel point in the first feature map, determine a reference pixel point in the second feature map with the same position as the first pixel point; Offset the reference pixel point based on each offset to obtain a second pixel point corresponding to each offset.
4. The method according to claim 2, wherein The fusing the first feature value corresponding to each first pixel point in the first feature map with the second feature values of the corresponding plurality of second pixel points in the second feature map to obtain a fused feature map includes: Input the third feature value into a weight prediction network to obtain an attention weight corresponding to each offset; Based on the attention weight, fuse the second feature values corresponding to the plurality of second pixel points to obtain a fourth feature value corresponding to the first pixel point; Fuse the first feature value corresponding to each first pixel point in the first feature map with the corresponding fourth feature value to obtain the fused feature map.
5. The method according to claim 4, wherein The fusing the first feature value corresponding to each first pixel point in the first feature map with the corresponding fourth feature value to obtain the fused feature map includes: Add the first feature value corresponding to each first pixel point in the first feature map to the fourth feature value to obtain the fused feature map.
6. The method according to claim 1, characterized in that Before the determining a plurality of second pixel points corresponding to each first pixel point in the first feature map from the second feature map, it further includes: Scale the second feature map based on the scale of the first feature map so that the scaled second feature map has the same scale as the first feature map.
7. A map generation device, characterized in that, The device includes: A first acquisition module, configured to acquire a local navigation map and a bird's-eye view map around the vehicle, where the bird's-eye view map is generated based on the vehicle perimeter information collected by sensors on the vehicle; A second acquisition module, configured to perform feature extraction on the bird's-eye view map and the local navigation map respectively to obtain a first feature map corresponding to the bird's-eye view map and a second feature map corresponding to the local navigation map; A determination module, configured to determine, from the second feature map, a plurality of second pixel points corresponding to each first pixel point in the first feature map; A fusion module, configured to fuse the first feature value corresponding to each first pixel point in the first feature map with the second feature values of the corresponding plurality of second pixel points in the second feature map to obtain a fused feature map; A detection module, configured to perform road element detection on the fused feature map to obtain a road element detection result; A generation module, configured to generate a target local map based on the road element detection result.
8. The device according to claim 7, wherein, The determination module is configured to: Perform a linear transformation on the first feature value corresponding to the first pixel point to obtain a third feature value; Input the third feature value into an offset prediction network to obtain a plurality of offsets; Based on the plurality of offsets and the position of the first pixel point in the first feature map, obtain, from the second feature map, a plurality of second pixel points corresponding to the first pixel point.
9. A vehicle, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; wherein the processor is configured to implement the steps of the method according to any one of claims 1-6.
10. A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enabling the mobile terminal to execute a map generation method, the method including: Acquire a local navigation map and a bird's-eye view map around the vehicle, where the bird's-eye view map is generated based on the vehicle perimeter information collected by sensors on the vehicle; Perform feature extraction on the bird's-eye view map and the local navigation map respectively to obtain a first feature map corresponding to the bird's-eye view map and a second feature map corresponding to the local navigation map; Determine, from the second feature map, a plurality of second pixel points corresponding to each first pixel point in the first feature map; Fuse the first feature value corresponding to each first pixel point in the first feature map with the second feature values of the corresponding plurality of second pixel points in the second feature map to obtain a fused feature map; Perform road element detection on the fused feature map to obtain a road element detection result; Generate a target local map based on the road element detection result.