Intersection aerial view generation method and device, electronic equipment and readable storage medium

By determining the vehicle's peripheral image characteristics, generating the initial bird's-eye feature map and combining map elements and vehicle information, using feature interaction parameters to generate the bird's-eye feature map of the intersection, solving the problem of inaccurate generation of the bird's-eye view of the intersection caused by sensor perception range and occlusion, and achieving high accuracy perception in complex intersection environments.

CN120279084APending Publication Date: 2025-07-08CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410026245.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When generating a bird's-eye feature map of the intersection, the prior art is affected by the sensor's perception range limitations and obstacle occlusion, resulting in insufficient expressiveness of the model and unable to adapt to complex large intersections, and the perception information is not matched when stitching images.

Method used

By determining the image characteristics of the vehicle's peripheral vision image, an initial bird's-eye feature map is generated, map elements and feature interaction parameters are determined, and the vehicle position information and posture information is combined, residual network and feature pyramid network are used for feature extraction, and image projection is used for geometric guided nuclear transformer, and finally a bird's-eye feature map of the intersection is generated.

Benefits of technology

It improves the accuracy of the generation of bird's-eye feature maps at the intersection, solves the sensor's perception range and occlusion problems, and ensures the matching and accuracy of perceived information in complex intersection environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279084A_ABST
    Figure CN120279084A_ABST
Patent Text Reader

Abstract

The invention relates to an intersection aerial view generation method and apparatus, an electronic device and a readable storage medium. The method comprises the steps of determining image features for a panoramic image; generating an initial aerial view feature map for the vehicle based on the image features; determining map elements for the initial aerial view feature map; determining feature interaction parameters for the initial aerial view feature map based on the map elements; acquiring position information and attitude information of the vehicle; and generating an intersection aerial view feature map based on the initial aerial view feature map, the position information, the attitude information and the feature interaction parameters, thereby improving the accuracy of generating the aerial view.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of assisted driving lane environment perception, and particularly to a method for generating an intersection bird's-eye view, a device for generating an intersection bird's-eye view, an electronic device, and a computer-readable storage medium. Background Art

[0002] The assisted driving technology of automobiles is a technology that realizes assisted driving of automobiles through a computer system. The assisted driving of automobiles mainly includes four parts, namely, environment perception, positioning, planning, and control. Among them, environment perception is an important part of the assisted driving system and has been increasingly emphasized as a key step in assisted driving. At the same time, good perception performance is a prerequisite for planning and control. An important step in environment perception is to obtain a bird's-eye view feature map. Usually, multiple wide-angle cameras that can cover all the field-of-view ranges around the vehicle are installed around the vehicle, and the multi-channel video images collected at the same time are processed to synthesize a 360-degree bird's-eye view feature map around the vehicle. Currently, the common methods for generating a panoramic bird's-eye view of in-vehicle surround cameras mainly include generating a bird's-eye view of a single camera and stitching multiple bird's-eye view feature maps. However, the current methods for generating bird's-eye view feature maps are limited by the sensor perception range or blocked by obstacles, etc., resulting in insufficient model expressiveness and being unable to adapt to complex large intersections. At the same time, during the process of image stitching, problems of mismatched perception information will occur due to differences in the sensor perception range. Summary of the Invention

[0003] One of the purposes of the present invention is to provide a method for generating an intersection bird's-eye view to solve the problem of how to generate an intersection bird's-eye view feature map; the second purpose is to provide a device for generating an intersection bird's-eye view; the third purpose is to provide an electronic device; the fourth purpose is to provide a computer-readable storage medium.

[0004] To achieve the above purposes, the technical solutions adopted by the present invention are as follows:

[0005] The present invention provides a method for generating an intersection bird's-eye view. The method is applied to a vehicle, and vehicle-mounted cameras are arranged around the vehicle. The vehicle-mounted cameras are used to obtain the panoramic images of the vehicle, and may include:

[0006] Determine the image features of the panoramic images;

[0007] Generate an initial bird's-eye view feature map for the vehicle based on the image features;

[0008] Determine the map elements of the initial bird's-eye view feature map;

[0009] Determine the feature interaction parameters of the initial bird's-eye view feature map based on the map elements;

[0010] Obtain the position information and attitude information for the vehicle;

[0011] Generate an intersection bird's-eye view feature map based on the initial bird's-eye view feature map, the position information, the attitude information, and the feature interaction parameters.

[0012] Optionally, the method has a corresponding residual network and a feature pyramid network, and the step of determining the image features for the panoramic image may include:

[0013] Use the residual network and the feature pyramid network to determine the image features for the panoramic image.

[0014] Optionally, the method has a corresponding geometric guidance kernel transformer, and the step of generating the initial bird's-eye view feature map for the vehicle based on the image features may include:

[0015] Determine the internal and external parameters for the on-vehicle camera;

[0016] Use the geometric guidance kernel transformer to project the image points of the panoramic image into a preset bird's-eye view range based on the internal and external parameters and the image features, and generate the initial bird's-eye view feature map for the vehicle.

[0017] Optionally, the map elements include lane dividing lines, road boundary lines, and crosswalk areas, and the step of determining the feature interaction parameters for the initial bird's-eye view feature map based on the map elements may include:

[0018] Respectively determine the query vectors for the lane dividing lines, the road boundary lines, and the crosswalk areas;

[0019] Generate key vectors and value vectors based on the initial bird's-eye view feature map;

[0020] Use deformable attention to determine the feature interaction parameters for the initial bird's-eye view feature map based on the query vectors, the key vectors, and the value vectors.

[0021] Optionally, the step of generating the intersection bird's-eye view feature map based on the initial bird's-eye view feature map, the position information, the attitude information, and the feature interaction parameters may include:

[0022] Determine the range parameters for the vehicle;

[0023] Determine a plurality of initial target bird's-eye view maps from the initial bird's-eye view feature map based on the range parameters;

[0024] Perform position adjustment on the initial target bird's-eye view maps based on the position information and the attitude information to generate a plurality of target bird's-eye view maps;

[0025] Stitch multiple of the target bird's-eye views to generate a historical bird's-eye feature map;

[0026] Update the historical bird's-eye feature map based on the feature interaction parameters to generate an intersection bird's-eye feature map.

[0027] Optionally, the step of updating the historical bird's-eye feature map based on the feature interaction parameters to generate an intersection bird's-eye feature map may include:

[0028] Establish a bird's-eye view query for the historical bird's-eye feature map based on the feature interaction parameters;

[0029] Generate an intersection bird's-eye feature map based on the bird's-eye view query and the historical bird's-eye feature map.

[0030] Optionally, the method may be applied to an intersection bird's-eye feature map generation model, the model includes a high-precision map element generation module, and the high-precision map element generation module is configured with a detection head, and the detection head is used to output probability scores for the lane demarcation line, the road boundary line, and the crosswalk area.

[0031] An embodiment of the present invention further provides an intersection bird's-eye view generation device, the device is applied to a vehicle, and vehicle-mounted cameras are arranged around the vehicle, and the vehicle-mounted cameras are used to obtain panoramic images of the vehicle, and may include:

[0032] An image feature determination module, configured to determine image features for the panoramic image;

[0033] An initial bird's-eye feature map generation module, configured to generate an initial bird's-eye feature map for the vehicle based on the image features;

[0034] A map element determination module, configured to determine map elements for the initial bird's-eye feature map;

[0035] A feature interaction parameter determination module, configured to determine feature interaction parameters for the initial bird's-eye feature map based on the map elements;

[0036] An attitude information acquisition module, configured to acquire position information and attitude information for the vehicle;

[0037] An intersection bird's-eye feature map generation module, configured to generate an intersection bird's-eye feature map based on the initial bird's-eye feature map, the position information, the attitude information, and the feature interaction parameters.

[0038] Optionally, the device has a corresponding residual network and a feature pyramid network, and the image feature determination module may include:

[0039] An image feature determination sub-module, configured to determine image features for the panoramic image by using the residual network and the feature pyramid network.

[0040] Optionally, the device has a corresponding geometric guidance kernel transformer, and the initial bird's-eye feature map generation module may include:

[0041] An internal and external parameter determination sub-module, configured to determine the internal and external parameters for the vehicle-mounted camera;

[0042] An initial bird's-eye feature map generation sub-module, configured to project the image points of the panoramic image into a preset bird's-eye view range based on the internal and external parameters and the image features by using the geometric guidance kernel transformer, and generate an initial bird's-eye feature map for the vehicle.

[0043] Optionally, the map elements include lane dividing lines, road boundary lines, and crosswalk areas, and the feature interaction parameter determination module may include:

[0044] A query vector determination sub-module, configured to determine query vectors for the lane dividing lines, the road boundary lines, and the crosswalk areas respectively;

[0045] A value vector generation sub-module, configured to generate a key vector and a value vector based on the initial bird's-eye feature map;

[0046] A feature interaction parameter determination sub-module, configured to determine feature interaction parameters for the initial bird's-eye feature map by using deformable attention based on the query vectors, the key vector, and the value vector.

[0047] Optionally, the intersection bird's-eye feature map generation module may include:

[0048] A range parameter determination sub-module, configured to determine range parameters for the vehicle;

[0049] An initial target bird's-eye view determination sub-module, configured to determine a plurality of initial target bird's-eye views from the initial bird's-eye feature map based on the range parameters;

[0050] A target bird's-eye view generation sub-module, configured to perform position adjustment on the initial target bird's-eye views based on the position information and the attitude information, and generate a plurality of target bird's-eye views;

[0051] A historical bird's-eye feature map generation sub-module, configured to splice the plurality of target bird's-eye views to generate a historical bird's-eye feature map;

[0052] An intersection bird's-eye feature map generation sub-module, configured to update the historical bird's-eye feature map based on the feature interaction parameters to generate an intersection bird's-eye feature map.

[0053] Optionally, the intersection bird's-eye view feature map generation sub-module may include:

[0054] A bird's-eye view query establishment unit, configured to establish a bird's-eye view query for the historical bird's-eye view feature map based on the feature interaction parameter;

[0055] An intersection bird's-eye view feature map generation unit, configured to generate an intersection bird's-eye view feature map based on the bird's-eye view query and the historical bird's-eye view feature map.

[0056] Optionally, the device may be applied to an intersection bird's-eye view feature map generation model, and the model includes a high-precision map element generation module, and the high-precision map element generation module is configured with a detection head, and the detection head is used to output probability scores for the lane dividing line, the road boundary line, and the crosswalk area.

[0057] The present invention also provides an electronic device, including:

[0058] At least one processor; and,

[0059] A memory communicatively connected to at least one processor; wherein,

[0060] The memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor so that at least one processor can implement the above-mentioned intersection bird's-eye view generation method.

[0061] The present invention also provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, it can implement the above-mentioned intersection bird's-eye view generation method.

[0062] Advantages of the present invention:

[0063] In the embodiment of the present invention, by determining the image features of the panoramic image; generating an initial bird's-eye view feature map for the vehicle based on the image features; determining map elements for the initial bird's-eye view feature map; determining feature interaction parameters for the initial bird's-eye view feature map based on the map elements; obtaining position information and attitude information for the vehicle; and generating an intersection bird's-eye view feature map based on the initial bird's-eye view feature map, the position information, the attitude information, and the feature interaction parameters, the accuracy of generating the bird's-eye view is improved. Description of the Drawings

[0064] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings, and these exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the drawings do not constitute a proportional limitation.

[0065] Figure 1It is a flowchart of the steps of a method for generating an aerial view of an intersection provided in an embodiment of the present invention;

[0066] Figure 2 It is a flowchart of the steps of another method for generating an aerial view of an intersection provided in an embodiment of the present invention;

[0067] Figure 3 It is a schematic flowchart of the first module provided in an embodiment of the present invention;

[0068] Figure 4 It is a schematic flowchart of the second module provided in an embodiment of the present invention;

[0069] Figure 5 It is a schematic flowchart of the third module provided in an embodiment of the present invention;

[0070] Figure 6 It is a schematic flowchart of the fourth module provided in an embodiment of the present invention;

[0071] Figure 7 It is a structural block diagram of a device for generating an aerial view of an intersection provided in an embodiment of the present invention;

[0072] Figure 8 It is a hardware structural block diagram of an electronic device provided in various embodiments of the present invention. Detailed implementation manners

[0073] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0074] Environmental perception provides rich and valuable information for assisted driving and is one of the key steps. The deep learning method based on the bird's-eye view feature map brings a new paradigm of multi-sensor feature fusion for perception, greatly improving the perception ability. Due to this, the dynamic target detection and static element generation in the three-dimensional driving scene have been further developed. The perception method based on the bird's-eye view feature map usually encodes the data collected by cameras, radars, etc. from each perspective respectively, and then converts these features to the range of the bird's-eye view to form the bird's-eye view features. Finally, the target task results are output through the decoder. For example, obtain the point cloud detection data set of the lidar; input the point cloud detection data set into the pre-constructed sparse convolution network module to convert the sparse three-dimensional features in the point cloud into a dense feature map, and splice all the dense feature maps to obtain the bird's-eye view feature map; then input the bird's-eye view feature map into the confidence correction module to output the corrected three-dimensional detection box, and use the corrected three-dimensional detection box to detect the target, or, form a data set with the camera image and the corresponding lidar point cloud; use the data set to train the online high-precision map construction framework; input the collected lidar point cloud to be reconstructed into the trained online high-precision map construction framework and then output. The above methods all have the problems of not considering the sensor perception range and occlusion, and for large or complex intersections in the urban scene, single data collection is difficult to cover, which may lead to the lack of key information. At the same time, due to the possible differences in the ranges of the lidar point cloud and the camera bird's-eye view feature map, directly splicing the two feature maps may lead to mismatching of the perception information, resulting in insufficient accuracy of the output result. The embodiment of the present invention provides a method for generating a bird's-eye view of an intersection, which generates a bird's-eye view feature map of the intersection based on the pose information of the vehicle and the feature interaction parameters to improve the accuracy of generating the bird's-eye view feature map of the intersection.

[0075] Referring to Figure 1 , a flowchart of the steps of a method for generating a bird's-eye view of an intersection provided in an embodiment of the present invention is shown, which may specifically include the following steps:

[0076] Step 101, determine the image features of the panoramic image;

[0077] Step 102, generate an initial bird's-eye view feature map for the vehicle based on the image features;

[0078] Step 103, determine the map elements of the initial bird's-eye view feature map;

[0079] Step 104, determine the feature interaction parameters of the initial bird's-eye view feature map based on the map elements;

[0080] Step 105, obtain the position information and attitude information of the vehicle;

[0081] Step 106: Generate a bird's-eye view feature map of the intersection based on the initial bird's-eye view feature map, the position information, the attitude information, and the feature interaction parameters.

[0082] In practical applications, the embodiments of the present invention can be applied to vehicles. Vehicle-mounted cameras can be configured around the vehicle, which can be used to obtain panoramic images of the vehicle. At the same time, since the feature maps generated under similar pose conditions of the vehicle also have similar values, multiple acquisitions can be performed at a single intersection. The panoramic images can be driving scene images captured at fixed time intervals.

[0083] In a specific implementation, the embodiments of the present invention can determine the image features of the panoramic image; generate an initial bird's-eye view feature map for the vehicle based on the image features; determine the map elements for the initial bird's-eye view feature map; determine the feature interaction parameters for the initial bird's-eye view feature map based on the map elements; obtain the position information and attitude information of the vehicle; and generate a bird's-eye view feature map of the intersection based on the initial bird's-eye view feature map, the position information, the attitude information, and the feature interaction parameters. Exemplarily, the panoramic image can be a panoramic camera image captured by a vehicle-mounted camera configured on the vehicle. The image features of the panoramic camera image can be determined, and an initial bird's-eye view feature map for the vehicle can be generated based on the image features. The initial bird's-eye view feature map can include map elements. The map elements can be determined from the initial bird's-eye view feature map, and then the feature interaction parameters for the initial bird's-eye view feature map can be determined based on the map elements. Next, the position information and attitude information of the vehicle can be obtained. For example, the current position of the vehicle can be used as the position information of the vehicle, and the driving direction of the vehicle can be used as the attitude information of the vehicle. Then, a bird's-eye view feature map of the intersection can be generated based on the initial bird's-eye view map, the position information, the attitude information, and the feature interaction parameters.

[0084] In the embodiments of the present invention, by determining the image features of the panoramic image; generating an initial bird's-eye view feature map for the vehicle based on the image features; determining the map elements for the initial bird's-eye view feature map; determining the feature interaction parameters for the initial bird's-eye view feature map based on the map elements; obtaining the position information and attitude information of the vehicle; and generating a bird's-eye view feature map of the intersection based on the initial bird's-eye view feature map, the position information, the attitude information, and the feature interaction parameters, the accuracy of generating the bird's-eye view map is improved.

[0085] Based on the above embodiments, variant embodiments of the above embodiments are proposed. Here, it should be noted that for the sake of brevity of description, only the differences from the above embodiments are described in the variant embodiments.

[0086] In an optional embodiment of the present invention, the step of determining the image features of the panoramic image includes:

[0087] Determine the image features for the panoramic image by using the residual network and the feature pyramid network.

[0088] In practical applications, the embodiments of the present invention may have corresponding residual networks and feature pyramid networks. A convolutional neural network can be composed of a residual network and a feature pyramid network, and the convolutional neural network is used to extract the image features of the panoramic image. Exemplarily, the residual network can be ResNet50, and the feature pyramid network can be FPN. Among them, ResNet50 is a deep learning model. The full name of "ResNet" is "Residual Network", which means "residual network", and "50" indicates that this network contains 50 layers. The main feature of ResNet50 is the introduction of the "residual block" (ResidualBlock). In traditional neural networks, each layer adds a new transformation on the basis of the previous layer, while in ResNet, each layer adds a new transformation on the basis of the previous layer and also retains the original input of the previous layer. This is the so-called "residual". This design enables the network to better learn the difference between the input and the output, rather than directly learning the output, which helps to improve the performance of the model. The full name of FPN is feature pyramid networks, which is a multi-scale object detection algorithm applicable to. FPN fuses features of different scales to achieve accurate detection of multi-scale objects. FPN starts from the high-resolution feature map through top-down feature propagation and uses the upsampling operation to gradually propagate the features to lower scales. Then, through bottom-up feature fusion, FPN fuses the high-resolution features with the low-resolution features to obtain a multi-scale feature pyramid with rich semantic information. Finally, FPN uses these fused features for the object detection task to achieve accurate detection of objects of different scales. FPN can effectively fuse features of different scales to obtain a multi-scale feature pyramid in the object detection task. In this way, the network can not only detect small objects but also large objects, improving the accuracy of object detection. And through feature propagation and feature fusion, FPN combines high-level semantic information with low-level detail information, enabling the network to obtain rich semantic information and detail information simultaneously when performing object detection, thus improving the robustness and generalization ability of object detection. Compared with other multi-scale feature fusion methods, FPN has a lightweight network structure and does not bring excessive computational and storage overheads. This enables FPN to achieve efficient feature fusion and object localization in the object detection task.

[0089] In a specific implementation, embodiments of the present invention may use a residual network and a feature pyramid network to determine image features for a panoramic image. Exemplarily, when the residual network is ResNet50 and the feature pyramid network is FPN, the residual network ResNet50 may be used to extract features of the panoramic image, and at the same time, the feature pyramid network FPN may be used to extract multi-scale information for the panoramic image, thereby improving the model's detection ability for targets of different scales.

[0090] Embodiments of the present invention determine image features for the panoramic image by using the residual network and the feature pyramid network, thereby improving the accuracy and adaptability of image feature extraction for the panoramic image and providing support for subsequent data processing.

[0091] In an optional embodiment of the present invention, the step of generating an initial bird's-eye view feature map for the vehicle based on the image features includes:

[0092] Determine the internal and external parameters for the vehicle-mounted camera;

[0093] Use the geometric guidance kernel transformer to project the image points of the panoramic image into a preset bird's-eye view range based on the internal and external parameters and the image features, and generate an initial bird's-eye view feature map for the vehicle.

[0094] In practical applications, embodiments of the present invention may have a corresponding geometric guidance kernel transformer. Exemplarily, the geometric guidance kernel transformer may be Geometry-guided Kernel Transformer, abbreviated as GKT. GKT uses geometric priors to guide the transformer to focus on discriminative regions, thereby generating a BEV (Bird's-eye-view) representation with panoramic image features, also known as a bird's-eye view representation. GKT is based on kernel-wise attention and is efficient. Especially in the case of LUT indexing (Lookup Table, a commonly used image processing technology that maps pixel values in an image through a pre-defined lookup table to achieve pixel-level transformation operations), GKT is robust to camera deviations, making the 2D to BEV transformation more stable and reliable.

[0095] In a specific implementation, embodiments of the present invention can determine the internal and external parameters of an in-vehicle camera; and use a geometric-guided kernel transformer to project the image points of the panoramic image into a preset bird's-eye view range based on the internal and external parameters and image features, generating an initial bird's-eye view feature map for the vehicle. Exemplarily, the internal and external parameters of the in-vehicle camera can be the internal and external camera parameters. The internal and external camera parameters are important parameters in camera measurement, representing the internal structure and external pose of the camera. The internal parameters refer to the internal imaging structure of the camera, including focal length, pixel coordinate system, distortion, etc.; the external parameters refer to the description of the pose and position of the camera in the external space, including the rotation and translation of the camera. The determination of the internal and external camera parameters is of great significance for fields such as vision research, robot navigation, and 3D reconstruction. The internal and external parameters of the in-vehicle camera can be determined through manual calibration, automatic calibration based on planar rigid body motion, multi-camera calibration, camera self-calibration, etc. After determining the internal and external parameters of the in-vehicle camera, a geometric-guided kernel transformer can be used to project the image points of the panoramic image into a preset bird's-eye view range. For example, by focusing on a local area through a geometric prior-guided transformer, an initial bird's-eye view feature map for the vehicle can be generated.

[0096] In the embodiments of the present invention, by determining the internal and external parameters of the in-vehicle camera; using the geometric-guided kernel transformer to project the image points of the panoramic image into a preset bird's-eye view range based on the internal and external parameters and the image features, generating an initial bird's-eye view feature map for the vehicle, the dependence on geometric priors is reduced, and at the same time, the computational consumption caused by global interaction is avoided.

[0097] In an optional embodiment of the present invention, the step of determining the feature interaction parameters for the initial bird's-eye view feature map based on the map elements includes:

[0098] Determine query vectors for the lane demarcation line, the road boundary line, and the crosswalk area respectively;

[0099] Generate a key vector and a value vector based on the initial bird's-eye view feature map;

[0100] Use deformable attention to determine the feature interaction parameters for the initial bird's-eye view feature map based on the query vector, the key vector, and the value vector.

[0101] In practical applications, the map elements in the embodiments of the present invention may include lane demarcation lines, road boundary lines, and crosswalk areas. Among them, the lane demarcation lines may be white dotted lines, which are used to divide the traffic flow moving in the same direction and are set on the lane demarcation lines for vehicles moving in the same direction. Under the condition of ensuring safety, vehicles are allowed to cross the line to change lanes. The road boundary lines may be composed of white solid lines and dotted lines, which are generally used to mark the boundaries of lanes, indicating that vehicles are not allowed to cross this line. The crosswalk area may be the walking range specified for pedestrians to cross the lane marked with zebra stripes or other markings on the roadway, which is an area marked on the roadway to prevent vehicles from hurting pedestrians when driving fast and designates the area where vehicles need to slow down to give way to pedestrians crossing the street. Different map elements have different functions, so they also have different requirements for vehicle driving. Accurately dividing the lane demarcation lines, road boundary lines, and crosswalk areas can effectively improve the safety and reliability of assisted driving.

[0102] In a specific implementation, the embodiments of the present invention can respectively determine query vectors for the lane demarcation lines, road boundary lines, and crosswalk areas; generate key vectors and value vectors based on the initial bird's-eye view feature map; and use deformable attention to determine the feature interaction parameters for the initial bird's-eye view feature map based on the query vectors, key vectors, and value vectors. Exemplarily, the query vector may be a Query Vector. The Query Vector refers to the representation of the input sequence that the model currently wants to focus on or process, usually the last hidden state obtained by the encoder or the current decoder state. The query vector can represent the information that the model is currently focusing on, can be regarded as a question or content that needs to be judged, and can be generated by the encoder. The key vector can be used to calculate the similarity between the query vector and each element in all the remaining input sequences, and can be the information provided to the model for retrieval or search. The value vector can correspond to each key vector and contain the importance weight of the input sequence element represented by this key in the model, can be generated by the encoder, and can give an answer or decision for specific content. When the map elements include lane demarcation lines, road boundary lines, and crosswalk areas, query vectors for the lane demarcation lines, road boundary lines, and crosswalk areas can be determined, key vectors and value vectors can be generated based on the initial bird's-eye view feature map, and feature interaction between the query vector and the key vector and value vector can be realized through deformable attention, and the feature interaction parameters can be determined.

[0103] In the embodiments of the present invention, by respectively determining the query vectors for the lane demarcation line, the road boundary line, and the crosswalk area; generating the key vector and the value vector based on the initial bird's-eye view feature map; and using deformable attention to determine the feature interaction parameters for the initial bird's-eye view feature map based on the query vector, the key vector, and the value vector, the decoding of the bird's-eye view feature map is realized, and high-precision map elements are generated, which provides convenience for subsequent use.

[0104] In an alternative embodiment of the present invention, the step of generating an intersection bird's-eye view feature map based on the initial bird's-eye view feature map, the position information, the attitude information, and the feature interaction parameter includes:

[0105] Determine a range parameter for the vehicle;

[0106] Based on the range parameter, determine a plurality of initial target bird's-eye view maps from the initial bird's-eye view feature map;

[0107] Based on the position information and the attitude information, perform position adjustment on the initial target bird's-eye view map to generate a plurality of target bird's-eye view maps;

[0108] Stitch together a plurality of the target bird's-eye view maps to generate a historical bird's-eye view feature map;

[0109] Update the historical bird's-eye view feature map based on the feature interaction parameter to generate an intersection bird's-eye view feature map.

[0110] In a specific implementation, the embodiment of the present invention can determine a range parameter for the vehicle; determine a plurality of initial target bird's-eye view maps from the initial bird's-eye view feature map based on the range parameter; perform position adjustment on the initial target bird's-eye view map based on the position information and the attitude information to generate a plurality of target bird's-eye view maps; stitch together the plurality of target bird's-eye view maps to generate a historical bird's-eye view feature map; update the historical bird's-eye view feature map based on the feature interaction parameter to generate an intersection bird's-eye view feature map. Exemplarily, the range parameter of the vehicle can be set to 120 meters × 120 meters. Based on the range parameter of 120 meters × 120 meters and with the vehicle itself as the center, select all the initial bird's-eye view feature maps whose acquisition positions are within the 120 meters × 120 meters range parameter from the plurality of initial bird's-eye view feature maps as the initial target bird's-eye view maps, and based on the position information and the attitude information, perform position adjustment on the initial target bird's-eye view map, such as translating and / or rotating the initial target bird's-eye view map according to the vehicle's own position information and attitude information, and use the position-adjusted initial target bird's-eye view map as the target bird's-eye view map, and stitch together the target bird's-eye view maps to generate a historical bird's-eye view feature map.

[0111] Preferably, before stitching together the target bird's-eye view maps, considering the low reliability of feature learning at the edge of the feature map, the target bird's-eye view maps can be cropped. For example, set the cropping retention range to 80% of the target bird's-eye view map, and then perform stitching after cropping to improve the reliability of feature learning.

[0112] Then, the historical bird's-eye view feature map can be updated based on the feature interaction parameter to generate an intersection bird's-eye view feature map.

[0113] Preferably, pixel replication can be performed on the historical bird's-eye view feature map. For pixels where multiple feature maps overlap, considering the similarity of the feature values of the same local scene, the average value of the overlapping pixels can be used for representation.

[0114] In an embodiment of the present invention, by determining range parameters for the vehicle; determining a plurality of initial target bird's-eye views from the initial bird's-eye view feature map based on the range parameters; performing position adjustment on the initial target bird's-eye views based on the position information and the attitude information to generate a plurality of target bird's-eye views; splicing the plurality of target bird's-eye views to generate a historical bird's-eye view feature map; and updating the historical bird's-eye view feature map based on the feature interaction parameters to generate an intersection bird's-eye view feature map, it is thus achieved that the acquisition of the intersection bird's-eye view feature map can support multiple data inputs, improving the accuracy of generating the bird's-eye view.

[0115] In an optional embodiment of the present invention, the step of updating the historical bird's-eye view feature map based on the feature interaction parameters to generate an intersection bird's-eye view feature map includes:

[0116] Establishing a bird's-eye view query for the historical bird's-eye view feature map based on the feature interaction parameters;

[0117] Generating an intersection bird's-eye view feature map based on the bird's-eye view query and the historical bird's-eye view feature map.

[0118] In a specific implementation, an embodiment of the present invention can establish a bird's-eye view query for the historical bird's-eye view feature map based on the feature interaction parameters; generate an intersection bird's-eye view feature map based on the bird's-eye view query and the historical bird's-eye view feature map. Exemplarily, a bird's-eye view query can be established for the current historical bird's-eye view feature map. For each unit feature in the query, find the corresponding position in the historical bird's-eye view feature map, extract the features of several adjacent units, perform cross-attention calculation on the two, obtain the updated historical bird's-eye view feature map, and then the updated historical bird's-eye view feature map can be placed in the map according to the position information and the attitude information to generate an intersection bird's-eye view feature map.

[0119] In an embodiment of the present invention, by establishing a bird's-eye view query for the historical bird's-eye view feature map based on the feature interaction parameters; generating an intersection bird's-eye view feature map based on the bird's-eye view query and the historical bird's-eye view feature map, it is thus achieved that further constructing an intersection bird's-eye view feature map based on map elements, improving the convenience of the bird's-eye view feature map in practical applications.

[0120] In an optional embodiment of the present invention, the model includes a high-precision map element generation module, and the high-precision map element generation module is configured with a detection head, and the detection head is used to output probability scores for the lane dividing line, the road boundary line, and the crosswalk area.

[0121] In a specific implementation, the embodiments of the present invention can be applied to an intersection aerial feature map generation model. The model includes a high-precision map element generation module, and the high-precision map element generation module is configured with a detection head. The detection head is used to output probability scores for lane dividing lines, road boundary lines, and crosswalk areas. Exemplarily, the detection head can be composed of a fully connected layer and a softmax layer. Each node of the fully connected layer (FC) is connected to all nodes of the previous layer, which is used to synthesize the features extracted previously and plays the role of a "classifier" in the entire convolutional neural network. Softmax is an activation function that can normalize a numerical vector into a probability distribution vector, and the sum of all probabilities is 1. Softmax can be used as the last layer of a neural network for the output of multi-classification problems. The detection head can output probability scores for lane dividing lines, road boundary lines, and crosswalk areas. For example, the probability scores that the map element belongs to a lane dividing line, or a road boundary line, or a crosswalk area.

[0122] Preferably, the detection head can be made to output information on the positions of the points corresponding to the map elements simultaneously.

[0123] In the embodiments of the present invention, by making the model include a high-precision map element generation module, the high-precision map element generation module is configured with a detection head, and the detection head is used to output probability scores for the lane dividing line, the road boundary line, and the crosswalk area, the accuracy of generating an intersection aerial feature map is further improved.

[0124] To enable those skilled in the art to better understand the embodiments of the present invention, the following uses a complete example to illustrate the embodiments of the present invention.

[0125] Refer to Figure 2 , Figure 2 which is the flowchart of steps of another intersection aerial view generation method provided in the embodiments of the present invention; refer to Figure 3 , Figure 3 which is the schematic flowchart of the first module provided in the embodiments of the present invention; refer to Figure 4 , Figure 4 which is the schematic flowchart of the second module provided in the embodiments of the present invention; refer to Figure 5 , Figure 5 which is the schematic flowchart of the third module provided in the embodiments of the present invention; refer to Figure 6 , Figure 6 which is the schematic flowchart of the fourth module provided in the embodiments of the present invention.

[0126] First module: Camera image feature extraction module

[0127] The main purpose of this module is to extract features from the panoramic camera images through a convolutional neural network. This convolutional neural network consists of a Residual Network (ResNet50) and a Feature Pyramid Network (FPN). The former is mainly used for image feature extraction, and the latter is mainly used to extract multi-scale information to improve the model's detection ability for targets of different scales.

[0128] The Second Module: Bird's-eye View Feature Transformation Module

[0129] The main purpose of this module is to transform the panoramic camera image features into the bird's-eye view range, which is achieved by using a Geometry-guided Kernel Transformer (GKT). The GKT roughly projects the image points into the bird's-eye view range according to the internal and external camera parameters, that is, it uses geometric priors to guide the Transformer to focus on local areas, reducing the dependence on geometric priors while avoiding the computational consumption caused by global interactions. Based on this, the panoramic image features in the same pose are transformed into the bird's-eye view space, forming the bird's-eye view feature map at the current moment.

[0130] The Third Module: High-precision Map Element Generation Module

[0131] The main purpose of this module is to decode and generate high-precision map elements, mainly including lane dividing lines, road boundary lines, and crosswalk areas. The lines and areas are expressed by several points and the topological relationships between the points. This module mainly uses a Transformer structure to generate query vectors for each map element, such as lane dividing lines, and then generates key vectors and value vectors through the bird's-eye view feature map, and uses deformable attention to achieve feature interaction between the query vector and the key vector and value vector. In addition, this module also includes a detection head for outputting the class scores of map elements and the positions of their constituent points, which is mainly composed of a fully connected layer and a softmax layer.

[0132] The Fourth Module: Bird's-eye View Feature Map Stitching Module

[0133] The main purpose of this module is to splice the bird's-eye view feature maps of the same intersection output by the second module to generate a bird's-eye view feature map of the intersection with rich semantic information. The range is tentatively set to 120 meters × 120 meters, and all bird's-eye view feature maps whose acquisition positions are within this range need to be spliced. Before splicing, considering the low reliability of feature learning at the edges of the feature map, it is cropped, and the tentative cropping retention range is about 80% of the original bird's-eye view feature map. For the method of splicing the feature maps, a feature map with the initial intersection range size is proposed. For the feature maps that fall within this range, the feature maps are translated and rotated according to the position and pose information of the ego vehicle, and then each pixel of the bird's-eye view feature map of the intersection is assigned a value at the corresponding position. For the pixels where multiple feature maps overlap, considering the similarity of the feature values of the same local scene, the average value of the overlapping pixels is also used to represent them. Optionally, another way to generate a learnable bird's-eye view feature map of the intersection is provided. A feature map of 120 meters × 120 meters is also established. For any bird's-eye view feature map, a feature map of the same size is sampled at the corresponding position in the feature map according to the pose, denoted as the historical bird's-eye view feature map. A bird's-eye view query is established for the current bird's-eye view feature map. For each unit feature in the query, the corresponding position is found in the historical bird's-eye view feature map, the features of several adjacent units are extracted, and cross-attention calculation is performed on the two to obtain the updated bird's-eye view feature map, and the updated bird's-eye view feature map is placed in the feature map according to the pose. At this time, after the model is trained, the learned feature map can support multiple acquisitions of data input, complete the task of splicing the feature maps, and finally, after the intersection position is located, the bird's-eye view feature map of the intersection can be extracted.

[0134] Data Definition and Format:

[0135] Camera Image: It refers to the image data obtained by using on-vehicle cameras. The acquisition angles of this image cover all around the vehicle, and driving scene images are taken at fixed time intervals as the acquisition vehicle moves. It is assumed that the cameras around the acquisition vehicle take pictures simultaneously. The dimension of the image data obtained at the current moment can be expressed as (N, H, W, C), where N represents the number of panoramic cameras, H and W represent the width and height of the image respectively, and C represents the number of channels of the picture. For an RGB three-channel image, C is 3.

[0136] Bird's-eye View Feature Map: The tensor data generated after the camera image is processed by the first module and the second module represents the bird's-eye view feature map. This feature map is centered on the ego vehicle and is constructed according to the vehicle's forward direction and the vertical direction. The range can be appropriately adjusted according to the actual visible range of the sensor. Each pixel in the bird's-eye view feature map reflects the features of the corresponding position around the ego vehicle and is an aggregation of the image features related to this position. The dimension of the bird's-eye view feature map generated from the image features of the panoramic images at the same moment can be expressed as (H b , W b , C b ), where Hb and W b respectively represent the forward direction of the vehicle and the width and height perpendicular to the forward direction. C b represents the dimension of the feature map. It can be assumed that each pixel of the bird's-eye view feature image represents an actual size of 0.3 meters, covering an actual range of 30 meters in front and behind the vehicle and 15 meters on the left and right. At this time, H b and W b are 200 and 100 respectively. In addition, C b is taken as 256.

[0137] Bird's-eye view feature map of the intersection: The bird's-eye view feature map of the intersection is composed of the feature map located at a certain intersection through cropping and splicing. It is a high-dimensional feature representation of the given intersection information, and its dimension can be expressed as (H m , W m , C m ), where H m and W m respectively represent the width and height of the bird's-eye view feature map of the intersection, and C m represents the dimension of the feature map. The intersection size can be taken as 120 meters × 120 meters. According to the same resolution of 0.3 meters / pixel as above, H m and W m are both taken as 400. In addition, C m is also taken as 256.

[0138] The specific process is as follows:

[0139] Perform preliminary feature extraction on the camera image to provide an information source for subsequent transfer to the bird's-eye view space. The input of this module is a group of panoramic images with the same or similar time, and the output is the features of each image.

[0140] Adopt the pre-trained ResNet50 and FPN parameters and the corresponding structures.

[0141] Each picture is first preprocessed, including resampling, random rotation augmentation, unifying the picture size, and making it meet the network input requirements.

[0142] The processed pictures are fed into ResNet50 and enter stages one to four in sequence, outputting four feature maps with gradually decreasing sizes and increasing dimensions. Each stage contains a module composed of several convolutional, ReLU, and BN layers.

[0143] In order to enhance the learning ability of the image feature extraction module for features of targets at different scales, it is necessary to represent the feature maps at different scales through the network structure design of FPN.

[0144] The FPN network uses the bilinear interpolation method multiple times to interpolate the feature map with the smallest size output by ResNet50. Each interpolation is performed on the result of the previous interpolation, and the feature dimension and feature map size of the interpolation result correspond to the output of ResNet50.

[0145] In addition, before interpolation, feature maps of corresponding sizes will be added element-wise as the input for subsequent processes.

[0146] Then, the image features of multiple cameras at the same moment can be transformed according to the spatial relationship to generate a bird's-eye view feature map centered on the ego vehicle. The input of this module is the image features output by the first module, and the panoramic images at the same moment are regarded as a group, and the output is the corresponding bird's-eye view feature map of this group.

[0147] The GKT method is used to generate the bird's-eye view feature map.

[0148] First, a bird's-eye view query is constructed, with a range equal to the bird's-eye view feature map, which is 200×100.

[0149] Determine the rough position of each unit of the bird's-eye view query in the image according to the internal and external parameters of the camera.

[0150] Centered on this position, capture the kernel features within a local range, and then perform cross-attention calculation with the bird's-eye view query to obtain the bird's-eye view feature map, with a dimension of 200×100×256.

[0151] 256 is determined according to the number of channels of the feature map after extraction and fusion by ResNet50 and FPN.

[0152] Next, high-precision map element prediction can be achieved based on the bird's-eye view feature map output by the second module, especially for lane dividing lines, road boundary lines, and crosswalk areas. At the same time, after obtaining relatively accurate output results, a bird's-eye view feature map with rich semantic information is extracted.

[0153] For each map element, it is explicitly encoded as a set of query collections, and the set size is equal to the maximum number of points that make up the map element, which can be 20.

[0154] Then, the multi-head attention mechanism of Transformer is used to realize the information interaction between queries.

[0155] Deformable attention is used to realize the interaction between queries and bird's-eye view features. Each item in the query collection predicts a coordinate within the bird's-eye view range, and then the query is updated with the bird's-eye features around this coordinate.

[0156] The detection head at the end of the module has two outputs. One is the category score of the map element, which can represent the probability scores of the generated element belonging to three categories such as lane dividers. The other is the positions of the points that make up the element. This detection head is mainly composed of a fully connected layer and a softmax layer.

[0157] Then, the bird's-eye view feature maps at the intersection can be stitched together to obtain a complete bird's-eye view feature map of the intersection with rich semantic information.

[0158] For a certain intersection, a range of 120 meters × 120 meters is established with the intersection center, and the bird's-eye view feature maps that fall into it are screened out.

[0159] Considering the unreliability of the information at the edge of the bird's-eye view feature map, the feature map is cropped at a ratio of 10% in both the horizontal and vertical directions.

[0160] Initialize a bird's-eye view feature map of the intersection with a size of 400×400 (the length and width of each unit correspond to an actual ground distance of 0.3 meters). Translate and rotate the processed bird's-eye view feature map that falls into this intersection according to the position of the ego vehicle at the intersection and the ego vehicle's attitude, and place it in the corresponding bird's-eye view feature map of the intersection, discarding the features that exceed the range.

[0161] For the pixels where several feature maps overlap, their average value can be used as the final feature value.

[0162] Optionally, another way to generate a learnable bird's-eye view feature map of the intersection can be provided.

[0163] Similarly, a feature map of 120 meters × 120 meters is established, and its range is still larger than the bird's-eye view feature map output by the second module.

[0164] Given a bird's-eye view feature map, sample a feature map of the same size at the corresponding position in the feature map according to the pose, and denote it as the historical bird's-eye view feature map.

[0165] Establish a bird's-eye view map query for the current bird's-eye view feature map. For each unit feature in the query, find the corresponding position in the historical bird's-eye view feature map, extract the features of several adjacent units, and perform cross-attention calculation on the two to obtain the updated bird's-eye view feature map. The number of adjacent units that can be taken is 24.

[0166] Similarly, place the updated bird's-eye view feature map in the feature map according to the pose.

[0167] Through the above method, the function of generating a bird's-eye view feature map of the intersection based on map elements is realized, encoding information with rich semantic information, and avoiding the influence brought by the occlusion and field of view limitation problems that are prone to occur in single acquisition.

[0168] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0169] Referring to Figure 7 , a structural block diagram of an intersection bird's-eye view generation device provided in an embodiment of the present invention is shown, which may specifically include the following modules:

[0170] An image feature determination module 701, configured to determine the image features of the panoramic image;

[0171] An initial bird's-eye view feature map generation module 702, configured to generate an initial bird's-eye view feature map for the vehicle based on the image features;

[0172] A map element determination module 703, configured to determine the map elements for the initial bird's-eye view feature map;

[0173] A feature interaction parameter determination module 704, configured to determine the feature interaction parameters for the initial bird's-eye view feature map based on the map elements;

[0174] A pose information acquisition module 705, configured to acquire the position information and pose information of the vehicle;

[0175] An intersection bird's-eye view feature map generation module 706, configured to generate an intersection bird's-eye view feature map based on the initial bird's-eye view feature map, the position information, the pose information, and the feature interaction parameters.

[0176] Optionally, the device has a corresponding residual network and a feature pyramid network. The image feature determination module may include:

[0177] An image feature determination sub-module, configured to determine the image features of the panoramic image by using the residual network and the feature pyramid network.

[0178] Optionally, the device has a corresponding geometric guidance kernel transformer. The initial bird's-eye view feature map generation module may include:

[0179] An internal and external parameter determination sub-module, configured to determine the internal and external parameters of the vehicle-mounted camera;

[0180] An initial bird's-eye view feature map generation sub-module, which is used to project the image points of the panoramic image within a preset bird's-eye view range based on the internal and external parameters and the image features by using the geometric guidance kernel transformer, and generate an initial bird's-eye view feature map for the vehicle.

[0181] Optionally, the map elements include lane demarcation lines, road boundary lines, and crosswalk areas, and the feature interaction parameter determination module may include:

[0182] A query vector determination sub-module, which is used to determine query vectors for the lane demarcation lines, the road boundary lines, and the crosswalk areas respectively;

[0183] A value vector generation sub-module, which is used to generate key vectors and value vectors based on the initial bird's-eye view feature map;

[0184] A feature interaction parameter determination sub-module, which is used to determine the feature interaction parameters for the initial bird's-eye view feature map based on the query vectors, the key vectors, and the value vectors by using deformable attention.

[0185] Optionally, the intersection bird's-eye view feature map generation module may include:

[0186] A range parameter determination sub-module, which is used to determine range parameters for the vehicle;

[0187] An initial target bird's-eye view determination sub-module, which is used to determine a plurality of initial target bird's-eye views from the initial bird's-eye view feature map based on the range parameters;

[0188] A target bird's-eye view generation sub-module, which is used to perform position adjustment on the initial target bird's-eye views based on the position information and the attitude information, and generate a plurality of target bird's-eye views;

[0189] A historical bird's-eye view feature map generation sub-module, which is used to splice a plurality of the target bird's-eye views to generate a historical bird's-eye view feature map;

[0190] An intersection bird's-eye view feature map generation sub-module, which is used to update the historical bird's-eye view feature map based on the feature interaction parameters to generate an intersection bird's-eye view feature map.

[0191] Optionally, the intersection bird's-eye view feature map generation sub-module may include:

[0192] A bird's-eye view query establishment unit, which is used to establish a bird's-eye view query for the historical bird's-eye view feature map based on the feature interaction parameters;

[0193] An intersection bird's-eye view feature map generation unit, which is used to generate an intersection bird's-eye view feature map based on the bird's-eye view query and the historical bird's-eye view feature map.

[0194] Optionally, the device can be applied to an intersection bird's-eye view feature map generation model, which includes a high-precision map element generation module configured with a detection head for outputting probability scores for the lane boundary lines, road boundary lines, and crosswalk areas.

[0195] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the related parts, refer to the partial description of the method embodiments.

[0196] In addition, an embodiment of the present invention also provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the above-mentioned intersection bird's-eye view generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0197] An embodiment of the present invention also provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by the processor, it implements each process of the above-mentioned intersection bird's-eye view generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0198] Figure 8 It is a schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.

[0199] The electronic device 800 includes, but is not limited to: a radio frequency unit 801, a network module 802, an audio output unit 803, an input unit 804, a sensor 805, a display unit 806, a user input unit 807, an interface unit 808, a memory 809, a processor 810, and a power supply 811, etc. Those skilled in the art can understand that Figure 8 the electronic device structure shown in does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, mobile phones, tablet computers, laptop computers, handheld computers, vehicle-mounted terminals, wearable devices, and pedometers, etc.

[0200] It should be understood that in the embodiments of the present invention, the radio frequency unit 801 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink data from the base station, it is given to the processor 810 for processing; in addition, the uplink data is sent to the base station. Generally, the radio frequency unit 801 includes, but is not limited to, antennas, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc. In addition, the radio frequency unit 801 can also communicate with the network and other devices through a wireless communication system.

[0201] The network module 802 provides the electronic device with wireless broadband Internet access, such as helping users send and receive emails, browse the web, and access streaming media, etc.

[0202] The audio output unit 803 can convert the audio data received by the radio frequency unit 801 or the network module 802 or stored in the memory 809 into an audio signal and output it as sound. Moreover, the audio output unit 803 can also provide audio output related to the specific functions executed by the electronic device 800 (for example, call signal receiving sound, message receiving sound, etc.). The audio output unit 803 includes a speaker, a buzzer, and a receiver, etc.

[0203] The input unit 804 is used to receive audio or video signals. The input unit 804 can include a graphics processing unit (GPU) 8041 and a microphone 8042. The graphics processing unit 8041 processes the image data of the still pictures or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The processed image frames can be displayed on the display unit 806. The processed image frames can be stored in the memory 809 (or other storage media) or sent via the radio frequency unit 801 or the network module 802. The microphone 8042 can receive sounds and can process such sounds into audio data. The processed audio data can be converted into a format that can be output via the radio frequency unit 801 to the mobile communication base station in the case of a phone call mode.

[0204] The electronic device 800 further includes at least one sensor 805, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 8061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 8061 and / or the backlight when the electronic device 800 is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used to identify the posture of the electronic device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; the sensor 805 can also include a fingerprint sensor, a pressure sensor, an iris sensor, a molecular sensor, a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be elaborated here.

[0205] The display unit 806 is used to display the information input by the user or the information provided to the user. The display unit 806 may include a display panel 8061, and the display panel 8061 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0206] The user input unit 807 can be used to receive input numerical or character information, and generate key signal inputs related to the user settings and function controls of the electronic device. Specifically, the user input unit 807 includes a touch panel 8071 and other input devices 8072. The touch panel 8071, also known as a touch screen, can collect the touch operations of the user on or near it (such as the operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 8071). The touch panel 8071 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 810, and receives and executes the commands sent by the processor 810. In addition, the touch panel 8071 can be implemented in a variety of types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 8071, the user input unit 807 can also include other input devices 8072. Specifically, the other input devices 8072 can include but are not limited to a physical keyboard, function keys (such as volume control buttons, power on / off buttons, etc.), a trackball, a mouse, a joystick, which will not be elaborated here.

[0207] Further, the touch panel 8071 can be covered on the display panel 8061. After the touch panel 8071 detects a touch operation on or near it, it is transmitted to the processor 810 to determine the type of touch event. Subsequently, the processor 810 provides a corresponding visual output on the display panel 8061 according to the type of touch event. Although in Figure 8 , the touch panel 8071 and the display panel 8061 are implemented as two independent components to realize the input and output functions of the electronic device, but in some embodiments, the touch panel 8071 and the display panel 8061 can be integrated to realize the input and output functions of the electronic device, and the specific implementation here is not limited.

[0208] The interface unit 808 is an interface for connecting an external device to the electronic device 800. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headset port, and so on. The interface unit 808 can be used to receive inputs from an external device (such as data information, power, etc.) and transmit the received inputs to one or more components within the electronic device 800 or can be used to transmit data between the electronic device 800 and the external device.

[0209] The memory 809 can be used to store software programs and various data. The memory 809 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 809 can include a high-speed random access memory and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0210] The processor 810 is the control center of the electronic device, connecting various parts of the entire electronic device using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 809, and by calling data stored in the memory 809, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. The processor 810 can include one or more processing units; preferably, the processor 810 can integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 810 either.

[0211] The electronic device 800 may further include a power supply 811 (such as a battery) for powering each component. Preferably, the power supply 811 can be logically connected to the processor 810 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.

[0212] In addition, the electronic device 800 includes some functional modules not shown here and will not be elaborated further.

[0213] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0214] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0215] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention's claims, and all of them fall within the protection scope of the present invention.

Claims

1. A method for generating an aerial view of an intersection, characterized in that, The method is applied to a vehicle, and vehicle-mounted cameras are arranged around the vehicle. The vehicle-mounted cameras are used to obtain panoramic images of the vehicle, including: Determine image features of the panoramic image; Generate an initial bird's-eye view feature map for the vehicle based on the image features; Determine map elements of the initial bird's-eye view feature map; Determine feature interaction parameters for the initial bird's-eye view feature map based on the map elements; Obtain position information and attitude information of the vehicle; Generate an intersection bird's-eye view feature map based on the initial bird's-eye view feature map, the position information, the attitude information, and the feature interaction parameters.

2. The method according to claim 1, characterized in that The method has a corresponding residual network and a feature pyramid network. The step of determining image features of the panoramic image includes: Use the residual network and the feature pyramid network to determine image features of the panoramic image.

3. The method according to claim 1, characterized in that, The method has a corresponding geometric guidance kernel transformer. The step of generating an initial bird's-eye view feature map for the vehicle based on the image features includes: Determine internal and external parameters of the vehicle-mounted camera; Use the geometric guidance kernel transformer to project image points of the panoramic image into a preset bird's-eye view range based on the internal and external parameters and the image features, and generate an initial bird's-eye view feature map for the vehicle.

4. The method according to claim 1, wherein The map elements include lane dividing lines, road boundary lines, and crosswalk areas. The step of determining feature interaction parameters for the initial bird's-eye view feature map based on the map elements includes: Determine query vectors for the lane dividing lines, the road boundary lines, and the crosswalk areas respectively; Generate key vectors and value vectors based on the initial bird's-eye view feature map; Use deformable attention to determine feature interaction parameters for the initial bird's-eye view feature map based on the query vectors, the key vectors, and the value vectors.

5. The method according to claim 1, wherein The step of generating an intersection bird's-eye view feature map based on the initial bird's-eye view feature map, the position information, the attitude information, and the feature interaction parameters includes: Determine range parameters for the vehicle; Determine a plurality of initial target bird's-eye view maps from the initial bird's-eye view feature map based on the range parameters; Perform position adjustment on the initial target bird's-eye view maps based on the position information and the attitude information to generate a plurality of target bird's-eye view maps; Stitch the plurality of target bird's-eye view maps to generate a historical bird's-eye view feature map; Update the historical bird's-eye view feature map based on the feature interaction parameters to generate an intersection bird's-eye view feature map.

6. The method according to claim 5, wherein The step of updating the historical bird's-eye view feature map based on the feature interaction parameters to generate an intersection bird's-eye view feature map includes: Establish a bird's-eye view map query for the historical bird's-eye view feature map based on the feature interaction parameters; Generate an intersection bird's-eye view feature map based on the bird's-eye view map query and the historical bird's-eye view feature map.

7. The method according to claim 4, wherein The method is applied to an intersection bird's-eye view feature map generation model. The model includes a high-precision map element generation module, and the high-precision map element generation module is configured with a detection head. The detection head is used to output probability scores for the lane dividing lines, the road boundary lines, and the crosswalk areas.

8. An intersection aerial view generation device, characterized in that The device is applied to a vehicle, and vehicle-mounted cameras are arranged around the vehicle. The vehicle-mounted cameras are used to obtain panoramic images of the vehicle, including: an image feature determination module, configured to determine image features for the panoramic images; an initial bird's-eye feature map generation module, configured to generate an initial bird's-eye feature map for the vehicle based on the image features; a map element determination module, configured to determine map elements for the initial bird's-eye feature map; a feature interaction parameter determination module, configured to determine feature interaction parameters for the initial bird's-eye feature map based on the map elements; an attitude information acquisition module, configured to acquire position information and attitude information for the vehicle; an intersection bird's-eye feature map generation module, configured to generate an intersection bird's-eye feature map based on the initial bird's-eye feature map, the position information, the attitude information, and the feature interaction parameters.

9. An electronic device, characterized in that, Including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the intersection bird's-eye view generation method according to any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the intersection bird's-eye view generation method according to any one of claims 1-7.