Intersection topological feature point detection method and device, terminal and storage medium
By generating bird's-eye feature maps and performing object detection through multiple vehicle peripheral video data, the problem of difficulty in constructing lane-free line intersection topological information in the existing technology is solved, and accurate detection of the location and topological relationship of the intersection topological feature point positions and topological relationships is achieved.
Patent Information
- Application Number
- CN202311739815.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-20
AI Technical Summary
It is difficult for prior art to construct the intersection topological information of some intersections through lane lines, especially in lane-free edge areas.
By obtaining the vehicle perimeter video data of multiple vehicles and multiple times, a bird's-eye feature map is generated, and the target detection method is used to detect the topological feature points of the intersection in the intersection feature map to obtain their location information and topological relationship information.
Without reference to lane lines, the location information and topological relationship information of the topological characteristic points of the intersection can be accurately constructed. It is suitable for intersections where there are lane-free edge lines, improving the accuracy of path planning.
Smart Images

Figure CN120183170A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of assisted driving technology, and specifically relates to a method, device, terminal and storage medium for detecting intersection topological feature points. Background Technique
[0002] Intersection topological information can reflect the lane connection relationship within an intersection. In the field of assisted driving, accurately generating intersection topological information helps improve the accuracy of path planning. Currently, intersection topological information is usually characterized based on lane lines using a high-precision map, that is, first generating a center line through the lane lines, and then generating intersection topological information through the center line. However, there are areas without lane boundaries in some intersections, making it difficult to construct intersection topological information through lane lines. Therefore, the existing methods for characterizing intersection topological information have a small scope of application.
[0003] Therefore, the existing technology still needs to be improved. Summary of the Invention
[0004] The present application provides a method, device, terminal and storage medium for detecting intersection topological feature points to solve the problem in related technologies that there are areas without lane boundaries in some intersections, making it difficult to construct intersection topological information through lane lines.
[0005] To achieve the above object, the present application adopts the following technical solutions:
[0006] A method for detecting intersection topological feature points includes the following steps:
[0007] Obtain a plurality of vehicle panoramic video data corresponding to a target intersection;
[0008] For each of the vehicle panoramic video data, generate an aerial view feature map based on a plurality of captured images at each moment in the vehicle panoramic video data to obtain a plurality of the aerial view feature maps, where each of the captured images corresponds to a different panoramic camera;
[0009] Generate an intersection feature map corresponding to the target intersection based on all the aerial view feature maps obtained from the vehicle panoramic video data;
[0010] Detect intersection topological feature points of the target intersection according to the intersection feature map to obtain position information and topological relationship information of a plurality of the intersection topological feature points.
[0011] According to the above technical means, in the embodiment of the present application, the intersection feature map can be accurately constructed by obtaining the vehicle panoramic video data of multiple vehicles multiple times, and the topological information of the intersection is generated based on the intersection feature map by using the object detection method. The present invention does not need to refer to the lane lines, but generates the position information and topological relationship information of the intersection topological feature points through the bird's-eye view feature map and the object detection method, and can be applied to intersections with areas without lane lines. It solves the problem in the prior art that there are areas without lane lines in some intersections, and it is difficult to construct the intersection topological information through the lane lines.
[0012] Optionally, in an embodiment of the present application, generating a bird's-eye view feature map according to a plurality of captured images at each moment in the vehicle panoramic video data includes:
[0013] Generating a feature map according to each of the captured images at this moment to obtain a plurality of feature maps corresponding to this moment;
[0014] Converting all the feature maps to the bird's-eye view space for fusion to obtain the bird's-eye view feature map corresponding to this moment in the vehicle coordinate system.
[0015] According to the above technical means, the captured images of each panoramic camera in the embodiment of the present application can reflect the scenes at different shooting angles. Feature extraction is performed on the captured images of each panoramic camera, and then they are converted to the bird's-eye view space for fusion, so as to obtain a bird's-eye view feature map that accurately represents the vehicle environment and can effectively avoid occlusion.
[0016] Optionally, in an embodiment of the present application, generating the intersection feature map corresponding to the target intersection according to all the bird's-eye view feature maps obtained based on the vehicle panoramic video data includes:
[0017] Determining the intersection range according to the target intersection;
[0018] Converting all the bird's-eye view feature maps to the global coordinate system, and determining a plurality of target bird's-eye view feature maps according to the intersection range and all the converted bird's-eye view feature maps;
[0019] Fusing all the target bird's-eye view feature maps to obtain the intersection feature map.
[0020] According to the above technical means, after the coordinate systems of all the bird's-eye view feature maps are unified in the embodiment of the present application, multiple target bird's-eye view feature maps located within the intersection range can be accurately selected. And a feature fusion method is used to fuse each target bird's-eye view feature map into one feature map, so as to obtain a more real and reliable intersection feature map.
[0021] Optionally, in an embodiment of the present application, fusing all the target bird's-eye view feature maps to obtain the intersection feature map includes:
[0022] Based on all the target bird's-eye view feature maps, determine the feature vectors corresponding to different positions within the intersection range. For each position, fuse the feature vectors corresponding to this position in each of the target bird's-eye view feature maps to obtain the feature vector corresponding to this position;
[0023] Based on the feature vectors of all the positions, obtain the intersection feature map.
[0024] According to the above technical means, in the embodiment of the present application, for each position within the intersection range, if this position is covered by multiple target bird's-eye view feature maps, fuse the multiple feature vectors corresponding to this position. After fusion, a more accurate feature vector corresponding to this position can be obtained, thereby improving the reliability of the intersection feature map.
[0025] Optionally, in an embodiment of the present application, the detecting the intersection topology feature points of the target intersection according to the intersection feature map to obtain the position information and topological relationship information of several intersection topology feature points includes:
[0026] Perform target detection according to the intersection feature map to obtain the position information of several intersection topology feature points;
[0027] Based on the position information of all the intersection topology feature points and the intersection feature map, determine the topological relationship information of all the intersection topology feature points.
[0028] According to the above technical means, in the embodiment of the present application, by analyzing the intersection feature map through a target detection method, the position information of the intersection topology feature points can be quickly generated. By analyzing the position information of each intersection topology feature point and the feature vectors on the intersection feature map, the connection relationship between each intersection topology feature point can be accurately determined.
[0029] Optionally, in an embodiment of the present application, the determining the topological relationship information of all the intersection topology feature points based on the position information of all the intersection topology feature points and the intersection feature map includes:
[0030] Generate a bird's-eye view feature matrix corresponding to all the intersection topology feature points based on the position information of all the intersection topology feature points and the intersection feature map through an attention mechanism;
[0031] Based on the bird's-eye view feature matrix, determine the topological relationship information of all the intersection topology feature points.
[0032] According to the above technical means, in the embodiment of the present application, the information interaction and fusion between the topological feature points of each intersection can be realized through the attention mechanism, so as to more accurately obtain the feature vectors of the topological feature points of each intersection and generate an aerial view feature matrix. The connection relationship between the topological feature points of each intersection can be accurately judged through the aerial view feature matrix.
[0033] Optionally, in an embodiment of the present application, the determining the topological relationship information of all the intersection topological feature points according to the aerial view feature matrix includes:
[0034] For two of the intersection topological feature points, stack the feature vectors of the two intersection topological feature points in the aerial view feature matrix to obtain a stacked vector;
[0035] Judge whether there is a connection relationship between the two intersection topological feature points according to the stacked vector;
[0036] Determine the topological relationship information of all the intersection topological feature points according to the judgment results between every two of all the intersection topological feature points.
[0037] According to the above technical means, in the embodiment of the present application, the feature vectors of the intersection topological feature points in the aerial view feature matrix are superimposed pairwise, and through the stacked vector, it can be quickly and accurately analyzed whether there is a connection relationship between the corresponding two intersection topological feature points.
[0038] An embodiment of the second aspect of the present application provides a detection device for intersection topological feature points, including:
[0039] A data acquisition module, configured to acquire a plurality of vehicle panoramic video data corresponding to a target intersection;
[0040] A feature generation module, configured to generate an aerial view feature map for each of the vehicle panoramic video data according to a plurality of captured images at each moment in the vehicle panoramic video data, so as to obtain a plurality of the aerial view feature maps, wherein each of the captured images corresponds to a different panoramic camera;
[0041] A feature fusion module, configured to generate an intersection feature map corresponding to the target intersection according to all the aerial view feature maps obtained based on the vehicle panoramic video data;
[0042] A target detection module, configured to detect the intersection topological feature points of the target intersection according to the intersection feature map, and obtain the position information and topological relationship information of a plurality of the intersection topological feature points.
[0043] In the third aspect of the embodiments of the present application, a terminal device is provided. The terminal device includes a memory, a processor, and a detection program for intersection topological feature points stored in the memory and executable on the processor. When the processor executes the detection program for intersection topological feature points, the steps of the detection method for intersection topological feature points as described in any one of the above are implemented.
[0044] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. A detection program for intersection topological feature points is stored on the computer-readable storage medium. When the detection program for intersection topological feature points is executed by a processor, the steps of the detection method for intersection topological feature points as described in any one of the above are implemented.
[0045] Advantages of the present application:
[0046] (1) By acquiring the vehicle panoramic video data of multiple vehicles multiple times, the present application can accurately construct an intersection feature map. Through the bird's-eye view feature map and the target detection method, the position information and topological relationship information of the intersection topological feature points are generated, without referring to lane lines, and can be applied to intersections with areas without lane boundaries. The problem in the prior art that it is difficult to construct intersection topological information through lane lines due to the existence of areas without lane boundaries in some intersections is solved.
[0047] (2) The captured images of each panoramic camera in the present application can reflect the scenes at different shooting angles. Through the captured images of each panoramic camera, a bird's-eye view feature map that accurately represents the vehicle environment can be obtained. Multiple target bird's-eye view feature maps within the intersection range are screened out, and through the feature fusion method, the feature vectors at different positions within the intersection range can be accurately obtained, and then a real and reliable intersection feature map is generated.
[0048] (3) By analyzing the intersection feature map through the target detection method, the position information of the intersection topological feature points can be quickly generated. By analyzing the position information of each intersection topological feature point and the feature vectors on the intersection feature map, the connection relationship between each intersection topological feature point can be accurately determined.
[0049] (4) Through the attention mechanism, the present application can realize the information interaction and fusion between each intersection topological feature point, so as to more accurately obtain the feature vectors of each intersection topological feature point. The feature vectors of each intersection topological feature point are superimposed pairwise, and through the stacked vectors, it can be quickly and accurately analyzed whether there is a connection relationship between two intersection topological feature points.
[0050] (5) Instead of adopting the road topology prediction paradigm based on lane lines / center lines, this application uses a brand-new road section topology representation method of intersection topology feature points. The intersection topology feature points are at the road level accuracy. Therefore, the accuracy requirements for vehicle positioning and feature stitching are relatively low, and it has a certain robustness to positioning errors, which can effectively reduce the difficulty of the prediction task, and can better match the bird's-eye view perception paradigm. At the same time, this representation method is also more in line with human driving habits.
[0051] Additional aspects and advantages of this application will be given in part in the following description, will become apparent in part from the following description, or will be learned through the practice of this application. Brief Description of the Drawings
[0052] The above-mentioned and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0053] Figure 1 is a schematic flowchart of the method for detecting intersection topology feature points in an embodiment of this application;
[0054] Figure 2 is a schematic flowchart of the generation process of the bird's-eye view feature map in an embodiment of this application;
[0055] Figure 3 is a schematic flowchart of the generation process of the intersection feature map in an embodiment of this application;
[0056] Figure 4 is a schematic flowchart of the generation process of the position of the intersection topology feature points in an embodiment of this application;
[0057] Figure 5 is a schematic flowchart of the attention mechanism and connection relationship recognition in an embodiment of this application;
[0058] Figure 6 is a schematic structural diagram of the device for detecting intersection topology feature points in an embodiment of this application;
[0059] Figure 7 is a schematic structural diagram of the target detection module in an embodiment of this application;
[0060] Figure 8 is a schematic block diagram of the internal structure principle of the terminal device provided in an embodiment of this application. Detailed Description of the Embodiments
[0061] The embodiments of this application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain this application and should not be construed as a limitation of this application.
[0062] The following describes a method, apparatus, terminal, and storage medium for detecting intersection topological feature points according to an embodiment of the present application. Regarding the problem mentioned in the above background technology that there are areas without lane lines at some intersections, making it difficult to construct intersection topological information through lane lines, the present application provides a method for detecting intersection topological feature points. In this method, a plurality of vehicle panoramic video data corresponding to a target intersection is obtained; for each of the vehicle panoramic video data, a bird's-eye view feature map is generated based on a plurality of captured images at each moment of the vehicle panoramic video data to obtain a plurality of the bird's-eye view feature maps, where each of the captured images corresponds to a different panoramic camera; an intersection feature map corresponding to the target intersection is generated based on all the bird's-eye view feature maps obtained from the vehicle panoramic video data; and intersection topological feature points of the target intersection are detected according to the intersection feature map to obtain position information and topological relationship information of the plurality of intersection topological feature points. The present application does not need to refer to lane lines, but generates position information and topological relationship information of intersection topological feature points through bird's-eye view feature maps and object detection methods, and can be applied to intersections with areas without lane lines. It solves the problem in the prior art that there are areas without lane lines at some intersections, making it difficult to construct intersection topological information through lane lines.
[0063] For example, taking intersection A as the target intersection, vehicle panoramic video data of vehicles a, b, c, d, e, and f passing through intersection A is obtained. Taking vehicle a as an example, the vehicle panoramic video data of vehicle a has a total of twenty moments, and five panoramic cameras are arranged on vehicle a. Then, at each moment of the vehicle panoramic video data of vehicle a, there are captured images respectively collected by five panoramic cameras. A bird's-eye view feature map of vehicle a at this moment is generated jointly by the five captured images corresponding to each moment. Based on the vehicle panoramic video data of vehicle a, a total of twenty bird's-eye view feature maps can be generated. Combining all the bird's-eye view feature maps generated from the vehicle panoramic video data of vehicles a, b, c, d, e, and f, an intersection feature map of intersection A is constructed, which can better cover intersection A. Performing object detection on this intersection feature map can obtain position information and topological relationship information of multiple intersection topological feature points within intersection A.
[0064] Specifically, Figure 1 It is a schematic flowchart of a method for detecting intersection topological feature points provided by an embodiment of the present application.
[0065] As Figure 1 shown, the method for detecting intersection topological feature points includes the following steps:
[0066] Step S100: Obtain a plurality of vehicle panoramic video data corresponding to a target intersection;
[0067] Specifically, the target intersection can be any intersection. For example, an intersection can be selected from a navigation map as the target intersection. In order to obtain the road conditions within the target intersection, this embodiment needs to acquire multiple vehicle panoramic video data of the target intersection. Multiple panoramic cameras are installed on a vehicle, and each vehicle panoramic video includes the captured images taken by each panoramic camera on the vehicle at different times. Therefore, each vehicle panoramic video data can reflect the road scenes captured at different times and from different angles when the vehicle passes through the target intersection. These vehicle panoramic video data can be obtained from different vehicles or from different trips of the same vehicle.
[0068] Step S200: For each of the vehicle panoramic video data, generate an aerial view feature map according to a plurality of captured images at each moment in the vehicle panoramic video data, so as to obtain a plurality of the aerial view feature maps, where each of the captured images corresponds to a different panoramic camera.
[0069] Taking a moment of a vehicle panoramic video data as an example, there are images captured by different panoramic cameras on the same vehicle at this moment. An aerial view feature map of the vehicle at this moment is jointly generated according to all the captured images corresponding to this moment. The aerial view perspective means observing the scene from an aerial view angle. Therefore, the aerial view feature map can be regarded as the position feature of each point in the scene. Multiple aerial view feature maps can be generated through the vehicle panoramic video data of multiple vehicles and multiple times, so as to comprehensively present the road conditions of the target intersection.
[0070] In one embodiment, the generating an aerial view feature map according to a plurality of captured images at each moment in the vehicle panoramic video data includes:
[0071] Generate a feature map according to each of the captured images at this moment, so as to obtain a plurality of feature maps corresponding to this moment;
[0072] Convert all the feature maps to the aerial view space for fusion, so as to obtain the aerial view feature map corresponding to this moment in the ego-vehicle coordinate system.
[0073] Specifically, taking a moment in a vehicle panoramic video data as an example, there are multiple captured images corresponding to this moment, and each captured image is used to reflect the road scene captured by a panoramic camera on the vehicle. Feature extraction is respectively performed on each captured image to obtain the feature maps of each captured image. These feature maps are converted to the aerial view space (BEV space) in the ego-vehicle coordinate system for feature fusion, that is, an aerial view feature map is obtained. In this embodiment, by performing feature extraction and perspective conversion on the captured images of each panoramic camera, an aerial view feature map represented by tensor data can be obtained, and this aerial view feature map has certain position features.
[0074] For example, Figure 2As shown in the figure, a feature extraction network is pre-constructed using ResNet50 and FPN, and a BEV stitching network is constructed using an inverse perspective method (such as GKT or LSS, etc.). For the captured image of each panoramic camera at each moment, the feature extraction network extracts the image features of the captured image and generates a feature map through a multi-scale fusion method. The multi-scale fusion method is specifically as follows: Extract the feature maps of the 0th, 1st, 2nd, and 3rd layers of ResNet50. The sizes of the four feature maps of ResNet50 are: 56×56×256, 28×28×512, 14×14×1024, 7×7×2048. Given that the sizes and number of channels of the four feature maps are different, the number of channels is unified through FPN, and the sizes of the above four feature maps are unified to: 56×56×256, 28×28×256, 14×14×256, 7×7×256. For the three feature maps with sizes of 28×28, 14×14, and 7×7, bilinear interpolation is used to upsample them to 56×56. Finally, addition is performed in the channel dimension to obtain the final feature map of the captured image, with a size of 56×56×256. After the feature maps of each panoramic camera are generated, through the BEV stitching network and the internal and external parameter matrices of the panoramic camera, the feature maps of each panoramic camera at the same moment are projected into the bird's-eye view space for stitching and fusion. Among them, each panoramic camera has only one internal and external parameter matrix, which is used for the conversion between the camera coordinate system and the world coordinate system, and the conversion between the camera coordinate system and the image coordinate system. After the feature maps are fused in the bird's-eye view space, the bird's-eye view feature map in the ego-vehicle coordinate system is output, and a bird's-eye view feature map in the ego-vehicle coordinate system is generated at each time step.
[0075] In one embodiment, each of the vehicle panoramic video data is an original image data matrix, and the original image data matrix is composed of image data matrices respectively corresponding to each panoramic camera. Each of the image data matrices is composed of image data collected by a panoramic camera at each moment.
[0076] Specifically, the image data directly collected by the panoramic camera is three-channel image data with time attributes, which can be represented as a group of four-dimensional matrices. A vehicle can generate N four-dimensional matrices equal to the number of panoramic cameras at one moment. Let the dimension of the originally collected image data be (C, H, W), where C represents the number of channels, H represents the height of the image, and W represents the width of the image. Suppose there are T moments of image data and N panoramic cameras, then the original image data matrix is represented as: X = [X1, X2,..., X N , where X i is the image data matrix of the i-th panoramic camera at T moments, with a dimension of (C, H, W).
[0077] For example, taking two panoramic cameras and three moments as an example, the image data matrices of the two panoramic cameras are respectively: X1 = [X 11 , X 12 , X 13 , with dimensions (C, H, W); X2 = [X 21 , X 22 , X 23 , with dimensions (C, H, W). The original image data matrix is then represented as: X = [X1, X2], with dimensions (N, T, C, H, W), where N = 2 indicates there are two panoramic cameras, T = 3 indicates there are image data at three moments, C represents the number of channels (for example, it can be taken as 3), H represents the height of the image, and W represents the width of the image.
[0078] In one embodiment, the size of each of the bird's-eye view feature maps is determined based on the bird's-eye view perception range and camera parameters.
[0079] Specifically, the center point of the bird's-eye view feature map is the center of the vehicle, and the range along the vehicle's forward direction and the vertical direction is the bird's-eye view perception range (BEV perception range), which can be flexibly configured according to the capabilities of the panoramic cameras. The bird's-eye view feature map corresponding to one moment is a three-dimensional tensor with dimensions Hb × Wb × C, where Hb and Wb respectively represent the dimensions along the forward direction of the vehicle itself and the vertical direction, and C represents the number of channels. It can be understood that the bird's-eye view feature map in the vehicle's coordinate system has spatial characteristics, and its coverage range is consistent with the BEV perception range. For the relative two-dimensional spatial positions within the BEV perception range, the corresponding feature vectors can be determined according to the resolution and the range of the bird's-eye view feature map.
[0080] For example, assuming the actual BEV perception range is 50 meters in front and behind the vehicle itself and 30 meters on the left and right, and the resolution is set to 0.2 meters, then Hb can be calculated as 100 meters / 0.2 meters = 500, and Wb can be calculated as 60 meters / 0.2 meters = 300. Since the bird's-eye view feature map has a time attribute, the bird's-eye view feature map can be represented in the following form: F = [F1, F2,..., F T , where F i represents the bird's-eye view feature map at the i-th moment, with dimensions (Hb, Wb, C), for example, 500 × 300 × 256 (500 and 300 are determined based on the resolution and the bird's-eye view perception range, and 256 is determined based on the number of channels of the feature map). Taking the bird's-eye view feature map data at three moments as an example, it can be represented as: F = [F1, F2, F3], with dimensions (T, Hb, Wb, C), where T represents the number of moments, Hb and Wb respectively represent the dimensions along the forward direction of the vehicle itself and the vertical direction, and C represents the number of channels.
[0081] Step S300: Generate the intersection feature map corresponding to the target intersection based on all the bird's-eye view feature maps obtained from the vehicle panoramic video data of each vehicle.
[0082] Specifically, each bird's-eye view feature map can reflect the position features of some positions within the target intersection. Therefore, integrating all the bird's-eye view feature maps together can more comprehensively present the road conditions of the target intersection, and then accurately generate the intersection feature map corresponding to the target intersection.
[0083] In one embodiment, the generating the intersection feature map corresponding to the target intersection based on all the bird's-eye view feature maps obtained from the vehicle panoramic video data of each vehicle includes:
[0084] Determine the intersection range according to the target intersection;
[0085] Convert all the bird's-eye view feature maps to the global coordinate system, and determine several target bird's-eye view feature maps according to the intersection range and all the converted bird's-eye view feature maps;
[0086] Fuse all the target bird's-eye view feature maps to obtain the intersection feature map.
[0087] Specifically, the range obtained by expanding the target intersection by a certain distance can be used as the intersection range. In this embodiment, it is limited that the feature fusion of multiple generated bird's-eye view feature maps is only carried out within the intersection range. As Figure 3 shown, first, each bird's-eye view feature map can be converted from the ego-vehicle coordinate system to the global coordinate system according to the vehicle pose information at the current moment, and multiple bird's-eye view feature maps with overlapping areas are obtained after conversion. Then, the bird's-eye view feature maps located within the intersection range are screened out. If the coverage range of a certain bird's-eye view feature map is partially within the intersection range and partially outside the intersection range, only the partial feature maps located within the intersection range need to be screened out. After screening, multiple target bird's-eye view feature maps are obtained. Finally, each target bird's-eye view feature map is merged into one feature map through the method of feature fusion, that is, a real and reliable intersection feature map is obtained.
[0088] For example, assume that the expansion distance is 50 meters, then the intersection range is a range of 100 meters × 100 meters. The intersection feature map obtained by fusion is a high-dimensional representation tensor for reflecting intersection information and has an absolute position attribute. The intersection feature map can be expressed as:
[0089]
[0090] Among them, B is the bird's-eye view feature map of the target intersection, Hr and Wr are the feature map grid sizes within the intersection range. For example, if the intersection is 100 meters by 100 meters in size and the resolution is 0.2 meters, then Hr is 100 / 0.2 = 500, and Wr is 100 / 0.2 = 500. dr is the feature size corresponding to each position (for example, taking the value 256).
[0091] In one embodiment, the step of fusing all the target bird's-eye view feature maps to obtain the intersection feature map includes:
[0092] Determine the feature vectors corresponding to different positions within the intersection range according to all the target bird's-eye view feature maps. Among them, for each position, fuse the feature vectors corresponding to this position in each of the target bird's-eye view feature maps to obtain the feature vector corresponding to this position;
[0093] Obtain the intersection feature map according to the feature vectors of all the positions.
[0094] Specifically, if the same position within the intersection range is covered by multiple bird's-eye view feature maps, feature fusion needs to be performed at this position. The method of feature fusion is to fuse the feature vectors of this position in different bird's-eye view feature maps into one feature vector. The feature vectors of different positions within the intersection range can form the intersection feature map of the target intersection. In this embodiment, through the method of feature fusion, the feature vectors of each position within the intersection range can be obtained more accurately, thereby improving the reliability of the intersection feature map.
[0095] In one embodiment, the step of fusing the feature vectors corresponding to this position in each of the target bird's-eye view feature maps to obtain the feature vector corresponding to this position includes:
[0096] Perform data screening on each feature vector;
[0097] Perform an averaging operation on the screened feature vectors to obtain the feature vector corresponding to this position.
[0098] Specifically, as Figure 3 shown, before performing feature fusion on multiple feature vectors at the same position, it is also necessary to perform data screening on each feature vector to eliminate outliers, for example, eliminating outliers outside 5%. Take the average of the multiple feature vectors screened after eliminating outliers to obtain the specific value of the feature vector at this position.
[0099] In one embodiment, if the actual coverage range of the intersection feature map is smaller than the intersection range, adjust the intersection range according to the actual coverage range.
[0100] Specifically, if the aerial view feature maps are fused into an intersection feature map and the intersection range cannot be fully covered, the intersection range can be cropped according to the actual coverage range of the intersection feature map. For example, the intersection range can be cropped into the circumscribed rectangle of the actual coverage range of the intersection feature map.
[0101] Step S400: Detect the intersection topological feature points of the target intersection according to the intersection feature map, and obtain the position information and topological relationship information of a plurality of the intersection topological feature points.
[0102] Specifically, in this embodiment, there is no need to refer to lane lines. Instead, through the aerial view feature map and the object detection method, the position information and topological relationship information of the intersection topological feature points are generated. The intersection topological feature points are a type of point element, and their geometric positions are located at the intersections of each section of the intersection. For example, if there is a hard isolation between the main and auxiliary roads or the same road section, two intersection topological feature points should be generated. The intersection topological feature points also have topological relationship attributes, which reflect the connection relationship between each intersection topological feature point and other intersection topological feature points in the intersection. Since this embodiment does not need to generate road topological information based on lane lines, it can be applied to intersections with areas without lane boundaries.
[0103] In one embodiment, the detecting the intersection topological feature points of the target intersection according to the intersection feature map and obtaining the position information and topological relationship information of a plurality of the intersection topological feature points includes:
[0104] Performing object detection according to the intersection feature map to obtain the position information of a plurality of the intersection topological feature points;
[0105] Determining the topological relationship information of all the intersection topological feature points according to the position information of all the intersection topological feature points and the intersection feature map.
[0106] Specifically, object detection is equivalent to the detection task in conventional image tasks, that is, inputting a picture and outputting several objects (classifying and detecting the categories of objects and / or the bounding boxes of objects). As Figure 4 shown, input the intersection feature map into a pre-constructed object detection algorithm, and quickly obtain the position information of the intersection topological feature points through the object detection algorithm. By analyzing the position information of each intersection topological feature point and the feature vectors on the intersection feature map, the connection relationship between each intersection topological feature point can be accurately judged.
[0107] Illustrating with an example, the intersection topological feature points can be point entities, and their position information is composed of (x, y) coordinates; the topological relationship information describes the ordered connection relationship between the intersection topological feature points and can be represented by a two-dimensional matrix of size N×N, that is, obtaining the adjacent matrix of connected feature points, where N is the number of intersection topological feature points, and the matrix element Eij Indicates whether the $i$-th topological feature point of the intersection is connected to the $j$-th topological feature point of the intersection. If it is connected, the value is 1; otherwise, it is 0.
[0108] In one embodiment, performing target detection according to the intersection feature map to obtain the position information of a plurality of the topological feature points of the intersection, including:
[0109] Performing target detection according to the intersection feature map to obtain the position information, category information, and confidence level corresponding to a plurality of topological feature points within the intersection range;
[0110] For each of the topological feature points, if the category information of the topological feature point is the target category and the confidence level is higher than a preset confidence threshold, then the topological feature point is used as the topological feature point of the intersection.
[0111] Specifically, the target detection algorithm generates multiple topological feature points based on the intersection feature map. Each topological feature point also has its own position information, category information, and confidence level. The topological feature points are screened through the category information and confidence level, so as to obtain the topological feature points of the intersection that can effectively represent the road topological information.
[0112] For example, the target detection can be implemented using the YOLO series model. The YOLO series model learns the positions of the topological feature points according to the intersection feature map. In an actual application scenario, the intersection feature map is input into the model, and the model generates the position information, category information (such as main road feature points, secondary road feature points, non-feature points), and the confidence level of the corresponding classification of the topological feature points. If the category information of a topological feature point is a main road feature point or a secondary road feature point and the confidence level is higher than 0.8, then the position of the topological feature point is used as the topological feature point of the intersection.
[0113] In one embodiment, determining the topological relationship information of all the topological feature points of the intersection according to the position information of all the topological feature points of the intersection and the intersection feature map, including:
[0114] Generating an aerial view feature matrix corresponding to all the topological feature points of the intersection through an attention mechanism based on the position information of all the topological feature points of the intersection and the intersection feature map;
[0115] Determining the topological relationship information of all the topological feature points of the intersection according to the aerial view feature matrix.
[0116] Specifically, in this embodiment, the attention mechanism can achieve information interaction and fusion among the topological feature points of each intersection, so as to more accurately obtain the feature vectors of the topological feature points of each intersection. By constructing a matrix with the feature vectors of the topological feature points of each intersection, the bird's-eye view feature matrix is obtained. Since the bird's-eye view feature matrix covers the feature vectors of all the topological feature points of the intersections, the connection relationship between the topological feature points of each intersection can be accurately judged through the bird's-eye view feature matrix, and then reliable topological relationship information can be obtained.
[0117] In one embodiment, the generating, by the attention mechanism, of the bird's-eye view feature matrix corresponding to all the topological feature points of the intersections based on the position information of all the topological feature points of the intersections and the intersection feature map includes:
[0118] According to the position information of each topological feature point of the intersections and the intersection feature map, obtain the feature vectors of each topological feature point of the intersections, and construct a query matrix according to the feature vectors of each topological feature point of the intersections;
[0119] Construct a key-value pair matrix according to the intersection feature map;
[0120] Based on the query matrix and the key-value pair matrix, use the attention mechanism to iteratively update the feature vectors in the query matrix;
[0121] Take the updated query matrix as the bird's-eye view feature matrix.
[0122] Specifically, in this embodiment, a query matrix is constructed using the feature vectors corresponding to the positions of the topological feature points of each intersection. That is, the query matrix is actually a feature matrix corresponding to the topological feature points of each intersection, with a dimension of N×256, where N represents the number of topological feature points passing through the intersections, and each element in the query matrix is a numerically value that can be learned and updated. And a key-value pair matrix is constructed using the intersection feature map, with a dimension of 600×300×256. The attention mechanism adopts a query-key-value mode, which can fully realize the information fusion and transmission among the topological feature points of each intersection, so as to iteratively update the feature vectors in the query matrix. The updated query matrix is the bird's-eye view feature matrix of the topological feature points of each intersection.
[0123] In one embodiment, the attention mechanism includes a multi-head attention mechanism, and each attention head of the multi-head attention mechanism includes a deformable attention mechanism.
[0124] Specifically, the dimension of each attention head can be calculated according to the dimension of the query matrix and the number of multi-heads. For example, the number of multi-heads is 8, and the dimension of each attention head is 32. Each attention head adopts a deformable attention mechanism, such as a deformable cross self-attention mechanism, so as to further improve the information interaction and fusion among the topological feature points of each intersection.
[0125] For example, a Transformer structure can be constructed to achieve information interaction between the topological feature points of each intersection. The Transformer structure first performs a linear projection on the query matrix to convert its dimension to N×32. Then, the intersection feature map is projected through a linear projection matrix to convert its dimension to 600×300×32. For each query matrix, the position coordinates of the reference point and the coordinates and weights of the corresponding four sampling points are determined through linear projection. For each attention head, the features of the reference point and the sampling points are interacted with the intersection feature map to obtain feature vectors. The feature vectors obtained by all attention heads are stacked and converted into a feature vector of N×256 through linear projection. As Figure 5 shown, by repeatedly using the Transformer structure multiple times, the fusion and transmission of information can be fully realized.
[0126] In one embodiment, determining the topological relationship information of all the intersection topological feature points according to the bird's-eye view feature matrix includes:
[0127] For two of the intersection topological feature points, the feature vectors of the two intersection topological feature points in the bird's-eye view feature matrix are stacked to obtain a stacked vector;
[0128] According to the stacked vector, it is judged whether there is a connection relationship between the two intersection topological feature points;
[0129] According to the judgment results between every two of all the intersection topological feature points, the topological relationship information of all the intersection topological feature points is determined.
[0130] Specifically, taking two intersection topological feature points as an example, the feature vectors of the two intersection topological feature points in the bird's-eye view feature matrix are stacked in the channel direction to obtain a stacked vector. Since this stacked vector fuses the feature vectors of the two intersection topological feature points, it is possible to quickly and accurately analyze whether there is a connection relationship between the corresponding two intersection topological feature points through this stacked vector. After judging the connection relationship between every two intersection topological feature points, the topological relationship information is obtained, which can be used for auxiliary driving planning and control.
[0131] In one embodiment, judging whether there is a connection relationship between every two of the intersection topological feature points according to the stacked vectors includes:
[0132] For each stacked vector, the stacked vector is projected into a low-dimensional space to obtain a low-dimensional stacked vector;
[0133] According to the low-dimensional stacked vector, the probability value that the corresponding two intersection topological feature points have a connection relationship is determined;
[0134] If the probability value is greater than a preset probability threshold, it is determined that there is a connection relationship between the corresponding two intersection topological feature points.
[0135] Specifically, the feature vectors of any two intersection topological feature points in the bird's-eye view feature matrix are superimposed in the channel direction to form a stacked vector. The stacked vector is projected into a low-dimensional space to predict the connection relationship between the two intersection topological feature points, and a probability value is calculated. If the probability value is greater than the preset probability threshold, indicating that the possibility of a connection relationship between the two intersection topological feature points is relatively large, it is determined that there is a connection relationship between them; if the probability value is less than the preset probability threshold, indicating that the possibility of a connection relationship between the two intersection topological feature points is relatively small, it is determined that there is no connection relationship between them.
[0136] For example, two multi-layer perceptrons (MLPs) can be pre-constructed. The first MLP is a two-dimensional matrix with a size of 512×64, where each element is a pre-trained parameter. The input data of the first MLP is a query matrix with a size of N×64, and its function is to project the input data into a low-dimensional space. The second MLP is a two-dimensional matrix with a size of 64×1, where each element is a pre-trained parameter. The input data of the second MLP is the output data of the first multi-layer perceptron, and its function is to predict the connection relationship / topological relationship between the two intersection topological feature points. According to the output value (i.e., the probability value) of the second MLP and in combination with the preset threshold, it can be determined whether the two intersection topological feature points are orderly connected.
[0137] In summary, in the present application, the intersection feature map can be accurately constructed by obtaining the vehicle panoramic video data of multiple vehicles multiple times. The position information and topological relationship information of the intersection topological feature points are generated by the bird's-eye view feature map and the target detection method, without referring to the lane lines, and can be applied to intersections with areas without lane boundaries. Secondly, the captured images of each panoramic camera in the present application can reflect the scenes at different shooting angles, and the bird's-eye view feature map for accurately representing the vehicle environment can be obtained through the captured images of each panoramic camera. Moreover, in the present application, multiple target bird's-eye view feature maps located within the intersection range are screened out, and the feature vectors at different positions within the intersection range can be accurately obtained through the feature fusion method, thereby generating a real and reliable intersection feature map. In addition, in the present application, the intersection feature map is analyzed by the target detection method, and the position information of the intersection topological feature points can be quickly generated. By analyzing the position information of each intersection topological feature point and the feature vectors on the intersection feature map, the connection relationship between each intersection topological feature point can be accurately determined. In addition, in the present application, the information interaction and fusion between each intersection topological feature point can be realized through the attention mechanism, so as to more accurately obtain the feature vectors of each intersection topological feature point. The feature vectors of each intersection topological feature point are superimposed pairwise, and whether there is a connection relationship between two intersection topological feature points can be quickly and accurately analyzed through the stacked vectors. Finally, the present application does not adopt the road topology prediction paradigm based on lane lines / center lines, but adopts a new section topology representation method of intersection topological feature points. The intersection topological feature points are at the road-level accuracy, so the accuracy requirements for vehicle positioning and feature stitching are relatively low, and it has a certain robustness to positioning errors, which can effectively reduce the difficulty of the prediction task, and can better match the bird's-eye view perception paradigm. At the same time, this representation method is also more in line with human driving habits.
[0138] Next, a detection device for intersection topological feature points according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0139] As Figure 6 shown, the detection device 10 for intersection topological feature points includes: a data acquisition module 100, a feature generation module 200, a feature fusion module 300, and a target detection module 400.
[0140] Specifically, the data acquisition module 100 is configured to acquire a plurality of vehicle panoramic video data corresponding to a target intersection;
[0141] The feature generation module 200 is configured to generate a bird's-eye view feature map for each of the vehicle panoramic video data according to a plurality of captured images at each moment in the vehicle panoramic video data, so as to obtain a plurality of the bird's-eye view feature maps, wherein each of the captured images corresponds to a different panoramic camera;
[0142] A feature fusion module 300, configured to generate an intersection feature map corresponding to the target intersection according to all the bird's-eye view feature maps obtained based on the vehicle panoramic video data.
[0143] A target detection module 400, configured to detect intersection topology feature points of the target intersection according to the intersection feature map, and obtain position information and topology relationship information of a plurality of the intersection topology feature points.
[0144] In one embodiment, as Figure 7 shown, the target detection module 400 includes:
[0145] An intersection topology feature point position generation module, configured to perform target detection according to the intersection feature map, and obtain position information of a plurality of the intersection topology feature points;
[0146] An intersection topology feature point connection relationship generation module, configured to determine connection relationship information between the intersection topology feature points according to the position information of each intersection topology feature point and the intersection feature map.
[0147] It should be noted that the foregoing explanation of the embodiment of the method for detecting intersection topology feature points is also applicable to the device for detecting intersection topology feature points in this embodiment, and will not be elaborated here.
[0148] Figure 8 The following is a schematic structural diagram of a terminal device provided by an embodiment of the present application. The terminal device may include:
[0149] A memory 801, a processor 802, and a computer program stored on the memory 801 and executable on the processor 802.
[0150] When the processor 802 executes the program, it implements the method for detecting intersection topology feature points provided in the foregoing embodiment.
[0151] Further, the terminal device further includes:
[0152] A communication interface 803, configured to communicate between the memory 801 and the processor 802.
[0153] The memory 801 is used to store a computer program executable on the processor 802.
[0154] The memory 801 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0155] If the memory 801, the processor 802, and the communication interface 803 are implemented independently, the communication interface 803, the memory 801, and the processor 802 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 only a thick line is used in Figure 8 , but it does not mean that there is only one bus or one type of bus.
[0156] Optionally, in a specific implementation, if the memory 801, the processor 802, and the communication interface 803 are integrated on a single chip, the memory 801, the processor 802, and the communication interface 803 can communicate with each other through an internal interface.
[0157] The processor 802 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0158] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the detection method of the intersection topological feature points as described above is implemented.
[0159] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0160] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0161] Any process or method description represented in a flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0162] The logic and / or steps represented in a flowchart or described otherwise herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can read instructions from and execute instructions by the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or N wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, as the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0163] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0164] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0165] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0166] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for detecting intersection topological feature points, characterized in that, Including the following steps: Obtain a plurality of vehicle panoramic video data corresponding to the target intersection; For each of the vehicle panoramic video data, generate a bird's-eye view feature map based on a plurality of captured images at each moment in the vehicle panoramic video data, so as to obtain a plurality of the bird's-eye view feature maps, wherein each of the captured images corresponds to a different panoramic camera; Generate an intersection feature map corresponding to the target intersection according to all the bird's-eye view feature maps obtained based on the vehicle panoramic video data; Detect the intersection topological feature points of the target intersection according to the intersection feature map, and obtain the position information and topological relationship information of the plurality of intersection topological feature points.
2. The method for detecting intersection topological feature points according to claim 1, characterized in that, The generating a bird's-eye view feature map based on a plurality of captured images at each moment in the vehicle panoramic video data includes: Generate a feature map based on each of the captured images at this moment, so as to obtain a plurality of feature maps corresponding to this moment; Convert all the feature maps to the bird's-eye view space for fusion, and obtain the bird's-eye view feature map corresponding to this moment in the ego-vehicle coordinate system.
3. The method for detecting intersection topological feature points according to claim 1, characterized in that, The generating an intersection feature map corresponding to the target intersection according to all the bird's-eye view feature maps obtained based on the vehicle panoramic video data includes: Determine the intersection range according to the target intersection; Convert all the bird's-eye view feature maps to the global coordinate system, and determine a plurality of target bird's-eye view feature maps according to the intersection range and all the converted bird's-eye view feature maps; Fuse all the target bird's-eye view feature maps to obtain the intersection feature map.
4. The method for detecting intersection topological feature points according to claim 3, characterized in that, The fusing all the target bird's-eye view feature maps to obtain the intersection feature map includes: Determine the feature vectors corresponding to different positions within the intersection range according to all the target bird's-eye view feature maps, wherein for each of the positions, fuse the feature vectors corresponding to the position in each of the target bird's-eye view feature maps to obtain the feature vector corresponding to the position; Obtain the intersection feature map according to the feature vectors of all the positions.
5. The method for detecting intersection topological feature points according to claim 1, characterized in that, The detecting the intersection topological feature points of the target intersection according to the intersection feature map and obtaining the position information and topological relationship information of the plurality of intersection topological feature points includes: Perform target detection according to the intersection feature map to obtain the position information of the plurality of intersection topological feature points; Determine the topological relationship information of all the intersection topological feature points according to the position information of all the intersection topological feature points and the intersection feature map.
6. The method for detecting intersection topological feature points according to claim 5, characterized in that, The determining the topological relationship information of all the intersection topological feature points according to the position information of all the intersection topological feature points and the intersection feature map includes: generating a bird's-eye view feature matrix corresponding to all the intersection topological feature points through an attention mechanism based on the position information of all the intersection topological feature points and the intersection feature map; Determine the topological relationship information of all the intersection topological feature points according to the bird's-eye view feature matrix.
7. The method for detecting intersection topological feature points according to claim 6, characterized in that, The determining the topological relationship information of all the intersection topological feature points according to the bird's-eye view feature matrix includes: For two of the intersection topological feature points, stack the feature vectors of the two intersection topological feature points in the bird's-eye view feature matrix to obtain a stacked vector; Judge whether there is a connection relationship between two of the intersection topological feature points according to the stacked vectors; Determine the topological relationship information of all the intersection topological feature points according to the judgment results between every two of all the intersection topological feature points.
8. A device for detecting intersection topological feature points, characterized in that, Including: A data acquisition module, configured to acquire a plurality of vehicle panoramic video data corresponding to a target intersection; A feature generation module, configured to generate an aerial view feature map for each of the vehicle panoramic video data according to a plurality of captured images at each moment in the vehicle panoramic video data, so as to obtain a plurality of the aerial view feature maps, wherein each of the captured images corresponds to a different panoramic camera; A feature fusion module, configured to generate an intersection feature map corresponding to the target intersection according to all the aerial view feature maps obtained based on all the vehicle panoramic video data; A target detection module, configured to detect intersection topological feature points of the target intersection according to the intersection feature map, and obtain position information and topological relationship information of a plurality of the intersection topological feature points.
9. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a detection program for intersection topological feature points stored in the memory and executable on the processor. When the processor executes the detection program for intersection topological feature points, the steps of the detection method for intersection topological feature points according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, A detection program for intersection topological feature points is stored on the computer-readable storage medium. When the detection program for intersection topological feature points is executed by a processor, the steps of the detection method for intersection topological feature points according to any one of claims 1-7 are implemented.