A semantic map construction method, a visual positioning method and related devices
By constructing a semantic map, the semantic visual features of road image frames are obtained and their three-dimensional spatial relationships are determined, which solves the problem of insufficient vehicle visual positioning accuracy, achieves higher-precision positioning, and supports vehicle assisted driving and autonomous driving.
Patent Information
- Application Number
- CN202110592199.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-28
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-05-28
AI Technical Summary
In vehicle driver assistance and autonomous driving systems, existing technologies struggle to achieve accurate visual positioning with limited onboard sensors and computing resources, and semantic maps lack sufficient positioning accuracy.
By acquiring the semantic visual features of road image frames, the three-dimensional spatial relationship between two road image frames is determined. A semantic map is constructed by combining the spatial location information of geographical location and semantic visual features, thereby improving positioning accuracy.
It improves the positioning accuracy of semantic maps, providing higher accuracy for vehicle visual positioning and supporting precise assisted driving and autonomous driving.
Smart Images

Figure CN115409910B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of map data technology, specifically to a semantic map construction method, a visual positioning method, and related equipment. Background Technology
[0002] Visual localization of vehicles on the road is a crucial component of vehicle driver assistance and autonomous driving systems. This visual localization can be achieved using semantic maps. Semantic maps are map data generated by fusing data from multiple sensors, including visual and geolocation sensors. Using semantic maps for vehicle visual localization provides a basis for decision-making in driver assistance and autonomous driving systems.
[0003] Vehicles have limited onboard sensors and computing resources, so semantic maps need to have high positioning accuracy to achieve accurate visual positioning with limited onboard sensors and computing resources. Summary of the Invention
[0004] In view of this, embodiments of this application provide a semantic map construction method, a positioning method, an apparatus, and related devices to improve the positioning accuracy of semantic maps.
[0005] To achieve the above objectives, the embodiments of this application provide the following technical solutions.
[0006] In a first aspect, embodiments of this application provide a semantic map construction method, including:
[0007] Acquire road image frames;
[0008] Extract the semantic visual features of the road image frame, where the semantic visual features are the feature information of road elements in the road image frame;
[0009] Determine the semantic visual features in two road image frames and their relationship in three-dimensional space;
[0010] At least the spatial location information of the semantic visual features of the road image frame is determined based on the aforementioned association.
[0011] A semantic map is obtained based at least on the geographical location, semantic visual features, and spatial location information of the semantic visual features corresponding to the road image frame.
[0012] Secondly, embodiments of this application provide a visual positioning method, including:
[0013] Obtain the current road image frame and the current geographical location of the vehicle;
[0014] Based on the current geographical location, obtain matching current map data from the semantic map; and extract current semantic visual features from the current road image frame;
[0015] At least spatial location information matching the current semantic visual feature is obtained from the current map data to obtain the initial spatial location information of the current semantic visual feature;
[0016] Based on the initial spatial location information, the current spatial location information of the current semantic visual feature is determined.
[0017] Thirdly, embodiments of this application provide a semantic map construction device, including: at least one memory and at least one processor, wherein the memory stores one or more computer-executable instructions, and the processor invokes the one or more computer-executable instructions to execute the semantic map construction method as described in the first aspect above.
[0018] Fourthly, embodiments of this application provide an in-vehicle device, including: at least one memory and at least one processor, wherein the memory stores one or more computer-executable instructions, and the processor invokes the one or more computer-executable instructions to execute the visual positioning method as described in the second aspect above.
[0019] Fifthly, embodiments of this application provide a storage medium that stores one or more computer-executable instructions. When the one or more computer-executable instructions are executed, they implement the semantic map construction method as described in the first aspect above, or the visual positioning method as described in the second aspect above.
[0020] The semantic map construction method provided in this application can acquire road image frames; extract semantic visual features from the road image frames, where the semantic visual features are feature information of road elements in the road image frames; determine the correlation between semantic visual features in two road image frames in three-dimensional space; determine the spatial location information of the semantic visual features of the road image frames based at least on the correlation; and obtain a semantic map based at least on the geographical location corresponding to the road image frames, the semantic visual features, and the spatial location information of the semantic visual features. Since the correlation between semantic visual features in two road image frames in three-dimensional space can represent the change relationship of road element feature information with vehicle movement in three-dimensional space, determining the spatial location information of the semantic visual features based on the correlation allows the spatial location information of the semantic visual features to be combined with the change relationship of road elements with vehicle movement, thereby obtaining a more accurate spatial location information of the semantic visual features in three-dimensional space. Furthermore, constructing a semantic map based on the accurately obtained spatial location information of the semantic visual features can improve the positioning accuracy of the semantic map, providing a possibility for achieving accurate vehicle visual positioning. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0022] Figure 1 System architecture diagram for building semantic maps.
[0023] Figure 2 A flowchart of a semantic map construction method provided in an embodiment of this application.
[0024] Figure 3 This is a schematic diagram illustrating the stages of the semantic map construction process provided in the embodiments of this application.
[0025] Figure 4 A flowchart illustrating the implementation of inter-frame road element matching provided in this application embodiment.
[0026] Figure 5 A flowchart of the visual positioning method provided in the embodiments of this application.
[0027] Figure 6 This is a schematic diagram illustrating the stages of the visual positioning process provided in an embodiment of this application.
[0028] Figure 7A block diagram of a semantic map construction apparatus provided in an embodiment of this application.
[0029] Figure 8 A block diagram of a semantic map building device provided in an embodiment of this application.
[0030] Figure 9 A block diagram of a visual positioning device provided in an embodiment of this application. Detailed Implementation
[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0032] Semantic maps, as a type of map data, can be constructed based on information collected by data collection vehicles traveling on roads. The information collected by these vehicles can be from multiple sensors mounted on them, such as visual sensors, geolocation sensors, and inertial navigation sensors. In some embodiments, if the data collection vehicle has sufficient computing resources, it can independently construct a semantic map based on the collected information and upload it to the cloud for storage. In other embodiments, the data collection vehicle can also upload the collected information to a cloud server, where the cloud server constructs the semantic map based on the information collected by the vehicle and stores it in the cloud.
[0033] Figure 1 An exemplary system architecture for building semantic maps is shown. For example... Figure 1 As shown, the system may include a data collection vehicle 100 and a cloud platform 200. The data collection vehicle 100 may be equipped with a visual sensor 110, a geolocation sensor 120, and an inertial navigation sensor 130. In some further embodiments, the data collection vehicle 100 may also be equipped with a processor chip for data computation and processing. The cloud platform 200 may include a cloud server 210 and a cloud database 220. If a semantic map is constructed by the cloud platform based on the information collected by the data collection vehicle 100, the semantic map construction process may be specifically executed by the cloud server 210. The cloud database 220 may store the semantic map constructed by the cloud server or the semantic map constructed by the data collection vehicle itself.
[0034] In this embodiment, the vision sensor 110 can be used to acquire road images while the vehicle is in motion. The road images acquired by the vision sensor 110 may include multiple road image frames. In some embodiments, the vision sensor 110 may include an onboard camera, such as a monocular camera.
[0035] The geolocation sensor 120 can be used to locate the vehicle's geographical location (e.g., the vehicle's latitude and longitude coordinates) while the vehicle is in motion. In some embodiments, the geolocation sensor 120 may include a GPS (Global Positioning System) positioning device. Of course, the geolocation sensor 120 may also support other positioning methods, not limited to GPS positioning methods, such as positioning methods based on BeiDou satellites.
[0036] The inertial navigation sensor 130 can be used to determine pose information during vehicle movement. In some embodiments, the inertial navigation sensor 130 may include an IMU (Inertial Measurement Unit). The IMU includes three single-axis accelerometers and three single-axis gyroscopes. The accelerometers detect the acceleration signals of the object along three independent axes of the carrier coordinate system, while the gyroscopes detect the angular velocity signals of the carrier relative to the navigation coordinate system. The IMU calculates the pose information of the object by measuring its angular velocity and acceleration in three-dimensional space.
[0037] In some embodiments, the visual sensor 110, the geolocation sensor 120, and the inertial navigation sensor 130 may be integrated into the same in-vehicle device. This in-vehicle device may be, for example, a dashcam. Alternatively, the visual sensor 110, the geolocation sensor 120, and the inertial navigation sensor 130 may also be directly mounted on the vehicle.
[0038] As an optional implementation, Figure 2 An optional flow of the semantic map construction method provided in the embodiments of this application is illustrated by way of example. Figure 2 The process shown can be implemented by a data collection vehicle or by a cloud server. If the semantic map is constructed by the cloud server, the cloud server can acquire information collected by multiple sensors on the vehicle (such as road image frames collected by the visual sensor 110, the geographical location of the vehicle determined by the geolocation sensor 120, the pose information determined by the inertial navigation sensor 130, etc.), and construct a semantic map based on the acquired information collected by multiple sensors.
[0039] It should be noted that, in addition to the traditional vehicle geolocation matching function, semantic maps also need to provide spatial location information of the road where the vehicle is located, based on the semantic visual features. Therefore, the construction of a semantic map can be achieved by combining the geographical location of the road image frame, the semantic visual features of the road image frame, and the spatial location information of the semantic visual features. Among them, semantic visual features can be considered as the feature information of road elements in the road image frame on the image pixels.
[0040] Reference Figure 2As shown, the process may include the following steps.
[0041] In step S210, a road image frame is acquired.
[0042] The road image frames can be provided by the vehicle's vision sensor 110 (e.g., a monocular camera). In some embodiments, the road image frames acquired in step S210 can be key road image frames. Key road image frames can be road image frames with high image quality acquired by the vision sensor 110. For example, in this embodiment, road image frames with image quality higher than a preset image quality can be used as key road image frames. In some embodiments, this embodiment can use a sparse SLAM (simultaneous localization and mapping) algorithm model based on ORB (Oriented FAST and Rotated BRIEF) feature points to filter key road image frames from the road image frames acquired by the vision sensor 110. Of course, this embodiment can also support the road image frames acquired in step S210 being any road image frame acquired by the vision sensor 110.
[0043] In step S211, the semantic visual features of the road image frame are extracted.
[0044] For each road image frame obtained in step S210, this embodiment of the application can extract feature information of road elements in the road image frame to obtain semantic visual features of the road image frame. In some embodiments, the road elements can be standard road elements in the road. Standard road elements can include standard ground elements and standard roadside elements. Standard ground elements refer to road elements that serve as traffic signs on the road surface, such as lane lines and ground markings (e.g., driving direction markings, speed limit signs, dedicated lane markings, etc.). Standard roadside elements can include road poles (e.g., streetlight poles, traffic signs, etc.) beside the road. It should be noted that the semantic visual features extracted in step S211 are not limited to feature information of standard road elements; semantic visual features can also cover other road elements (e.g., intersections, pedestrian overpasses, etc.).
[0045] In some embodiments, the semantic visual features may include at least the skeleton of a road element, which may represent the outline of the road element. In further embodiments, the semantic visual features may also include structural key points representing the key structures of the road element.
[0046] In optional implementations, due to the numerous types of road elements (e.g., lane lines, ground markings, road poles, etc.), this application embodiment can utilize a multi-task convolutional neural network based on the CenterNet algorithm to extract semantic visual features from road image frames in order to identify and detect these diverse types of road elements. For example, this application embodiment can utilize a multi-task convolutional neural network based on the CenterNet algorithm to extract the skeleton of road elements and structural key points representing key structures in road image frames. This multi-task convolutional neural network can support feature information extraction for different types of road elements through different top output layers.
[0047] In step S212, the semantic visual features in the two road image frames are determined to have a correlation in three-dimensional space.
[0048] In step S213, the spatial location information of the semantic visual features of the road image frame is determined at least based on the association relationship.
[0049] After extracting the semantic visual features of the road image frames in step S211, this embodiment of the application needs to further determine the spatial location information of the semantic visual features of the road image frames in order to construct a semantic map. To accurately obtain the spatial location information of the semantic visual features in order to construct a higher-precision semantic map, this embodiment of the application can first analyze the correlation between the semantic visual features in two road image frames in three-dimensional space. Since the visual sensor 110 is constantly acquiring road image frames during vehicle movement, the correlation between the semantic visual features in two road image frames in three-dimensional space can express the change relationship of the feature information of road elements in three-dimensional space as the vehicle moves, thus accurately reflecting the continuous change of road elements in the road image frames as the vehicle moves. Furthermore, this embodiment of the application determines the spatial location information of the semantic visual features of the road image frames based on the correlation relationship. This can be achieved by combining the change relationship of road elements as the vehicle moves when determining the spatial location information of the semantic visual features, thereby obtaining a more accurate spatial location information of the semantic visual features in three-dimensional space within the road image frames.
[0050] In some embodiments, since the semantic visual features of the road image frames extracted in step S211 are in a two-dimensional space, in an optional implementation of step S212, the semantic visual features of the road image frames can be inversely projected into a three-dimensional space; then, based on the pose information between the image acquisition time points corresponding to the two road image frames (i.e., the pose information acquired by the inertial navigation sensor 130 between the image acquisition time points of the two road image frames), the relative pose transformation of the semantic visual features of the two road image frames in the three-dimensional space is obtained; furthermore, based on this relative pose transformation, the correlation relationship of the semantic visual features of the two road image frames in the three-dimensional space is determined. In more specific embodiments, the two road image frames can be two adjacent road image frames. For example, two adjacent key road image frames, i.e., two adjacent key road image frames continuously detected by the visual sensor.
[0051] In some embodiments, based on the association determined in step S212, the embodiments of this application may use a nonlinear optimization method to solve the spatial location information of semantic visual features that have an association relationship in two road image frames, thereby obtaining the spatial location information of semantic visual features in the road image frames.
[0052] In step S214, a semantic map is obtained based at least on the geographical location corresponding to the road image frame, the semantic visual features of the road image frame, and the spatial location information of the semantic visual features.
[0053] Based on the semantic visual features extracted from the road image frame in step S211 and the spatial location information of the semantic visual features of the road image frame determined in step S213, embodiments of this application can associate the semantic visual features and spatial location information of the semantic visual features with the geographical location corresponding to the road image frame to obtain a semantic map. The obtained semantic map has both the function of vehicle geographical location matching and can provide the spatial location information of the semantic visual features of the road where the vehicle is located, providing a basis for vehicle visual positioning. In some embodiments, the geographical location corresponding to the road image frame can be considered as the geographical location at the time of acquisition of the road image frame, which can be provided by the geolocation sensor 120.
[0054] The semantic map construction method provided in this application can acquire road image frames; extract semantic visual features from the road image frames, where the semantic visual features are feature information of road elements in the road image frames; determine the correlation between semantic visual features in two road image frames in three-dimensional space; determine the spatial location information of the semantic visual features of the road image frames based at least on the correlation; and obtain a semantic map based at least on the geographical location corresponding to the road image frames, the semantic visual features, and the spatial location information of the semantic visual features. Since the correlation between semantic visual features in two road image frames in three-dimensional space can represent the change relationship of road element feature information with vehicle movement in three-dimensional space, determining the spatial location information of the semantic visual features based on the correlation allows the spatial location information of the semantic visual features to be combined with the change relationship of road elements with vehicle movement, thereby obtaining a more accurate spatial location information of the semantic visual features in three-dimensional space. Furthermore, constructing a semantic map based on the accurately obtained spatial location information of the semantic visual features can improve the positioning accuracy of the semantic map, providing a possibility for achieving accurate vehicle visual positioning.
[0055] In some further embodiments, since the data acquisition vehicle may repeatedly pass through the same road when collecting mapping data (i.e., when the multiple sensors of the data acquisition vehicle are collecting information), it is necessary to merge the road elements repeatedly detected on the same road and the related information of the road elements. Based on this, in some embodiments, after executing step S213, the same semantic visual features in multiple road image frames can be merged, and then step S214 is executed after merging the same semantic visual features. In other embodiments, after executing step S213, the same semantic visual features in multiple road image frames can be merged, and then the spatial location information of the semantic visual features of the road image frames is determined again to further improve the quality of the spatial location information of the semantic visual features of the road image frames, and then step S214 is executed. By merging the same semantic visual features in multiple road image frames, the embodiments of this application can eliminate redundant features in the semantic map, thereby reducing the amount of semantic map data and increasing the hit rate and reuse rate of feature matching between road image frames and the semantic map. In some embodiments, the present application embodiments may match the spatial location information of semantic visual features in multiple road image frames based on a cascaded greedy algorithm, and then merge the semantic visual features that are close in spatial location (e.g., the distance is less than a preset distance) and have the same semantic visual features in multiple road image frames through a disjoint-set data structure algorithm.
[0056] As a more specific optional implementation Figure 3This diagram illustrates the stages of the semantic map construction process provided in this embodiment. This embodiment can be implemented through... Figure 3 The stages shown construct a semantic map. It should be noted that whether the semantic map is constructed by the data collection vehicle itself or by a cloud server, the semantic map construction process can include... Figure 3 The stage shown.
[0057] like Figure 3 As shown, the semantic map construction process includes the following stages: keyframe determination stage 310, road element cascade detection stage 320, inter-frame road element matching stage 330, state estimation optimization mapping stage 340, duplicate road element merging stage 350, and semantic map acquisition stage 360. The content to be implemented in each stage will be explained below.
[0058] In the keyframe determination stage 310, embodiments of this application can determine key road image frames from multiple road image frames acquired by the vision sensor 110. In some embodiments, the key road image frame can be a road image frame with higher image quality acquired by the vision sensor 110, for example, a road image frame with image quality higher than a preset image quality is used as the key road image frame.
[0059] In other embodiments, the key road image frame may be a road image frame corresponding to a road of interest to the user. For example, a road image frame of a road the user prefers, or a road image frame of a tourist attraction. Based on the road image frames corresponding to roads of interest to the user, embodiments of this application construct a semantic map. This semantic map can provide spatial location information of semantic visual features in the roads of interest to the user, thereby making it easier for the user to perform assisted driving and autonomous driving on the roads of interest.
[0060] In some further embodiments, since the subsequent process of constructing a semantic map is mainly based on road elements in key road image frames, the embodiments of this application can first filter out non-road elements in key road image frames (for example, filter out backgrounds such as sky and mountains in key road image frames) before proceeding to the subsequent stages, so as to reduce the amount of data processing in the subsequent stages.
[0061] In the cascaded detection stage 320 of the road element, for each key road image frame in a single frame, the embodiments of this application can extract the semantic visual features of the key road image frame, that is, extract the feature information of road elements in the key road image frame. In some embodiments, the semantic visual features may be the feature information of standard road elements. In some embodiments, the feature information of road elements may at least include the skeleton of the road element. In further embodiments, the feature information of road elements may also include structural key points representing the key structure of the road element. The extraction process of semantic visual features is described below using the standard road elements divided into three categories: lane lines, ground markings, and road poles.
[0062] For ground signs and road poles, the feature information of road elements provided in this application embodiment may include a skeleton and structural key points representing key structures. In some embodiments, this application embodiment may determine the bounding box (e.g., a 2D bounding box) of the ground signs and road poles, and use this bounding box as the skeleton in the feature information of the ground signs and road poles. At the same time, this application embodiment may determine some structural key points in the ground signs and road poles that can represent the structure (e.g., the pole top point, the connection point of the ground sign, the polygon vertex, etc.).
[0063] For lane lines, since sampling points can represent lane lines, in this embodiment of the application, two sets of sampling points can be used to represent the left and right contours of each lane line. These two sets of sampling points can serve as the skeleton in the feature information of the lane line.
[0064] Furthermore, due to the special characteristics of virtual lane lines, this embodiment of the application also needs to determine the structural key points representing the structure within the virtual lane lines (e.g., the vertices of dashed lane lines). In some embodiments, after determining the skeleton of the lane lines, this embodiment of the application may further use a sliding window to slide along the left and right contours of the dashed lane lines to detect the vertices (i.e., the corner points) of the dashed lane lines. The specifications of this sliding window are not limited in this embodiment of the application.
[0065] It should be explained that, considering sparsity and map building efficiency, standard road elements on urban roads are suitable for detection and modeling as semantic landmarks. The main considerations here are: roadside poles and traffic signs can be captured by frontal cameras; although ground markings are sometimes obscured by vehicles, their area occupies nearly half of each road image frame, so they cannot be ignored; similarly, lane lines are also suitable for detection and modeling as semantic landmarks. These characteristics allow standard road elements to maintain effectiveness while reducing the size of the semantic map. Besides the standard road elements described above, other road elements are also worth considering, such as intersections, pedestrian bridges, and building skylines, but these are either not standardized detection methods or easily lead to ambiguity in association. Therefore, using lane lines, ground markings, and road poles as standard road elements for feature extraction is a preferred choice in this embodiment. Of course, this embodiment can also support other forms of standard road elements. Furthermore, although there are certain problems in extracting feature information from non-standard road elements, this does not affect the substantial effect of the embodiments of this application. Therefore, the embodiments of this application are not limited to extracting feature information only from standard road elements on the road. For other road elements (such as intersections, pedestrian overpasses, etc.), the embodiments of this application can also support the extraction of feature information. In some embodiments, the semantic visual features detected in the road element cascade detection stage 320 can also be set based on the user's overall habits. For example, the feature information of road elements that the user is generally interested in can be prioritized for detection.
[0066] In some embodiments, this application can utilize a multi-task convolutional neural network based on the CenterNet algorithm to extract semantic visual features of key road image frames during the cascaded detection stage of road elements 320. If the goal is to extract feature information of lane lines, ground markings, and road poles, this application can use different top output layers of the multi-task convolutional neural network to support the detection and feature extraction of these three types of standard road elements.
[0067] In an optional implementation, the multi-task convolutional neural network can first perform instance-level detection to obtain bounding boxes (which can be two-dimensional) containing structural key points for ground markings and road poles, and obtain lane contours (e.g., left and right contours of lane lines) as skeletons for lane lines. Then, for detected dashed lane lines, the multi-task convolutional neural network can use a sliding window containing candidate dashed corner points to extract the vertices of the dashed lane lines to obtain the structural key points of the dashed lane lines. In the above feature information extraction process, in order to reduce redundant computation in shareable processes such as feature information extraction, this application embodiment can use the CenterNet algorithm to separate the low-level feature extraction process in the multi-task convolutional neural network from the top-level top output layer, so that these top output layers can adapt to different types of road elements. In some embodiments, this application embodiment can use DLA (Deep Layer Aggregation) and DCN (Deformable Convolution) modules as the backbone of feature information extraction, and then after deconvolution, obtain downsampled feature maps of the top output layer adapted to different tasks (i.e., different types of road elements).
[0068] In some embodiments, the multi-task convolutional neural network described above, as a deep learning model structure, can be labeled and trained using a labeling tool platform. Compared to labeling the input data, labeling and training using a labeling tool platform can control the median pixel detection error of key road image frames to within 2 pixels.
[0069] In the inter-frame road element matching stage 330, this embodiment of the application can determine the semantic visual features of two key road image frames in three-dimensional space. It is understood that as the vehicle moves, the road elements in the key road image frames collected by the visual sensor 110 are dynamically changing, and the changes in road elements between key road image frames are continuous. For example, the vertex of a lane line continuously approaches the vehicle as it moves and then disappears from the vehicle's field of vision. Therefore, after extracting the feature information (i.e., semantic visual features) of road elements in the key road image frames in the road element cascade detection stage 320, this embodiment of the application can further determine the correlation of the feature information of road elements in two key road image frames in three-dimensional space in the inter-frame road element matching stage 330, so as to reflect the changing relationship of the feature information of road elements as the vehicle moves.
[0070] In one alternative implementation, Figure 4 A flowchart illustrating an optional implementation of inter-frame road element matching provided in an embodiment of this application is shown. For example... Figure 4 As shown, the process may include the following steps.
[0071] In step S410, the first semantic visual feature corresponding to the first road image frame and the second semantic visual feature corresponding to the second road image frame are obtained.
[0072] The first road image frame and the second road image frame can be two road image frames. For example, two key road image frames determined in the keyframe determination stage 210. In some embodiments, the first road image frame and the second road image frame can be two adjacent road image frames, such as two adjacent key road image frames. For ease of explanation, the semantic visual features corresponding to the first road image frame in this application embodiment can be referred to as the first semantic visual features, and the semantic visual features corresponding to the second road image frame can be referred to as the second semantic visual features.
[0073] In step S411, the first semantic visual feature and the second semantic visual feature are inversely projected into the three-dimensional space, respectively.
[0074] In some embodiments, this application embodiment can back-project the first semantic visual features corresponding to the first road image frame and the second semantic visual features corresponding to the second road image frame into a three-dimensional space based on the ground plane parameters and the relative pose of the visual sensor 110, thereby obtaining the back-projection information of the first and second semantic visual features in the three-dimensional space. It is understood that the first and second semantic visual features extracted in the road element cascade detection stage 320 are in a two-dimensional space state. Step S411, by back-projecting the first and second semantic visual features in the two-dimensional space state into a three-dimensional space, can provide a basis for subsequently determining the correlation between the first and second semantic visual features in the three-dimensional space.
[0075] In step S412, based on the pose information between the image acquisition time points of the first road image frame and the second road image frame, the relative pose transformation of the first semantic visual feature and the second semantic visual feature in three-dimensional space is obtained.
[0076] In some embodiments, this application can acquire pose information collected by the inertial navigation sensor 130 between the image acquisition time points of the first road image frame and the second road image frame, and then determine the relative pose transformation of the first semantic visual feature and the second semantic visual feature in three-dimensional space based on the pose information. It is understood that during vehicle operation, the inertial navigation sensor 130 is continuously acquiring pose information. After inversely projecting the first semantic visual feature and the second semantic visual feature into three-dimensional space, to obtain the pose relationship between the first semantic visual feature and the second semantic visual feature in three-dimensional space, this application can utilize the pose information between the image acquisition time points of the first road image frame and the second road image frame. That is, the relative pose transformation of the first semantic visual feature and the second semantic visual feature in three-dimensional space can be determined using the pose information between the image acquisition time points of the first road image frame and the second road image frame.
[0077] In step S413, based on the relative pose transformation, the correlation between the first semantic visual feature and the second semantic visual feature in three-dimensional space is determined.
[0078] The relative pose transformation represents the transformation relationship of the feature information of road elements in the first road image frame and the second road image frame in terms of relative pose as the vehicle moves. Therefore, based on the relative pose transformation, this application embodiment can determine the association relationship in three-dimensional space between the first semantic visual feature corresponding to the first road image frame and the second semantic visual feature corresponding to the second road image frame. In some embodiments, based on the relative pose transformation, this application embodiment can use a greedy matching algorithm to associate the first semantic visual feature and the second semantic visual feature in three-dimensional space, thereby obtaining the association relationship in three-dimensional space between the first semantic visual feature and the second semantic visual feature. The greedy matching algorithm (also known as the greedy algorithm) refers to always making the best choice at the moment when solving a problem, that is, without considering the overall best solution, the algorithm obtains a locally better solution in some sense.
[0079] It should be noted that, for any two key road image frames, the embodiments of this application can be achieved through... Figure 3 The process shown determines the correlation between the feature information of road elements in two key road image frames in three-dimensional space.
[0080] In some more specific implementations of the inter-frame road element matching stage 330, given two consecutively detected key road image frames, this embodiment can use an inertial navigation sensor 130 (e.g., an IMU) to accumulate the relative pose transformation between the two key road image frames. For ground road elements (e.g., ground markings and lane lines), this embodiment can first perform ray-ground intersection in the coordinate system of the visual sensor (e.g., the camera coordinate system) to obtain a rough three-dimensional position for each semantic key point, box vertex, and lane line sample point of the ground road element, so as to realize the inverse projection of the semantic visual features in the two key road image frames into three-dimensional space respectively. Then, we use a greedy matching algorithm to associate the semantic visual features back-projected into 3D space in pixel space: the semantic visual features of one key road image frame in 3D space are reprojected onto another key road image frame, and then the intersection of the semantic visual features in the two key road image frames on the union is calculated. If the semantic visual features in the two key road image frames intersect on the union, this embodiment can establish instance matching, and further considers intra-instance matching that reprojects the instance matching. In instance matching and intra-instance matching, feature information with a pixel percentage less than 50% or a pixel distance greater than 5.0 pixels in the union will be ignored.
[0081] It should be noted that for structural keypoints such as vertices in semantic visual features, this embodiment can use optical flow methods to track feature information between key road image frames. During feature information tracking, this embodiment retains classic keypoints extracted, described, and tracked by the GFTT extractor and anomaly descriptor, because these are not only part of visual inertial ranging but also stable feature points worth including in structured objects for tracking. Unlike the segmentation of the output mask, the bounding boxes of road elements detected in this embodiment may contain GFTT feature keypoints from the background region, especially in vertex instances.
[0082] In the state estimation optimization mapping stage 340, this embodiment of the application can solve the spatial location information of the semantic visual features that are related in two key road image frames to obtain the spatial location information of the semantic visual features of the key road image frames.
[0083] In some embodiments, the spatial location information of the semantic visual features may include at least one of the following: the three-dimensional spatial location of the structural key points of the ground road element, the three-dimensional spatial location of the sampling points of the lane line, the coefficient representing the three-dimensional planar spatial location of the ground road element, the position coefficient of the ground plane in the coordinate system of the visual sensor, the correlation coefficient of the lane line, and the pose of the road image frame (e.g., the pose of the key road image frame). The above types of spatial location information will be described below.
[0084] In some embodiments, for lane markings on a road, four consecutive sampling points determine the shape between two intermediate sampling points. For example, suppose the four consecutive sampling points are C. k-1 C k C k+1 and C k+2 Then the two intermediate sampling points C k and C k+1 The shape between them can be represented by C(t'), and C(t') can be expressed by the following formula:
[0085]
[0086] Where t'∈[0,1]; τ=0.5, describes the shape of the lane line curvature at the sampling point. On both sides of the lane line, the first and last sampling points are always offset from the lane line to adjust the direction of their endpoints. It can be seen that in the state estimation optimization mapping stage 340 of this application, the lane line of the conventional road surface can be fitted with a piecewise cubic Catmull-Rom spline curve to determine the shape of the lane line curve.
[0087] In some embodiments of this application, five optimizable variables may be introduced in the state estimation optimization mapping stage 340.
[0088] 1) This embodiment of the application can detect the three-dimensional spatial positions of structural key points representing structures from ground road elements and corner points of ground road elements in key road image frames. This embodiment of the application can use an inverse depth parameterization scheme to optimize the three-dimensional spatial positions of these structural key points.
[0089] 2) Ground coefficients in the coordinate system of the visual sensor, i.e., position coefficients in the ground plane space in the coordinate system of the visual sensor. In this embodiment, the observable area of the ground in each frame of a key road image can be approximated as a plane, and expressed using the following formula:
[0090]
[0091] Based on the above description, in the state estimation optimization mapping stage 340, the ground coefficients in the coordinate system of the visual sensor can be optimized to support online calibration of ground information. Where θ∈[0, π], d∈R3, where R3 represents the set of sampling points.
[0092] 3) The coefficient representing the three-dimensional planar spatial position of ground road elements, i.e., the vertical plane in global coordinates, can be expressed using the following formula:
[0093]
[0094] The above formula can be considered as the Z-axis of object α being horizontal with the direction of gravity after initializing the visual inertial measurement range. α∈[0,2π], e∈R3.
[0095] 4) The three-dimensional spatial location of the sampling points of the lane lines, which depicts the left or right contour of the road lane lines.
[0096] 5) Lane line correlation coefficient, also known as dynamic correlation parameter, is used to correlate sampling points and corner points with lane lines detected in key road image frames.
[0097] In some further embodiments, the variables that can be optimized in this application embodiment may also include: the pose of key road image frames.
[0098] As an optional implementation, this embodiment of the application can provide three types of constraints for the detected semantic visual features in the state estimation optimization mapping stage 340, including:
[0099] 1) Reprojection constraints of structural key points in road elements, also known as point observation factors. In this application, the structural key points in road elements can be triangulated and parameterized based on the following formula:
[0100]
[0101] Where, p aci It represents the pixel position corresponding to the detected structural keypoint, π(·) is the camera projection operator, and T c It is the pose of the key road image frame c. It is the allocated noise covariance, representing the accuracy of the detected sampling points and corner points.
[0102] 2) Lane line reprojection constraint, also known as lane line observation factor. In this embodiment, the constraint can be based on the following formula to dynamically associate sampling points and structural key points p through an explicit association parameter. aci , to be used as the measured value of lane line Ca:
[0103]
[0104] Among them, t aci These are the dynamically related parameters introduced for joint optimization. Similar to This indicates the precision of the detected sampling points and corner points.
[0105] 3) Coplanar constraints of structural key points in road elements, also known as coplanar prior factors. In some embodiments, based on the thickness or noise covariance of the ground road elements, this application embodiment may assume that these observed lane lines are locally planar in each camera view and have consistent coefficients.
[0106] Based on the optimizable variables and constraint types described above, in the state estimation optimization mapping stage 340, embodiments of this application can initialize the optimizable variables based on their order. In some embodiments, given a key road image frame, embodiments of this application can triangulate the structural key points of road elements in the key road image frame from the estimated pose of the key road image frame. Then, for vertical ground road elements, embodiments of this application use feature points of their depth to perform line fitting on the XOY plane to obtain V. a (α, e) represents the coefficients characterizing the three-dimensional planar spatial position of ground road elements; simultaneously, the thickness standard |V is used. a (α, e)·P ai |<σ1=0.3m, accept the pixel positions extracted from those successfully triangulated structural keypoints into the detected bounding box. Embodiments of this application include these successfully triangulated structural keypoints within the bounding box because these structural keypoints are typically detected stably and can provide useful geometric constraints.
[0107] In some embodiments, after obtaining the triangulated structural key points, the point set of these structural key points can be matched with a three-dimensional plane to obtain the coefficients in the camera coordinates, i.e., the ground coefficients G(θ, φ, d) in the coordinate system of the visual sensor. If no ground markings are found during the initialization process, the embodiments of this application can use the feature points corresponding to the lane convex hulls detected on each key road image frame, and select to use a three-dimensional plane fitting strategy to remove the feature points corresponding to the lane convex hulls on the moving vehicle.
[0108] For lane line initialization, this embodiment of the application can use a random sampling detection consistency algorithm to initialize the lane line sampling points on a conventional road surface. In this process, this embodiment of the application can consider the three-dimensional residual of the curve fitting of the lane line sampling points, and then use a regularization term to ensure uniform sampling of these sampling points.
[0109] In some embodiments of this application, during the state estimation optimization mapping stage 340, the lane lines of the conventional road surface are subjected to reprojection constraints to explicitly use the correlation coefficients of the lane lines for construction and optimization. The initialization of such lane line correlation coefficients is first completed by solving for stationary points using the derivative of the spatial three-dimensional distance function, and then by independent nonlinear optimization.
[0110] In the repeated road element merging stage 350, this embodiment of the application can merge semantic visual features repeatedly detected on the same road. It is understood that when the data acquisition vehicle is collecting mapping data, it may occasionally repeatedly pass through the same road segment, so it is necessary to merge the repeatedly detected semantic visual features on the same road. This embodiment of the application can use a cascaded greedy matching method to match the spatial location information of semantic visual features in multiple frames of key road image frames, and use a disjoint-set data structure algorithm to merge semantic visual features that are close in distance and detected as the same. Then, within each semantic visual feature, substructure matching is performed to ensure that the feature points in each semantic visual feature form correct associations. Finally, the state estimation optimization mapping stage 340 is re-executed for a second round of offline nonlinear optimization to improve the quality of the spatial location information of semantic visual features in the semantic map.
[0111] In the semantic map acquisition stage 360, this embodiment of the application can combine the semantic visual features of key road image frames, the spatial location information of the semantic visual features, and the geographical location of the key road image frames to obtain a semantic map. For example, after merging repeated semantic visual features, this embodiment of the application can store the state variables such as the pose of the optimized key road image frames, the spatial location information of the semantic features, and the geographical location of the key road image frames to form a semantic map for use during online vehicle visual positioning.
[0112] It should be noted that the semantic map construction process provided in this application embodiment is not limited to being based on key road image frames, but can be based on any road image frames collected by a visual sensor.
[0113] The semantic map construction scheme provided in this application embodiment can more accurately determine the spatial location information of semantic visual features in three-dimensional space by combining the changing relationship of road elements with vehicle movement; furthermore, it can eliminate redundant semantic visual features in the semantic map, increase the hit rate and reuse rate of feature matching between the image and the semantic map, and reduce the amount of semantic map storage data. Therefore, the semantic map construction scheme provided in this application embodiment can improve the positioning accuracy of the semantic map while reducing the amount of semantic map data. The semantic map construction scheme provided in this application embodiment requires sensors such as monocular cameras, Global Positioning System (GPS) receivers, and Inertial Navigation Units (IMUs), and can be deployed and used on conventional embedded platforms, such as NVIDIA TX2.
[0114] Based on the semantic map constructed above, embodiments of this application can achieve precise visual positioning of the vehicle during driving, thereby providing a decision-making basis for assisted driving and autonomous driving. As an optional implementation, Figure 5An optional flow of the visual positioning method provided in the embodiments of this application is shown. Figure 5 The method shown can be performed by an in-vehicle device. This in-vehicle device can be a user device placed in the vehicle (e.g., a dashcam, a user's mobile phone, etc.) or a device built into the vehicle. The in-vehicle device or the vehicle can be configured with at least the following settings: Figure 1 The visual and geolocation sensors shown are used to acquire road images and locate the vehicle's geographical position while the vehicle is in motion. In a further optional implementation, the onboard equipment or vehicle may also be equipped with inertial navigation sensors to acquire pose information while the vehicle is in motion.
[0115] like Figure 5 As shown, the process may include the following steps.
[0116] In step S510, the current road image frame and the current geographical location of the vehicle are acquired.
[0117] During vehicle operation, a visual sensor can acquire road images, and a geolocation sensor can determine the vehicle's geographical location. The road images acquired by the visual sensor can include multiple road image frames. When performing visual localization of the vehicle at the current moment, the road image frame acquired by the visual sensor at that moment can be called the current road image frame, and the geographical location of the vehicle determined by the geolocation sensor at that moment can be called the current geographical location.
[0118] In some embodiments, the current road image frame may be a key road image frame. However, the current road image frame is not limited to key road image frames. In other embodiments, the current road image frame may be the current road image frame corresponding to a road of interest to the user. For example, the current road image frame of a road the user prefers, or the current road image frame of a tourist attraction, etc.
[0119] In some further embodiments, the embodiments of this application may filter background information in the current road image frame before proceeding to subsequent steps. In other possible implementations, the information to be filtered in the current road image frame may also be set by the user, for example, by providing a settings page that allows the user to set image information they do not care about, so that the embodiments of this application can filter out image information that the user does not care about in the current road image frame before proceeding to subsequent steps.
[0120] In step S511, matching current map data is obtained from the semantic map based on the current geographical location.
[0121] In some embodiments, after constructing a semantic map based on the semantic map construction method described above, the semantic map can be stored in the cloud and divided into tiles according to geographical location ranges. For example, a semantic map may include multiple map tiles, with each map tile corresponding to a geographical location range. A map tile may include semantic visual features of roads within the corresponding geographical location range, spatial location information of the semantic visual features, and other data.
[0122] After obtaining the vehicle's current geographical location, this application embodiment can request a current map tile (which can be considered an optional form of current map data) from the cloud that matches the current geographical location. In some embodiments, after the cloud receives the request from the vehicle-mounted device, it can determine the geographical location range that matches the current geographical location, and then use the map tile corresponding to that geographical location range as the current map tile and feed it back to the vehicle-mounted device.
[0123] In step S512, the current semantic visual features are extracted from the current road image frame.
[0124] For the current road image frame obtained in step S510, this embodiment of the application can extract semantic visual features from the current road image frame. For ease of explanation, the semantic visual features in the current road image frame can be referred to as current semantic visual features. The specific implementation method for extracting semantic visual features from the road image frame can be referred to in the corresponding section above, and will not be repeated here.
[0125] In step S513, at least spatial location information matching the current semantic visual feature is obtained from the current map data to obtain the initial spatial location information of the current semantic visual feature.
[0126] In step S514, the current spatial location information of the current semantic visual feature is determined based on the initial spatial location information.
[0127] The current map data can record semantic visual features of roads within a geographic area matching the current geographic location, as well as spatial location information of these semantic visual features. After extracting the current semantic visual features from the current road image frame in step S512, this embodiment can match the spatial location information corresponding to the current semantic visual features from the current map data to achieve retrieval of the current map data, thereby obtaining the initial spatial location information of the current semantic visual features. Since the actual driving position, angle, and shape of the vehicle may differ from the data acquisition vehicle used for mapping, the spatial location information of the current semantic visual features matched from the current map data can be used as the initial spatial location information in this embodiment. This embodiment can further optimize this initial spatial location information to obtain accurate current spatial location information. In some embodiments, if the spatial location information of the semantic visual feature includes the following information: the three-dimensional spatial location of the structural key points of the ground road element, the three-dimensional spatial location of the sampling points of the lane line, the coefficient representing the three-dimensional planar spatial location of the ground road element, the position coefficient of the ground plane space in the coordinate system of the visual sensor, the correlation coefficient of the lane line, and the pose of the key road image frame; then the three-dimensional spatial location of the structural key points of the ground road element and the three-dimensional spatial location of the sampling points of the lane line can be set as constants in the semantic map. Thus, in this embodiment, the three-dimensional spatial location of the structural key points of the ground road element and the three-dimensional spatial location of the sampling points of the lane line corresponding to the current semantic visual feature as constants can be matched from the current map data, and then modeled and solved to obtain the spatial location information of other items of the current semantic visual feature, thereby obtaining the current spatial location information of the current semantic visual feature.
[0128] In one alternative implementation, Figure 6 A schematic diagram illustrating the stages of the visual positioning process provided in an embodiment of this application is shown. Figure 6 As shown, the stages of the visual localization process may include: current keyframe determination stage 610, road element cascade detection stage 620, map tile loading stage 630, map retrieval stage 640, and state estimation online localization stage 650.
[0129] In the current keyframe determination stage 610, this embodiment of the application can acquire the current key road image frame at the current moment. Furthermore, this embodiment of the application can also acquire the current geographical location of the vehicle.
[0130] In the road element cascade detection stage 620, the embodiments of this application can extract the current semantic visual features from the current key road image frame, that is, the feature information of road elements in the current key road image frame.
[0131] During the map tile loading stage 630, this embodiment of the application can obtain the current map tile that matches the current geographical location from the cloud based on the vehicle's current geographical location. For example, the vehicle-mounted device can use GPS data to query the map tile stored in the cloud that best matches the latitude and longitude spatial location, and pull it to the local device of the vehicle for matching the semantic visual features of the current key road image frame with the map tile.
[0132] In the map retrieval stage 640, embodiments of this application can retrieve the current map tile based on the current semantic visual features and match the initial spatial location information of the current semantic visual features. For example, the three-dimensional spatial locations of the structural key points of the ground road elements and the three-dimensional spatial locations of the sampling points of the lane lines, which are constants, can be matched from the current map tile. In some embodiments, for in-vehicle devices in operation, embodiments of this application can utilize the computing power of the embedded platform to perform a low-frequency depth perception, and retrieve the current map tile based on the semantic visual features obtained from these sparse detections.
[0133] In the online localization stage 650 of state estimation, this embodiment of the application can solve for the current spatial location information of the current semantic visual feature based on the initial spatial location information obtained in the map retrieval stage 640. For example, by taking the three-dimensional spatial location of the structural key points of the ground road element corresponding to the current semantic visual feature and the three-dimensional spatial location of the sampling points of the lane line as constants, at least one of the following spatial location information can be solved: the coefficient representing the three-dimensional planar spatial location of the ground road element corresponding to the current semantic visual feature, the position coefficient of the ground plane space in the coordinate system of the visual sensor, the correlation coefficient of the lane line, and the pose of the road image frame.
[0134] In some embodiments, this application can support visual positioning based on user-customized configuration of conditions such as road and time. For example, the in-vehicle device can provide user configuration options, allowing users to configure road, time, and other conditions for visual positioning. Different users may have different configurations, thereby achieving visual positioning based on user-customized configurations.
[0135] In some embodiments, the visual positioning scheme provided in this application can be installed as a service on in-vehicle equipment to enable users to better utilize the vehicle's assisted driving and autonomous driving functions. This service can be enabled or disabled by the user. When the service is enabled, this application can provide corresponding charging information; during the service process, this application can also support the insertion of recommended information (such as advertisements).
[0136] The visual positioning and semantic map construction schemes provided in this application can be implemented using a monocular camera, without relying on left-right eye matching of a wide-baseline binocular camera or distance perception information from a depth camera. This allows these embodiments to be directly used in small in-vehicle devices such as dashcams. The visual positioning and semantic map construction schemes provided in this application improve mapping and positioning accuracy, achieving higher-precision visual positioning results with less map storage capacity. These embodiments form a closed-loop technology encompassing semantic visual feature perception, semantic map construction, and visual positioning. Semantic maps can be constructed and used using the results of detection-based deep learning, and uncertainty estimation can be achieved during semantic map construction using constraint types, conforming to the design principles of maximum likelihood estimation and improving the quality of semantic map construction and positioning.
[0137] The foregoing describes multiple embodiment schemes provided by the embodiments of this application. The optional methods described in each embodiment scheme can be combined and cross-referenced with each other without conflict, thereby extending to a variety of possible embodiment schemes. These can all be considered as the embodiment schemes disclosed and published by the embodiments of this application.
[0138] The semantic map construction apparatus provided in the embodiments of this application is described below. The apparatus described below can be considered as a data collection vehicle or a cloud server, representing the functional modules required to implement the semantic map construction method provided in the embodiments of this application. The apparatus described below can be referenced in conjunction with the description above.
[0139] In the optional implementation, Figure 7 An optional block diagram of the semantic map construction apparatus provided in an embodiment of this application is shown. For example... Figure 7 As shown, the device may include:
[0140] Image frame acquisition module 710 is used to acquire road image frames;
[0141] The feature extraction module 711 is used to extract the semantic visual features of the road image frame, wherein the semantic visual features are the feature information of road elements in the road image frame;
[0142] The association determination module 712 is used to determine the association relationship between semantic visual features in two road image frames in three-dimensional space.
[0143] The spatial location determination module 713 is used to determine the spatial location information of the semantic visual features of the road image frame based at least on the association relationship;
[0144] The map acquisition module 714 is used to obtain a semantic map based at least on the geographical location, semantic visual features, and spatial location information of the semantic visual features corresponding to the road image frame.
[0145] In some embodiments, the two road image frames include: a first road image frame and a second road image frame. The association determination module 712, used to determine the association relationship of semantic visual features in the two road image frames in three-dimensional space, includes:
[0146] The first semantic visual features corresponding to the first road image frame and the second semantic visual features corresponding to the second road image frame are respectively inversely projected into three-dimensional space;
[0147] Based on the pose information between the image acquisition time points of the first road image frame and the second road image frame, the relative pose transformation of the first semantic visual feature and the second semantic visual feature in three-dimensional space is obtained.
[0148] Based on the relative pose transformation, the correlation between the first semantic visual feature and the second semantic visual feature in three-dimensional space is determined.
[0149] In some embodiments, the spatial location determination module 713 is used to determine the spatial location information of the semantic visual features of the road image frame, at least based on the association relationship, including:
[0150] By using a nonlinear optimization method, the spatial location information of semantic visual features that are related in two road image frames is solved simultaneously to obtain the spatial location information of semantic visual features in the road image frames.
[0151] In some embodiments, the spatial location information of the semantic visual features includes at least one of the following:
[0152] The three-dimensional spatial positions of the structural key points of ground road elements, the three-dimensional spatial positions of the sampling points of lane lines, the coefficients representing the three-dimensional planar spatial positions of ground road elements, the position coefficients of the ground plane in the coordinate system of the visual sensor, the correlation coefficients of lane lines, and the pose of road image frames.
[0153] In some embodiments, the feature extraction module 711 is used to extract semantic visual features of the road image frame, including:
[0154] The multi-task convolutional neural network, based on the central network algorithm, extracts the skeleton of road elements and key structural points representing key structures in the road image frame; the multi-task convolutional neural network supports feature extraction of different types of road elements through different top output layers.
[0155] In some further embodiments, before obtaining the semantic map based at least on the geographical location, semantic visual features, and spatial location information of the semantic visual features corresponding to the road image frame, the semantic map construction apparatus may also be used for:
[0156] Merge identical semantic visual features in multiple road image frames.
[0157] This application also provides a semantic map construction device, which can be a computing and processing device in a data collection vehicle or a cloud server. This semantic map construction device can execute the semantic map construction method provided in this application by loading the semantic map construction apparatus described above. As an optional implementation, Figure 8 A block diagram of a semantic map building device provided in an embodiment of this application is shown. Figure 8 As shown, the semantic map building device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4.
[0158] In this embodiment, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4.
[0159] Optionally, communication interface 2 can be an interface for a communication module used for network communication.
[0160] Optionally, processor 1 may be a CPU (Central Processing Unit), GPU (Graphics Processing Unit), NPU (Embedded Neural Network Processor), FPGA (Field Programmable Gate Array), TPU (Tensor Processing Unit), AI chip, ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of this application.
[0161] Memory 3 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0162] The memory 3 stores one or more computer-executable instructions, and the processor 1 calls the one or more computer-executable instructions to execute the semantic map construction method provided in this application embodiment.
[0163] This application also provides a storage medium that can store one or more computer-executable instructions. When the one or more computer-executable instructions are executed, the semantic map construction method provided in this application can be implemented.
[0164] The visual positioning device provided in the embodiments of this application will be described below. The device described below can be considered as an in-vehicle device, which includes the functional modules required to implement the visual positioning method provided in the embodiments of this application. The device described below can be referred to in correspondence with the description above.
[0165] Figure 9 A block diagram of a visual positioning device provided in an embodiment of this application is shown. Figure 9 As shown, the device may include:
[0166] The information acquisition module 910 is used to acquire the current road image frame and the current geographical location of the vehicle;
[0167] The map data acquisition module 911 is used to acquire matching current map data from the semantic map based on the current geographical location;
[0168] The semantic feature extraction module 912 is used to extract current semantic visual features from the current road image frame;
[0169] The map retrieval and matching module 913 is used to obtain at least the spatial location information that matches the current semantic visual feature from the current map data, so as to obtain the initial spatial location information of the current semantic visual feature;
[0170] The current spatial location determination module 914 is used to determine the current spatial location information of the current semantic visual feature based on the initial spatial location information.
[0171] In some embodiments, the map data acquisition module 911 is configured to acquire matching current map data from the semantic map based on the current geographical location, including:
[0172] Request the current map tile that matches the current geographical location from the cloud; wherein, the semantic map is divided into multiple map tiles according to the geographical location range, and each map tile has a corresponding geographical location range;
[0173] Get the current map tile corresponding to the geographical location range that matches the current geographical location, as fed back from the cloud.
[0174] In some embodiments, the map retrieval and matching module 913 is configured to obtain at least spatial location information matching the current semantic visual feature from the current map data, so as to obtain the initial spatial location information of the current semantic visual feature, including:
[0175] The three-dimensional spatial positions of the structural key points of the ground road elements corresponding to the current semantic visual features and the three-dimensional spatial positions of the sampling points of the lane lines are obtained from the current map data.
[0176] In some embodiments, the current spatial location determination module 914 is configured to determine the current spatial location information of the current semantic visual feature based on the initial spatial location information, including:
[0177] Using the three-dimensional spatial positions of the structural key points of the ground road element corresponding to the current semantic visual feature and the three-dimensional spatial positions of the sampling points of the lane lines as constants, at least one of the following spatial position information is obtained: the coefficient representing the three-dimensional planar spatial position of the ground road element corresponding to the current semantic visual feature, the position coefficient of the ground plane space in the coordinate system of the visual sensor, the correlation coefficient of the lane lines, and the pose of the road image frame.
[0178] This application also provides a vehicle-mounted device that can implement the visual positioning method provided in this application by mounting the visual positioning device described above. The optional hardware framework of this vehicle-mounted device can be combined with... Figure 8 As shown, the device includes at least one memory and at least one processor. The memory stores one or more computer-executable instructions, and the processor invokes the one or more computer-executable instructions to execute the visual positioning method provided in the embodiments of this application. In some further embodiments, the vehicle-mounted device may also include sensors such as a visual sensor, a geolocation sensor, and an inertial navigation sensor.
[0179] This application also provides a storage medium that can store one or more computer-executable instructions. When the one or more computer-executable instructions are executed, the visual positioning method provided in this application can be implemented.
[0180] While the embodiments disclosed above are described in this application, this application is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of this application; therefore, the scope of protection of this application should be determined by the scope defined in the claims.
Claims
1. A semantic map construction method, wherein, include: Acquire road image frames; Extract the semantic visual features of the road image frame, where the semantic visual features are the feature information of road elements in the road image frame; Determine the semantic visual features in two road image frames and their relationship in three-dimensional space; At least the spatial location information of the semantic visual features of the road image frame is determined based on the aforementioned association. A semantic map is obtained based at least on the geographical location, semantic visual features, and spatial location information of the semantic visual features corresponding to the road image frame; The extraction of semantic visual features from the road image frame includes: The multi-task convolutional neural network, based on the central network algorithm, extracts the skeleton of road elements and key structural points representing key structures in the road image frame; the multi-task convolutional neural network supports feature extraction of different types of road elements through different top output layers.
2. The semantic map construction method according to claim 1, wherein, The two road image frames include: a first road image frame and a second road image frame; determining the semantic visual features in the two road image frames in three-dimensional space includes: The first semantic visual features corresponding to the first road image frame and the second semantic visual features corresponding to the second road image frame are respectively inversely projected into three-dimensional space; Based on the pose information between the image acquisition time points of the first road image frame and the second road image frame, the relative pose transformation of the first semantic visual feature and the second semantic visual feature in three-dimensional space is obtained. Based on the relative pose transformation, the correlation between the first semantic visual feature and the second semantic visual feature in three-dimensional space is determined.
3. The semantic map construction method according to claim 1, wherein, The step of determining the spatial location information of the semantic visual features of the road image frame based at least on the association relationship includes: By using a nonlinear optimization method, the spatial location information of semantic visual features that are related in two road image frames is solved simultaneously to obtain the spatial location information of semantic visual features in the road image frames.
4. The semantic map construction method according to claim 1 or 3, wherein, The spatial location information of the semantic visual features includes at least one of the following: The three-dimensional spatial positions of the structural key points of ground road elements, the three-dimensional spatial positions of the sampling points of lane lines, the coefficients representing the three-dimensional planar spatial positions of ground road elements, the position coefficients of the ground plane in the coordinate system of the visual sensor, the correlation coefficients of lane lines, and the pose of road image frames.
5. The semantic map construction method according to claim 1, wherein, Before obtaining a semantic map based at least on the geographical location, semantic visual features, and spatial location information of the semantic visual features corresponding to the road image frame, the method further includes: Merge identical semantic visual features in multiple road image frames.
6. A visual positioning method, wherein, include: Obtain the current road image frame and the current geographical location of the vehicle; Based on the current geographical location, obtain matching current map data from the semantic map; And, extract current semantic visual features from the current road image frame; wherein, the semantic map is constructed using the semantic map construction method as described in claim 1; At least spatial location information matching the current semantic visual feature is obtained from the current map data to obtain the initial spatial location information of the current semantic visual feature; Based on the initial spatial location information, the current spatial location information of the current semantic visual feature is determined.
7. The visual positioning method according to claim 6, wherein, The step of obtaining matching current map data from the semantic map based on the current geographical location includes: Request the current map tile that matches the current geographical location from the cloud; wherein, the semantic map is divided into multiple map tiles according to the geographical location range, and each map tile has a corresponding geographical location range; Obtain the current map tile corresponding to the geographical location range that matches the current geographical location, as fed back from the cloud.
8. The visual positioning method according to claim 6, wherein, The step of obtaining at least spatial location information matching the current semantic visual feature from the current map data to obtain the initial spatial location information of the current semantic visual feature includes: Obtain the three-dimensional spatial positions of the structural key points of the ground road elements corresponding to the current semantic visual features from the current map data, and the three-dimensional spatial positions of the sampling points of the lane lines. The step of determining the current spatial location information of the current semantic visual feature based on the initial spatial location information includes: Using the three-dimensional spatial positions of the structural key points of the ground road element corresponding to the current semantic visual feature and the three-dimensional spatial positions of the sampling points of the lane lines as constants, at least one of the following spatial position information is obtained: the coefficient representing the three-dimensional planar spatial position of the ground road element corresponding to the current semantic visual feature, the position coefficient of the ground plane space in the coordinate system of the visual sensor, the correlation coefficient of the lane lines, and the pose of the road image frame.
9. A semantic map building device, wherein, include: At least one memory and at least one processor, the memory storing one or more computer-executable instructions, the processor invoking the one or more computer-executable instructions to perform the semantic map construction method as described in any one of claims 1-5.
10. A vehicle-mounted device, wherein, include: At least one memory and at least one processor, the memory storing one or more computer-executable instructions, the processor invoking the one or more computer-executable instructions to perform the visual positioning method as described in any one of claims 6-8.
11. A storage medium, wherein, The storage medium stores one or more computer-executable instructions, which, when executed, implement the semantic map construction method as described in any one of claims 1-5, or the visual positioning method as described in any one of claims 6-8.
Citation Information
Patent Citations
Position determination method and device based on data association
CN112446234A