A vehicle repositioning method and system based on a lightweight visual semantic map and an electronic device

CN122510356BActive Publication Date: 2026-09-25RATE TECH (CHONGQING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610983108.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-25
Estimated Expiration
2046-07-03

AI Technical Summary

Technical Problem

[0003]当前主流视觉重定位技术存在两类关键问题:一类依赖激光雷达(LiDAR)生成的稠密点云与高精度地图,设备成本高、数据存储量大,难以在量产车型普及;另一类基于传统视觉特征,易受光照、天气、季节变化影响,鲁棒性不足

Benefits of technology

[0019]经由上述的技术方案可知,与现有技术相比,本发明的有益效果包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122510356B_ABST
    Figure CN122510356B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle repositioning method and system based on a lightweight visual semantic map, which utilizes sparse semantic point clouds obtained by a vehicle-mounted camera and a search area delimited by coarse-precision GNSS to quickly screen candidate positions through semantic layered convolution; then converts the point clouds into a format with semantic label coding, inputs the format into a DH3D network to extract global and local descriptors of fused geometric and semantic information; finally, calculates the global descriptor cosine similarity between query point clouds and each candidate position, introduces a spatial distance constraint in local descriptor matching, combines the global descriptor similarity and the similarity calculation of the local descriptor with the constraint to obtain a comprehensive score, selects the candidate position with the highest score as the current accurate position of the vehicle, and realizes vehicle repositioning. The application does not require a laser radar or a high-precision map, relies on low-cost sensors, can realize millisecond-level high-robustness repositioning in a GNSS failure scene, and is suitable for mass-produced automatic driving vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence, smart transportation and intelligent positioning, and in particular to a vehicle relocation method, system and electronic device based on a lightweight visual semantic map, which can be widely used in autonomous vehicles and advanced driver assistance systems, and is especially suitable for complex urban scenarios where high-precision GNSS signals are unavailable or unreliable. Background Technology

[0002] Vehicle relocation is a core component of the autonomous driving positioning system, primarily addressing two scenarios: first, vehicle startup initialization, where a precise position of centimeters to decimeters can be quickly acquired given a local map covering a range of hundreds of meters; and second, positioning recovery during driving, where a precise position is regained through relocation after the vehicle enters an area without GNSS signals, providing an initial reference for subsequent inertial navigation and visual navigation.

[0003] Current mainstream visual relocalization technologies face two key challenges: one relies on dense point clouds and high-precision maps generated by LiDAR, resulting in high equipment costs and large data storage requirements, hindering widespread adoption in mass-produced vehicles; the other is based on traditional visual features, which are susceptible to changes in lighting, weather, and seasons, exhibiting insufficient robustness. Lightweight visual semantic maps, which have emerged in recent years, generate sparse semantic point clouds (containing semantic labels such as lane lines, stop lines, and arrows) using cameras, offering advantages in low cost and small storage requirements. However, when existing LiDAR-based dense point cloud-based localization algorithms are directly transferred to sparse semantic point clouds, insufficient geometric information leads to a sharp drop in localization accuracy and low success rate, failing to meet practical application needs. While existing technologies such as AVP-SLAM and sparse semantic SLAM attempt to incorporate semantic features, they still suffer from low candidate region selection efficiency, insufficient fusion of semantic information in descriptors, and a lack of consideration for spatial distance constraints during matching, resulting in slow relocalization speeds, poor accuracy, and difficulty in adapting to complex urban road scenarios.

[0004] Existing technologies suffer from high equipment and cost barriers. Mainstream high-precision positioning solutions rely on LiDAR and paid high-precision GNSS services, resulting in high hardware costs. Furthermore, the acquisition and updating of high-precision map data are expensive, making them unsuitable for mass-produced vehicles. Simultaneously, sparse point clouds have poor adaptability. Traditional positioning algorithms based on dense point clouds suffer from insufficient geometric information in sparse point cloud scenarios with lightweight semantic maps, leading to positioning accuracy dropping to meter-level accuracy, which fails to meet the requirements of autonomous driving. In addition, candidate region selection is inefficient. Existing methods often employ global traversal matching to select candidate locations, resulting in high computational cost and long processing time, making it difficult to meet the requirements of real-time vehicle positioning (millisecond-level response). Descriptors lack expressive power; traditional descriptors rely solely on point cloud geometric coordinates (x, y, z) without incorporating semantic label information, leading to low feature discrimination in sparse point cloud scenarios and a tendency for mismatches. Finally, the matching mechanism lacks constraints, relying solely on descriptor similarity without considering spatial distance constraints, making positioning drift prone to occur in scenarios with repetitive road elements (such as continuous straight road segments).

[0005] Therefore, how to achieve a more economical, accurate, and mass-production-ready vehicle visual repositioning technology, provide a reliable backup positioning solution for intelligent driving when GNSS fails, improve the vehicle's self-positioning capability in scenarios such as tunnels and tall buildings, and provide technical support for cost reduction and efficiency improvement of autonomous driving positioning systems are technical problems that urgently need to be solved by those skilled in the art. Summary of the Invention

[0006] The purpose of this invention is to provide a vehicle relocation method, system, and electronic device based on a lightweight visual semantic map, which can be applied to autonomous vehicles and driver assistance systems. In scenarios where high-precision GNSS signals are unavailable (such as in tunnels or when obstructed by tall buildings in cities) or where only coarse-precision GNSS signals can be obtained, vehicle relocation is achieved by relying on visual information collected by low-cost cameras and a lightweight semantic map, reducing dependence on expensive satellite positioning services and improving the robustness and economy of the positioning system.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] In a first aspect, the present invention provides a vehicle relocation method based on a lightweight visual semantic map, the method comprising the following steps: The system acquires query point clouds collected in real time by a camera and delineates the map search area based on GNSS signals. The query point clouds are divided into multiple semantic layers according to semantic labels. The point clouds of each semantic layer are projected onto a two-dimensional plane and quantized into binary images. After morphological dilation, a semantic region template is formed. The semantic region template is then convolved with the map point clouds of the corresponding semantic layers within the map search area in two dimensions to generate similarity response maps of each semantic layer. These maps are then superimposed to obtain a global response heatmap, and candidate locations are selected from the global response heatmap. The query point cloud and the map point cloud corresponding to each candidate location are converted into (x,y,label) format, where x and y are location information and label is semantic label encoding. The data are then input into the DH3D network to extract global and local descriptors that fuse geometric and semantic information. The similarity between the global descriptor of the query point cloud and the global descriptors of each candidate location is calculated to filter out candidate locations that meet the preset conditions. Then, for the candidate locations that meet the preset conditions, the similarity between the query local descriptor and the candidate local descriptor is calculated, and spatial distance constraints are introduced. Finally, the comprehensive score is calculated by combining the similarity of the global descriptor and the constrained similarity of the local descriptor. The candidate location with the highest score is selected as the current accurate location of the vehicle to achieve relocalization.

[0009] Furthermore, each of the multiple semantic layers corresponds to a type of road element.

[0010] Furthermore, the morphological dilation uses a 5×5 rectangular kernel to dilate the binary image, filling the gaps in the sparse point cloud.

[0011] Furthermore, the formula for calculating the global response heatmap is:

[0012] in, This indicates the global response heatmap at pixel coordinates. The response value at the location; These are the width and height dimensions of the semantic template, respectively. For semantic category indexing, The total number of semantic categories; Representing the Class semantic templates in position Pixel value at; Representing the Map-like semantic layer in location The pixel value at that location.

[0013] Furthermore, multiple locations that meet preset conditions are extracted from the global response heatmap. Candidate locations are selected by calculating the L2 distance between any two locations and removing duplicate locations with a distance less than a preset value.

[0014] Furthermore, the semantic tags include three categories: lane lines, arrows, and zebra crossings.

[0015] Furthermore, the formula for calculating the comprehensive score is as follows:

[0016] in, This represents the overall score of the candidate positions. Indicates querying point cloud, Indicates the i-th candidate position. These are the weighting coefficients. This represents the global descriptor for querying the point cloud. Indicates candidate position global descriptor, This represents the cosine similarity between the global descriptor of the query point cloud and the global descriptor of each candidate location. This represents the average similarity of constrained local descriptors.

[0017] Secondly, the present invention also provides a vehicle relocation system based on a lightweight visual semantic map. The system includes a candidate region selection module, a descriptor extraction module, and a fusion matching module. The above-mentioned method is used to realize vehicle relocation.

[0018] Thirdly, the present invention also provides an electronic device, including a processor and a memory, the memory storing machine-executable instructions executable by the processor, the processor executing the machine-executable instructions to perform the above-described method.

[0019] As can be seen from the above technical solution, compared with the prior art, the beneficial effects of the present invention include: This invention enables more economical, accurate, and mass-production-ready vehicle visual repositioning, provides a reliable backup positioning solution for intelligent driving when GNSS fails, helps improve the vehicle's self-positioning capability in scenarios such as tunnels and tall buildings, and provides technical support for cost reduction and efficiency improvement of autonomous driving positioning systems.

[0020] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.

[0021] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0024] Figure 1 This is a schematic diagram of the core process of the vehicle relocation method provided in the embodiments of the present invention. It describes the complete logic of taking a global map containing coarse GNSS information and query samples as input, sequentially filtering candidate locations through a candidate region selection module, generating feature descriptors through a descriptor extraction module, and then calculating the relocation result through a scoring module (fusion matching module) based on flexible descriptor matching.

[0025] Figure 2 This is a schematic diagram of the feature similarity score between the query samples and search areas of the lidar point cloud and sparse point cloud provided in the embodiments of the present invention, wherein the positioning effect is intuitively presented through color gradient.

[0026] Figure 3 This is a schematic diagram of the descriptor extraction process provided in an embodiment of the present invention.

[0027] Figure 4 This is a schematic diagram of the electronic device structure provided in an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0029] In the description of this invention, it should be noted that some processes described in this application specification and drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may be performed in any order or in parallel. Furthermore, various numbers are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0030] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0031] See Figures 1 to 3 As shown, this invention discloses a vehicle relocalization method based on a lightweight visual semantic map. The core design of this invention revolves around three main objectives: "lightweight, efficient, and accurate," and is divided into three main stages: candidate region selection, descriptor extraction, and fusion matching. Specifically: Candidate area selection process: Determining 100-meter range based on coarse-precision GNSS. Within a 100-meter search range, the query point cloud collected in real-time by the camera is layered according to semantic categories (such as lane lines and arrows), projected onto a two-dimensional plane, and quantized into a binary grid image. This is then processed through 5... A 5-state dilation kernel is used to expand the convolutional template, which is then convolved with the map point cloud within the search range to generate a response heatmap. An improved NMS (Non-Maximum Suppression) is employed, which, in the candidate region selection module, filters high-response locations in the global response heatmap, eliminates duplicate or redundant suboptimal candidate locations, and retains representative candidate regions, thus reducing computational load and improving efficiency for subsequent accurate matching. This invention improves the IoU judgment logic of traditional NMS (where the Intersection over Union (IoU) is replaced with L2 distance), using L2 distance as the filtering threshold: calculating the L2 distance between any two initial candidate locations, and filtering out 10-20 high-response candidate locations.

[0032] Descriptor extraction stage: Improve the DH3D network by replacing the input dimension of the point cloud from (x,y,z) to (x,y,label) (label is a semantic label encoding, such as lane line=1, arrow=2, zebra crossing=3). The FlexConv (flexible convolution) module captures local geometric features, and the SE (Squeeze-and-Excitation) module enhances semantically related features. Finally, the global descriptor (for preliminary matching) and the local descriptor (for precise localization) are output.

[0033] like Figure 3 As shown, the network takes sparse semantic point clouds (coordinates x, coordinates y, semantic labels) as input, and generates N×128 local features through a local feature encoder. Then, it uses three layers of flexible convolution (256, 8, 1) + batch normalization + ReLU activation to adapt the sparse point cloud distribution with dynamic convolution kernels, extracting local geometric features at multiple scales and supplementing sparse region information. After that, it enters the NetVLAD layer, which contains two parallel paths: one is the attention predictor outputting N×1 attention weights, and the other is a 1×1 convolution (1024×64) + normalized exponential function. Both are input into the VLAD core module, and after L2 normalization and internal normalization, it outputs 64×1024 aggregated features. Finally, it generates a 1×256 global descriptor through a fully connected layer.

[0034] Compared to the original DH3D, its core improvements are: the input dimension is reconstructed from (x,y,z) to (x,y,label), directly incorporating semantic information; the flexible convolutional layer optimizes the dynamic kernel strategy and sampling density to adapt to sparse point clouds; the attention weight calculation of the NetVLad layer combines geometric similarity and semantic consistency to strengthen the aggregation weight of semantically related features; and the full-process parameter optimization (such as the number of convolutional kernels and the dimension of fully connected layers) improves the computational efficiency of the vehicle while ensuring feature discriminability, ultimately achieving efficient and robust global descriptor extraction from sparse semantic point clouds. The SE module is a standard component of the original DH3D network used to optimize the weights between feature channels. In this invention, its function has been upgraded and integrated into the local feature encoder and subsequent flexible convolutional layers.

[0035] In the fusion matching and localization process, the cosine similarity between the query point cloud and the global descriptors of each candidate location is calculated. At the same time, a spatial distance constraint is introduced in the local descriptor matching process, which is the Euclidean distance between the query local features and the candidate local features. When the distance is greater than 2 meters, the penalty term takes effect. Finally, the candidate location with the highest similarity is selected as the vehicle's current accurate location.

[0036] The specific embodiments of the present invention will be described in detail below: In this embodiment, vehicle relocation consists of three core steps: candidate region selection, descriptor extraction, and fusion matching. Specifically: (1) In the candidate region selection step, the query point cloud collected by the camera in real time is first obtained and divided into three layers according to the label: “lane line”, “arrow” and “zebra crossing”. Each layer is processed separately. Then, each layer of point cloud is projected onto a two-dimensional plane (ignoring height) at a distance of 0.2 meters. The 0.2-meter raster is quantized into a binary image; then, each layer of the binary image is processed using a 5-meter raster. A rectangular dilation kernel is used to perform a dilation operation, filling the gaps in the sparse point cloud to form a complete semantic region template (convolutional template); then the dilated template is compared with a 100-meter area defined by coarse-precision GNSS. A 100-meter map area is convolved to calculate the semantic similarity response of each layer, and the results are superimposed to obtain a global response heatmap. Finally, the top 50 high-response locations in the response heatmap are selected, the L2 distance between the locations is calculated, duplicate locations with a distance of less than 6 meters are removed, and 10-20 candidate locations are retained.

[0037] In this embodiment, the core of the candidate region selection step is to generate a response heatmap through semantic hierarchical convolution, and then use an improved NMS to select high-confidence candidate locations. The specific processing flow is as follows: First, semantic layering and template construction are performed. The query point cloud collected in real time by the camera is divided into N layers according to semantic labels, denoted as... Each layer corresponds to a type of road element, specifically including three categories: arrows, lane lines, and zebra crossings. To simplify calculations, the height dimension of the point cloud is ignored, and only the planar coordinates are retained. Each point cloud layer is projected onto a two-dimensional plane and quantized into integer coordinates through magnification and normalization operations to ensure that the point cloud positions match the grid dimensions of subsequent convolution operations. Then, a 5×5 rectangular morphological dilation kernel is used to dilate each quantized point cloud layer, filling the gaps in the sparse point cloud and forming a complete semantic region template. This template will be used as a convolution kernel for subsequent similarity calculation with the map point cloud.

[0038] Next, semantic similarity response calculation is performed. Based on a 100m × 100m search area defined by coarse-precision GNSS, map point clouds within this area are extracted and divided into three layers of semantic point clouds according to semantic rules consistent with the query point cloud. Semantic templates for each layer With the corresponding map semantic layer Perform two-dimensional convolution operations. and For semantically related layers, a similarity response map is obtained. The core of the convolution operation is to calculate the overlap between the template and local map regions through template sliding; the higher the overlap, the larger the response value. The response maps of the three semantic layers are then pixel-wise overlaid to generate a global response heatmap. R The calculation formula is as follows:

[0039] In the formula, These are the pixel coordinates of the global response heatmap; For semantic templates Pixel dimensions (5×5 template corresponding) , ); For semantic category index ( (Corresponding to arrows, lane lines, and zebra crossings respectively). For the first Class semantic templates in The pixel value of the location; For the first Map-like semantic layer in The pixel value of the location. This formula comprehensively measures the matching degree between the query point cloud and the map point cloud in terms of semantic category and spatial distribution by superimposing the convolutional responses of each semantic layer pixel by pixel. The higher the response value, the higher the similarity between the location and the query point cloud.

[0040] Finally, candidate locations are filtered. The global response heatmap is traversed. Extract the top 50 positions with the highest response values ​​to form an initial candidate position set. , where each position Represented in plane coordinates as , Calculate the set any two positions The L2 distance between them is given by the formula:

[0041] Set a distance threshold to filter initial candidate locations: if the L2 distance between two locations is less than or equal to the distance between the two locations, then the candidate locations are considered to be in the same position. If a position is found to be duplicated, the position with the higher response value is retained, while the position with the lower response value is discarded. After removing duplicate positions, 10-20 high-confidence candidate positions are retained. This set will serve as the target region for subsequent descriptor matching, significantly reducing the amount of subsequent computation.

[0042] (2) In the descriptor extraction step, such as Figure 3 As shown, the query point cloud and the map point clouds of each candidate location are first converted to (x, y, label) format, with the label encoded as numerical values ​​as 1 (lane line), 2 (arrow), and 3 (zebra crossing). Then, the preprocessed point cloud features are input into a three-layer flexible convolutional (256, 8, 1) layer of the improved DH3D network. Combined with batch normalization and ReLU activation functions, the local geometric features of different semantic elements (such as the triangular structure of arrows and the linear structure of lane lines) are captured through dynamic convolutional kernels. Subsequently, the data is fed into the Ne... The tVLad layer contains two parallel paths: an attention predictor and a 1×1 convolution (1024×64). The attention predictor outputs N×1 attention weights, and the 1×1 convolution is processed by a normalized exponential function. Both are input into the VLAD core module, and after L2 normalization and internal normalization, they output 64×1024 aggregated features. Finally, a fully connected layer (64×1024→256) is used to generate a 1×256 global descriptor, while retaining the local features of each point cloud for subsequent matching.

[0043] (3) In the fusion matching and positioning step, the cosine similarity between the global descriptor of the query point cloud and the global descriptor of each candidate location is calculated first, and the top 5 candidate locations with the highest similarity are initially selected. Then, for the top 5 candidate locations, the similarity between the query local descriptor and the candidate local descriptor is calculated, and spatial distance constraints are introduced: when the Euclidean distance between the query local feature and the candidate local feature is greater than 2 meters, the probability of mismatch is reduced by the distance penalty term. Finally, the comprehensive score is calculated by combining the similarity of the global descriptor and the constrained similarity of the local descriptor, and the candidate location with the highest score is selected as the current accurate location of the vehicle. The positioning accuracy can reach 0.3-0.8 meters.

[0044] In this embodiment, the core of fusion matching localization is to achieve accurate localization through "global descriptor initial screening + local descriptor constraint matching + comprehensive score sorting". The specific processing steps are as follows: 1. Global descriptor cosine similarity calculation and initial screening: First, obtain the global descriptors for the query point cloud and candidate locations. Let the global descriptor for the query point cloud be... (128-dimensional vector, (Representing the global scope), the set of global descriptors for 10-20 candidate locations obtained after candidate region filtering is: All descriptors are output from the global descriptor generation layer of the improved DH3D network, and have incorporated semantic and geometric information.

[0045] Calculate the global descriptor of the query point cloud. With each candidate location global descriptor Cosine similarity:

[0046] The numerator is the dot product of the two vectors, and the denominator is the L2 norm product of the two vectors. The similarity result ranges from [value missing]. The closer the value is to 1, the higher the matching degree.

[0047] Candidate positions are sorted from highest to lowest based on cosine similarity, and the top 5 most similar candidate positions are selected and denoted as the set. This step significantly narrows the scope of subsequent precise matching and improves computational efficiency.

[0048] 2. Local descriptor matching with constraints (including spatial distance constraints): right For each candidate position, local descriptor matching is performed and spatial distance constraints are introduced to reduce the probability of false matches.

[0049] ① Local descriptor preparation: The set of local descriptors for querying point clouds is (M represents the number of local features in the query point cloud,) An index for querying local features of a point cloud; (representing a local area) Candidate positions The set of local descriptors is ( This represents the number of local features at the candidate location. (Representing the nth local feature of the candidate location), the local descriptor is output by the local descriptor generation layer of the improved DH3D network, which preserves the fine geometric and semantic features of each point.

[0050] ② Definition of spatial distance constraints: Let's define the query for local features. The corresponding spatial coordinates are The coordinates of the center of the query point cloud are Candidate positions The center coordinates are Candidate local features Corresponding spatial coordinates .

[0051] Calculate the spatial Euclidean distance between the query local features and the candidate local features. The formula is:

[0052] Right now:

[0053] Set a distance threshold, when When the local feature matching of this group is invalid, the penalty term takes effect, and the corresponding matching similarity is counted as 0.

[0054] ③ Constrained local similarity calculation: For each query local feature Calculate its relationship with candidate positions All effective local features The cosine similarity is calculated, and a distance penalty term is introduced. The formula is:

[0055] in: To query the cosine similarity with candidate local descriptors, the calculation method is the same as the cosine similarity of global descriptors; This is the distance penalty coefficient, used to adjust the weight of the influence of distance on matching similarity; if Then local similarity .

[0056] The candidate positions are obtained by averaging the constrained similarities of all local features of the query. Local descriptor comprehensive similarity:

[0057] in, This represents the average similarity of constrained local descriptors.

[0058] 3. Overall score calculation and final positioning: By fusing global descriptor similarity and constrained local descriptor similarity, a comprehensive score for candidate positions is calculated:

[0059] in, , which is a weighting coefficient used to balance the contributions of global descriptors (initial screening and localization) and local descriptors (precise correction).

[0060] Based on overall score From high to low The candidate positions are sorted, and the candidate position with the highest score is selected as the vehicle's current accurate positioning result. This process optimizes positioning accuracy and meets the needs of practical autonomous driving applications.

[0061] From the description of the above embodiments, those skilled in the art will understand that the present invention provides a vehicle relocation method based on a lightweight visual semantic map, the technical advantages of which include: This invention reduces hardware costs, eliminating the need for LiDAR and paid high-precision GNSS, relying solely on low-cost cameras and free coarse-precision GNSS, significantly lowering hardware costs and making it suitable for mass-produced vehicles. It also adapts to sparse point clouds, supplementing insufficient geometric information with semantic tags, and combining semantically enhanced descriptors with distance-constrained matching to improve sparse point cloud positioning accuracy from meter-level to decimeter-level. Furthermore, it enhances real-time performance, with the candidate region filtering module expanding the matching range from 100 meters. The number of potential locations (approximately 10,000) within 100 meters is reduced to 10-20, significantly reducing computational load and relocation time to meet real-time requirements. Robustness is enhanced, semantic features (lane lines, arrows) are unaffected by lighting or weather, and distance constraints are introduced during matching to avoid mismatches caused by repeated road elements, thus improving the success rate of urban road relocation.

[0062] This invention is economical, with low hardware costs and a small lightweight map storage size (approximately 10MB per square kilometer). The data acquisition and update costs are only 1 / 50th of those of traditional high-precision maps. It is versatile, adaptable to low-cost cameras from different brands, requiring no customized hardware, and compatible with existing autonomous driving positioning systems, allowing for direct integration as an initialization or backup module. It is robust, exhibiting high repositioning success rates in sunny, rainy, and nighttime scenarios, and maintaining stable operation even in environments without GNSS signals, such as tunnels or tall buildings. It is accurate, achieving a positioning accuracy of 0.3-0.8 meters, meeting the requirements of L2+ level autonomous driving for lateral (lane-level) and longitudinal positioning accuracy.

[0063] Furthermore, this invention also provides a vehicle relocation system based on a lightweight visual semantic map, applying the aforementioned method to achieve vehicle relocation. This system works collaboratively through three core modules: a candidate region selection module quickly filters 10-20 high-confidence candidate locations from a coarse-precision GNSS-defined range of 100 meters, significantly reducing subsequent matching computation; a descriptor extraction module generates a hybrid descriptor that integrates geometric and semantic information, enhancing semantics and improving the feature discrimination of sparse point clouds; and a fusion matching module (a scoring module based on flexible descriptor matching) achieves accurate location estimation within the candidate region. Ultimately, this achieves low-cost (requiring only a camera and readily available, free mobile navigation software-level precision GNSS signals), low-storage (lightweight semantic map), and high robustness (resistant to lighting / weather interference) vehicle relocation, providing initialization and positioning recovery capabilities for autonomous driving.

[0064] The system provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the system embodiment can be referred to the corresponding content in the aforementioned method embodiment, and will not be repeated here.

[0065] Additionally, refer to Figure 4 As shown, this embodiment of the invention also provides an electronic device, which may include a processor, a memory, a communication bus and a communication interface, and may also include a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the above-described method.

[0066] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, software systems, electronic devices, or computer program products, etc. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0067] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0068] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A vehicle relocalization method based on a lightweight visual semantic map, characterized in that, The method includes the following steps: The system acquires query point clouds collected in real time by a camera and delineates the map search area based on GNSS signals. The query point clouds are then divided into multiple semantic layers according to semantic labels, which include lane lines, arrows, and zebra crossings. The point clouds of each semantic layer are projected onto a two-dimensional plane and quantized into binary images. After morphological dilation, a semantic region template is formed. The semantic region template is then convolved with the map point clouds of the corresponding semantic layers within the map search area in two dimensions to generate similarity response maps for each semantic layer. These maps are then superimposed to obtain a global response heatmap, and candidate locations are selected from the global response heatmap. The formula for calculating the global response heatmap is: in, This indicates the global response heatmap at pixel coordinates. The response value at the location; These are the width and height dimensions of the semantic template, respectively. For semantic category indexing, The total number of semantic categories; Representing the Class semantic templates in position Pixel value at; Representing the Map-like semantic layer in location Pixel value at; Multiple locations that meet preset conditions are extracted from the global response heatmap. By calculating the L2 distance between any two locations and removing duplicate locations with a distance less than a preset value, 10-20 candidate locations are selected. The query point cloud and the map point cloud corresponding to each candidate location are converted into a format containing location information and semantic label encoding, input into the DH3D network, and global descriptors and local descriptors that fuse geometric and semantic information are extracted. The similarity between the global descriptor of the query point cloud and the global descriptors of each candidate location is calculated to filter out candidate locations that meet preset conditions. Then, the similarity between the query local descriptor and the candidate local descriptors is calculated for each candidate location that meets the preset conditions, and spatial distance constraints are introduced. Finally, a comprehensive score is calculated by combining the global descriptor similarity and the constrained similarity of the local descriptors. The formula for calculating the comprehensive score is as follows: in, This represents the overall score of the candidate positions. Indicates querying point cloud, Indicates the i-th candidate position. These are the weighting coefficients. This represents the global descriptor for querying the point cloud. Indicates candidate position global descriptor, This represents the cosine similarity between the global descriptor of the query point cloud and the global descriptor of each candidate location. This represents the average similarity of constrained local descriptors. The candidate position with the highest score is selected as the vehicle's current precise position to achieve relocation.

2. The method according to claim 1, characterized in that, The multiple semantic layers each correspond to a type of road element.

3. The method according to claim 1, characterized in that, The morphological dilation is performed on a binary image using a 5×5 rectangular kernel to fill the gaps in sparse point clouds.

4. A vehicle relocation system based on a lightweight visual semantic map, characterized in that, The system includes a candidate region selection module, a descriptor extraction module, and a fusion matching module, and applies the method described in any one of claims 1–3 to achieve vehicle relocation.

5. An electronic device, characterized in that, It includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to perform the method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Electronic map processing method, vehicle visual repositioning method and vehicle-mounted equipment

    CN110765224A

  • Long-term relocation method and device based on high-definition map

    CN115683129A

  • Relocation method, device and equipment for cross-modal feature fusion

    CN120451271A