Capsule network-based vehicle visual relocalization method, device and medium

By using a capsule network-based visual relocalization method, semantic map generation and feature extraction are simplified, solving the problems of high computing power requirements and insufficient storage space in existing technologies. This enables efficient and interpretable vehicle localization and improves the self-localization capability of intelligent driving systems.

CN116863427BActive Publication Date: 2026-04-28CHONGQING CHANGAN AUTOMOBILE CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING CHANGAN AUTOMOBILE CO LTD
Filing Date
2023-07-13
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing visual relocalization technologies require a lot of computing power to generate semantic maps and train descriptor extraction models. Furthermore, vehicle-mounted chips struggle to process the massive amounts of data, impacting storage space. Additionally, they suffer from poor interpretability, high computing power requirements, and high map accuracy requirements, making them unsuitable for practical applications.

Method used

We employ a capsule network-based approach to identify and locate road features directly through descriptor matching, construct a global map descriptor library, and utilize prior vehicle information for local map matching to reduce reliance on semantic information. We also use the PointNet structure and capsule network to optimize visual point cloud segmentation, simplifying the feature extraction and matching process.

Benefits of technology

It achieves more accurate, efficient, and interpretable road visual descriptor extraction and visual relocalization, improves vehicle self-localization capabilities, reduces reliance on satellite navigation, lowers operating costs, and improves positioning accuracy and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116863427B_ABST
    Figure CN116863427B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle visual repositioning method and device based on a capsule network and a medium. According to coarse positioning information of vehicle driving, a global map descriptor library is constructed by extracting features from global map visual information by a descriptor extraction network, and a local map descriptor library is obtained. The descriptor extraction network extracts descriptor features from current visual information of the vehicle, and obtains visual sensing descriptors. The local map descriptor library and the visual sensing descriptors are matched and positioned to determine the current position of the vehicle. The application can be used for visual information-based repositioning under GPS drift and without GPS signals, and reduces the dependence of the vehicle on a satellite navigation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the fields of artificial intelligence and intelligent driving technology, specifically vehicle visual relocation technology based on capsule networks. Background Technology

[0002] Vehicle relocation is a key component of intelligent driving systems, commonly used for initial vehicle positioning and addressing positioning loss during operation. Initial positioning: Upon vehicle startup, a local map of a relatively large area (e.g., 100m x 100m) is known, and relocation technology is used to obtain the vehicle's precise location. After startup, continuous self-localization can be achieved using inertial navigation, visual navigation, and other technologies. Positioning loss during operation: When a vehicle enters areas with no signal, such as tunnels or underground parking garages, positioning information is lost. Relocation technology is used to update its precise location; based on the precise location provided by the relocation algorithm, continuous self-localization is achieved using inertial navigation, visual navigation, and other technologies.

[0003] Vehicle visual relocalization technology is currently a major research hotspot. Traditional visual feature-based localization methods are affected by seasons, weather, lighting, viewing angle, and occlusion, making long-term, long-range, accurate, and robust localization based on visual features difficult to achieve. To address the problems of traditional visual features, researchers have proposed localization methods based on road semantic features. These road semantic features mainly include lane lines, stop lines, road markings (arrows, etc.), poles, traffic lights, and signs. Compared to traditional visual features, these road semantic features are widely and stably present in urban road scenarios and are more robust to changes in seasons, weather, lighting, and viewing angle. With the widespread application of deep learning, road semantic features are easy to detect and extract, and have a compact representation. High-precision maps use vector representations of road semantic features, requiring minimal storage resources. This can solve the problem of high-precision localization of autonomous vehicles under limited storage conditions. Early research, such as KAIST's Road-SLAM, utilizes IPM images to construct sub-maps for ICP matching. It calculates the transformation matrix (rotation and translation) between two point clouds, using the IPM semantic features of a single forward-looking camera to construct the map. Point cloud generation is limited to the region of interest near the camera. After point cloud semantic segmentation and random forest classification, the features composed of road markings and surrounding lanes are defined as sub-maps. ICP matching between sub-maps yields the relative pose of the local map to correct drift. Further, Huawei's AVP-SLAM technology uses a U-Net neural network for semantic segmentation of images, obtaining more accurate semantic features than Road-SLAM; it also employs four surround-view cameras to generate surround-view IPM maps, resulting in a wider perception range. Alibaba proposed a SLAM localization algorithm based on sparse visual semantic features. It parameterizes some common semantic entities on the road using different methods. By tracking the features of semantic entities, it uses a Hungarian matching strategy to associate ground features in pixel space at both instance and pixel levels, thereby achieving highly efficient map matching.

[0004] In existing visual relocalization technologies, the generation of semantic maps and the training of sub-models for extracting descriptions require significant computing power. Furthermore, on-vehicle chips struggle to handle the massive amounts of data and complex model training processes simultaneously, and the scale of semantic map data also impacts the device's limited storage space. Most cutting-edge algorithms suffer from poor interpretability, high computing power requirements, and stringent map accuracy requirements, making them difficult to apply in real-world scenarios.

[0005] For example, Chinese invention patent application CN114283397A, entitled "Global Relocation Method, Apparatus, Device, and Storage Medium," obtains the observation features of the vehicle's geographic space; compares the observation features with the semantic features of various reference locations in the semantic database to determine a target reference location that meets preset positioning requirements; and determines the vehicle's initial pose relative to the feature map based on the first semantic element of the feature map, the observation features, and the target reference location, thus determining the vehicle's current pose globally. This solves the technical problems of easy relocation failure and high computational cost in existing relocation technologies. However, this method's algorithm is affected by the accuracy of semantic element recognition, and the algorithm needs to be readjusted to recognize new road signs. Summary of the Invention

[0006] In view of this, this application proposes a vehicle visual relocalization method, device and medium based on capsule network. Utilizing a capsule network structure, relocalization matching does not require semantic information and can be directly identified and located through descriptor matching. This achieves more accurate, efficient and interpretable road visual descriptor extraction and visual relocalization, providing more stable and reliable road positioning for intelligent driving, improving vehicle self-localization capabilities, and providing technical support for GPS-based intelligent driving.

[0007] The technical solution of this application is as follows:

[0008] This application implements road feature extraction and relocation based on capsule networks. The network extracts descriptors for all maps of the vehicle and stores all map descriptors and their corresponding precise coordinates locally, forming a global map descriptor library (each descriptor consists of descriptive features extracted by the capsule network and corresponding latitude and longitude). Based on the vehicle's prior information, the local location of the vehicle is obtained, and local map descriptors and their corresponding locations are extracted from the global map descriptor library.

[0009] The method of this application specifically includes: constructing a global map descriptor library based on vehicle driving coarse positioning information and a descriptor extraction network extracting features from global map visual information to obtain a local map descriptor library; the descriptor extraction network extracting descriptor features based on the vehicle's current visual information to obtain a visual sensing descriptor; and matching the local map descriptor library with the visual sensing descriptor to determine the vehicle's current position.

[0010] Further preferably, the global map descriptor library consists of key positioning points and visual descriptors of key positioning points. Based on the key positioning points and visual descriptors of key positioning points in the global descriptor library where the vehicle is located, the mutual similarity of descriptors within a predetermined range is calculated, features with similarity below a threshold are retained, and the latitude and longitude information of the corresponding key positioning points is stored to form the global map descriptor library, wherein at least one key positioning point is included within the predetermined area.

[0011] Further optimization involves obtaining coarse positioning information for the vehicle based on its historical driving trajectory and previous satellite positioning information when the vehicle is in an initial startup state or a location loss state. Based on the coarse positioning information, a certain area around the vehicle is determined as a local map, and the local map description sub-library is extracted from the global map description sub-library based on the location information of the local map.

[0012] Further, the acquisition of the visual sensing descriptor further includes: building a 3D point cloud segmentation network based on the PointNet structure; after pre-training on a radar point cloud segmentation dataset, fine-tuning it on a visual semantic point cloud; taking the output features of the highest-dimensional intermediate layer of the network as the initial vector of the capsule network; using the semantic information of the point cloud as input labels; optimizing the ability to describe the 3D point cloud map; and combining network features to form basis vectors. The basis vectors include a scale basis vector describing the scale transformation of the target and a geometric basis vector describing the geometric appearance of the target. During capsule network training, when inputting the same category label, the scale basis vector is optimized; when inputting different category labels, all basis vectors are optimized simultaneously.

[0013] Further optimization involves inputting visual sensing information into the PointNet point cloud segmentation network, outputting an initial vector, training a capsule network, obtaining dynamic paths to construct a descriptor extraction network, and obtaining descriptors through the descriptor extraction network. Specifically, a 3D point cloud segmentation network is built using the PointNet structure as the basic framework. After pre-training on a radar point cloud segmentation dataset, it is fine-tuned on a visual semantic point cloud, and the output features of the highest-dimensional intermediate layer of the network are taken as the initial vector of the capsule network.

[0014] Further preferred, the construction of the global map descriptor library includes: performing high-density segmentation of the semantic point cloud in the global map, inputting it into a descriptor extraction network to obtain a high-density descriptor set, calling a descriptor filtering model to filter descriptors, and constructing the global map descriptor library. The high-density segmentation includes: obtaining the latitude and longitude contained in the visual point cloud of the global map as the positioning information of the visual point cloud; sampling the semantic point cloud of the global map along the latitude and longitude axis at predetermined intervals to obtain sampling points; cropping the point cloud in a fixed area around each sampling point to form a high-density visual point cloud subset; determining the latitude and longitude of the semantic point cloud; sequentially inputting the visual point cloud subset into a capsule network to extract descriptor features; concatenating the descriptor features with the positioning information of the visual point cloud to store a high-density visual point cloud descriptor set; calculating the similarity of descriptors within a predetermined range in the visual point cloud descriptor set; retaining features with similarity below a threshold; and storing the latitude and longitude of the corresponding point cloud.

[0015] Further preferred, the training of the capsule network further includes that the input and output of each layer of the capsule network are in vector form. If the actual number of points in the input point cloud does not meet the number of points n in the input PointNet, then upsampling or downsampling is performed to make it contain exactly n points. The features output by the capsule network are compressed to 18*1 to obtain the descriptor.

[0016] Further optimization involves adding a fully connected layer to the last layer of the capsule network for classification. The capsule network is trained using a classification loss function until convergence. The capsule network compresses the features to 18*1 for extracting 18-dimensional features. Among these, 6-dimensional features are used to describe the scaling transformation of the input data, and 12-dimensional features are used to describe the geometric representation of the input data. The 12-dimensional features used to describe the geometric representation of the input data serve as descriptors.

[0017] Further optimization involves training a capsule network using point clouds of basic road surface markers with semantic information in the road scene to obtain the weight matrix, and then using the weight matrix W... ij Multiply by the input vector, according to the formula: get The initial vector input is denoted as... When i = 1, 2, ... 32, call the formula:

[0018]

[0019] Obtain the output vector j = 1, 2, ... 24, where c ij The coupling coefficient is denoted by , which constrains the sum of all j connected to i to be 1.

[0020] According to another aspect of this application, an electronic device is also proposed, comprising: a processor; and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform the vehicle visual relocation method based on capsule networks as described above.

[0021] According to another aspect of this application, a non-transitory computer-readable storage medium storing computer instructions is also proposed, wherein the computer instructions are used to cause the computer to perform the vehicle visual relocation method based on capsule networks as described above.

[0022] The advantages of this application are as follows:

[0023] This application presents a visual point cloud descriptor based on capsule networks for vehicle localization and matching. Unlike traditional convolutional neural networks, capsule networks construct descriptors using linear combinations of basis vectors, which better reflects road features (composed of combinations of various geometric images); they also have strong interpretability, allowing for the reverse deduction of the basis vector combination method, and are trained and optimized based on a combination of existing databases and road point cloud data.

[0024] The capsule network-based visual descriptor computation method proposed in this application can accurately describe road markings and signs using simpler features, improving the computation speed of vehicle positioning systems, saving storage space, and enhancing vehicle positioning accuracy and stability. It can be used in vehicle driver assistance systems and autonomous vehicle systems for relocation relying on visual (camera) information in the absence of GPS drift or GPS signals, reducing vehicle dependence on satellite navigation systems, lowering operating costs (annual satellite positioning fees, etc.), and improving the robustness and positioning accuracy of vehicle intelligent systems. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the vehicle repositioning system process in an exemplary embodiment of this application;

[0026] Figure 2 This is a visual descriptor extraction algorithm framework based on capsule networks in an exemplary embodiment of this application;

[0027] Figure 3 This is a schematic diagram of the PointNet network structure and feature extraction locations;

[0028] Figure 4 This is a schematic diagram illustrating the construction process of the global map description sub-library in an exemplary embodiment of this application;

[0029] Figure 5 The process of vector feature change in a capsule network, which is an exemplary embodiment of this application;

[0030] Figure 6This is a schematic diagram of an electronic device that is an exemplary embodiment of this application.

[0031] Among them, 100 is electronic equipment, 101 is processor, 102 is bus, and 103 is memory.

[0032] The above is attached. Figures 1 to 3 This paper presents a method for road visual feature descriptor extraction and vehicle relocalization based on capsule networks. Figure 1 A flowchart of the overall process of vehicle relocation is provided, and vehicle relocation is implemented based on capsule network descriptors. Figure 2 To train the visual descriptor extraction module based on capsule networks; Figure 3 Construct a flow graph for the global map description sublibrary, including the extraction and construction of the description sublibrary in the relocation framework; Figure 4 This represents the change in vector features of the capsule network. Detailed Implementation

[0033] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. It should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. These embodiments are provided to provide a more thorough and complete understanding of this application. The accompanying drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application. Based on the embodiments in this application, technical solutions obtained by those skilled in the art without creative effort are all within the scope of protection of this application.

[0034] This application proposes a method for efficiently and accurately extracting vehicle visual sensing information and effective features from prior maps based on capsule networks. By constructing concise and efficient descriptors, it achieves high-precision feature matching and vehicle relocalization, thereby improving the self-localization capability of intelligent driving systems.

[0035] like Figure 1 The diagram shown is a schematic of a vehicle relocation system in an exemplary embodiment of this application. Based on the input vehicle driving coarse positioning information, the descriptor extraction network constructs a global map descriptor library (descr, position) based on the features extracted from the global map visual information, and obtains a local map descriptor library (descr, position). The descriptor extraction network obtains visual sensing descriptors based on the current visual information of the vehicle, and matches and locates the local map descriptor library with the visual sensing descriptors.

[0036] After obtaining the approximate location of the vehicle and visual input, descriptors based on the current visual sensor information are extracted. Simultaneously, to narrow down the matching calculation range, a local map descriptor library is searched within the global map descriptor library based on the coarse location. By calculating the Euclidean distance between the descriptor of the current visual sensor information and all descriptors in the local map descriptor library, the matching descriptor within the local map descriptor library is found according to the principle of minimum distance, and the corresponding position coordinates of this descriptor are extracted and output as the current vehicle location.

[0037] Based on the vehicle's current visual sensing information, a descriptor based on the visual sensing information is extracted. The visual sensing information descriptor is matched with the local map descriptor library, and the location of the local map descriptor with the highest matching degree is obtained as the ground truth value of the current vehicle position.

[0038] Construct a global map description sub-library, and extract local description sub-libraries based on visual information for matching and localization, wherein:

[0039] Construct a global map descriptor sub-library. The prerequisite for the positioning system to operate is a global descriptor sub-library covering a large known area (such as the current city). The global descriptor sub-library consists of key positioning points and their visual descriptors, ensuring that at least one key positioning point is included within a predetermined area to achieve a certain positioning point density.

[0040] Extracting a local map description sub-library. When the vehicle is in the initial startup state or the location is lost state, coarse positioning information of the vehicle can be obtained based on its driving trajectory history, the last satellite positioning history, etc. The range of this coarse positioning information can be selected within a certain range around the key positioning point (e.g., within about 50 meters). Based on the coarse positioning information, a certain area around it is determined as a local map, and the local map description sub-library is extracted from the global map description sub-library based on the location information of the local map.

[0041] Visual information-based matching and localization. A capsule network-based visual descriptor extraction model extracts descriptors from the vehicle's current visual sensor information and compares them with a local map descriptor library to find the closest local map descriptor. This local map descriptor is then used as the vehicle's current location.

[0042] like Figure 2 The diagram shown is a schematic representation of the visual feature extraction process framework based on capsule networks in an exemplary embodiment of this application, where the PointNet portion is enlarged as shown below. Figure 3 As shown. The abbreviations MLP stand for Multi-Layer Perception, referring to a multilayer perceptron architecture; Shared indicates that the network layers share weights; Transform represents network layer transformation; and Matrix Multiply represents matrix multiplication. For PointNet with a fixed input of n point coordinates, after... Figure 3 The network structure shown has multiple transformations and links, and is arranged according to... Figure 3 The three arrows at the top point are used to extract multi-layer features in parallel to obtain 32 sets of initial vectors.

[0043] A descriptor extraction network is constructed by extracting initial vectors and training a capsule network to extract descriptors. Visual sensing information (semantic point cloud) is input into the 3D point cloud segmentation network PointNet, which outputs an initial vector. This vector is then processed through a dynamic path and the capsule network to obtain descriptors.

[0044] A 3D point cloud segmentation network was built using the PointNet architecture as the basic framework. After pre-training on a radar point cloud segmentation dataset, it was fine-tuned on visual semantic point clouds. The output features of the highest-dimensional intermediate layers of the network were used as the initial vectors (primary capsules) of the capsule network. The semantic information of the point cloud was used as input labels to optimize its ability to describe 3D point cloud maps. The network features were combined to form 18 sets of basis vectors, including scale basis vectors describing the scale transformation of the target and geometric basis vectors describing the geometric appearance of the target. During capsule network training, when the input labels were of the same category, the scale basis vectors were optimized; when the input labels were of different categories, all basis vectors were optimized simultaneously.

[0045] The PointNet used in this embodiment is a classic 3D point cloud segmentation network structure, as follows: Figure 3 Location-based feature extraction yields 64, 64, and 1024-dimensional feature descriptions for n points, respectively. Since (64+64+1024) = 32*36, these three sets of features can be concatenated in parallel and then split to form 32 initial vectors (36n*32). The radar point cloud segmentation dataset used is the S3DIS database, a large-scale indoor scene semantic point cloud dataset publicly available from Stanford University. PointNet is trained directly using its point clouds and labels. The reason for initializing the network on the radar point cloud segmentation dataset is that this dataset has a large amount of data and diverse data formats similar to visual point cloud data, making it suitable as pre-training data for this task.

[0046] The global map descriptor library is constructed using a "descriptor extraction network". It involves a one-time collection and extraction of the global map to form the global map descriptor library, specifically including: security level segmentation, descriptor extraction, and descriptor filtering.

[0047] like Figure 4 The diagram shown is a flowchart of the construction of the global map description sublibrary in an exemplary embodiment of this application.

[0048] The semantic point cloud in the global map is densely segmented and input into a descriptor extraction network to obtain a dense set of descriptors. Each point cloud is ultimately recorded as (f1, f2, ..., f12, lat, ton), where f1-f12 represent the 12-dimensional descriptor features extracted by the capsule network; lat and ton represent the precision and latitude of the center point position of the corresponding point cloud, respectively. A descriptor filtering model is used to filter descriptors and construct a global map descriptor library. Based on the obtained dense set of descriptors and the corresponding descriptor filtering algorithm model, filtering and matching are performed in the global map library to obtain the descriptive information in the final global descriptor library, and then the final localization result in the library is obtained.

[0049] (1) Dense-level segmentation is used to determine the latitude and longitude of the semantic point cloud. Visual mapping algorithms can be used to construct the point cloud. The precise latitude and longitude (lon, lat) information contained in the visual point cloud of the global map is obtained as the positioning information of the point cloud. The semantic point cloud of the global map is then densely sampled along the latitude and longitude axes at predetermined intervals (e.g., 0.1 meters). For example, lane-level positioning accuracy is generally required to be 0.2 meters; setting it to 0.1 meters can meet the lane-level accuracy requirements of fully automated driving. A dense-level subset of the visual point cloud is formed by cropping the point cloud from a fixed area (e.g., 1 meter * 1 meter * 1 meter) centered on each sampling point.

[0050] (2) Descriptor Extraction. The descriptor extraction part sequentially inputs the visual point cloud subsets (all visual point clouds of size 1m*1m*1m) obtained from the dense-level segmentation into the capsule network to extract descriptor features. The descriptor features are concatenated with the localization information (lon, lat) of the visual point clouds and stored as a dense-level visual point cloud descriptor set. The dimension of the descriptor features is determined based on the size of the output features of the trained capsule network. When adjusting the dimension, it is necessary to retrain the capsule network with the corresponding length of descriptors. In this embodiment, 12-dimensional descriptor features are optimally extracted.

[0051] (3) Descriptor filtering. The set of density-level visual point cloud descriptors obtained from the extracted visual point cloud is simplified using a descriptor filtering algorithm to facilitate storage and accelerate matching.

[0052] Specifically, this involves calculating the mutual similarity of descriptors within a predetermined range. This range can be determined based on basic navigation requirements (e.g., a 1m x 1m range, which generally meets navigation needs). Features with similarity below a threshold are retained (e.g., features with significant differences (3%–10%), such as road arrows and markings). Simultaneously, the actual latitude and longitude information of the remaining points after filtering is stored, thus obtaining a global map descriptor library. This global map descriptor library is stored locally on the vehicle and contains numerous descriptors of the global map and their corresponding latitude and longitude locations. It is used to obtain a local map descriptor library and provide positioning matching.

[0053] Using the semantic information of point clouds as input labels, the network's point cloud semantic segmentation function is trained to optimize its ability to describe 3D point cloud maps.

[0054] Capsule network training. The capsule network is trained using point clouds of basic road landmarks with semantic information, such as lane lines, arrows, and zebra crossings. These landmark point clouds are input into the trained segmentation network to extract initial vectors, which are then used as input data to train the capsule network.

[0055] like Figure 5 The diagram illustrates the vector feature change process of a capsule network according to an exemplary embodiment of this application.

[0056] The capsule network used in this embodiment has 4 layers, specifically including: the first layer (36n)×32D, the second layer (18n)×32D, the third layer (18n)×24D, and the fourth layer (18n)×18D, outputting an 18-dimensional descriptor; the network can also be designed with other layers.

[0057] In the capsule network, the input and output of each layer are in vector form. The feature vector change process of the 4-layer network is shown in the figure. The network gradually reduces the input features from 32 dimensions and 36n points to 18 dimensions and 18n points (n is the number of points in the input PointNet, which is set to a fixed value. If the actual number of points in the input point cloud does not meet the fixed value, it is upsampled or downsampled to make it contain exactly n points). The capsule network finally directly compresses the features to 18*1 to obtain the descriptor required by this application.

[0058] The following section uses one of the layers as an example to illustrate the propagation process of vector features between layers.

[0059] A capsule network is trained using point clouds of basic road markers with semantic information in a road scene to obtain the weight matrix. Assume the input vector dimension of a certain network layer is N, and the output vector dimension is M: i = 1, 2, ..., N; j = 1, 2, ..., M. Assume the number of input and output vectors of this layer is X and Y, respectively. X N-dimensional vectors are stacked into an X×N matrix, and its row vectors are denoted as... Stack Y M-dimensional vectors into a Y×M matrix, and denote its row vectors as . The weight matrix W learned by the network layer ij ∈R X×Y The weight matrix W ij Multiply by the input vector, according to the formula: get

[0060] The initial row vector input is denoted as... When i = 1, 2, ..., N, call the formula:

[0061]

[0062] Obtain the output row vector j = 1, 2, ..., M.

[0063] Among them, c ij The coupling coefficient is such that the sum of all j connected to i is 1.

[0064]

[0065] In the above formula, i is the dimension index of the network layer input vector, and j is the dimension index of the network layer output vector. For all input values ​​in the i-th dimension All output values ​​for the j-th dimension The influence of the weight matrix W learned by the network. ij The decision is made. Since the number of input points in PointNet is fixed at n, the size of the matrix formed by the superimposed input features of the first layer of the network is 36n×32. Substituting into the above formula, we get X = 36n and N = 32.

[0066] During training, a fully connected layer is cascaded into the last layer of the capsule network for classification, and the capsule network is trained using a classification loss function until convergence. The capsule network ultimately compresses the features to 18*1, and the last layer is used to extract 18-dimensional feature descriptors. Of the resulting 18-dimensional features, 6 dimensions can be selected to describe the scaling transformation of the input data (e.g., when the device is fixed, target sensing information only involves translation and rotation 6DoF), and 12 dimensions are used to describe the geometric representation of the input data (assuming road signs can be composed of basic geometric structures and can be represented as a linear combination of different geometric structures). The 12-dimensional features used to describe the geometric representation of the input data serve as descriptors.

[0067] To ensure that each set of features can describe the corresponding data characteristics, during capsule network training, when the input is a label of the same category (such as an arrow, a crosswalk, a stop line, etc.), the feature descriptions of the geometric representation of the input data (the last 12 dimensions) in the 18-dimensional features are kept consistent, and only the features that control the scaling transformation (the first 6 dimensions) are optimized; when the input is a label of a different category, all 18-dimensional features are optimized at the same time.

[0068] After the capsule network is trained, a visual point cloud of a fixed-size region is input, and the trained capsule network outputs the corresponding descriptor. The last 12 dimensions of the 18-dimensional features output by the capsule network are used as the extracted visual descriptor.

[0069] like Figure 6 As shown, the electronic device 100 includes a processor 101 and a memory 103. The processor 101 and the memory 103 are connected, for example, via a bus 102.

[0070] The structure of the electronic device 100 does not constitute a limitation on the embodiments of this application.

[0071] Processor 101 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 101 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0072] Bus 102 may include a path for transmitting information between the aforementioned components. Bus 102 may be a PCI bus or an EISA bus, etc. Bus 102 may be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0073] The memory 103 may be a ROM or other type of static storage device capable of storing static information and instructions, RAM or other type of dynamic storage device capable of storing information and instructions, or it may be an EEPROM, CD-ROM or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0074] The applicant has provided a detailed description of the implementation examples of this application in conjunction with the accompanying drawings. However, those skilled in the art should understand that the above embodiments are merely preferred embodiments of this application. The detailed description is only intended to help better understand the spirit of this application and is not intended to limit the scope of protection of this application. Those skilled in the art to which this application pertains can make various modifications or additions to the specific embodiments described or use similar methods to replace them, but without departing from the spirit of this application or exceeding the scope defined by the appended claims.

[0075] Based on the embodiments described above, and through the above explanation, those skilled in the art can make various changes and modifications without departing from the inventive concept. The scope of this invention is not limited to the specific embodiments described, but should be determined according to the scope of the claims.

Claims

1. A vehicle visual relocalization method based on capsule networks, characterized in that, A global map descriptor library is constructed based on vehicle driving coarse positioning information and features extracted from global map visual information. A local map descriptor library is also extracted. The descriptor extraction network extracts descriptor features based on the vehicle's current visual information to obtain visual sensing descriptors. The local map descriptor library is matched with the visual sensing descriptors to determine the vehicle's current position. The method also includes: performing density-level segmentation on the semantic point cloud in the global map, inputting the segmentation into the descriptor extraction network to obtain a set of density-level descriptors, calling a descriptor filtering model to filter descriptors, and constructing the global map descriptor library. The density-level segmentation includes: obtaining the visual point cloud of the global map. The latitude and longitude included are used as the positioning information of the visual point cloud. The semantic point cloud of the global map is sampled at a predetermined interval along the latitude and longitude axis to obtain sampling points. The point cloud in the surrounding fixed area is cropped with each sampling point as the center to form a dense visual point cloud subset, and the latitude and longitude of the semantic point cloud are determined. The visual point cloud subset is sequentially input into the capsule network to extract descriptor features. The descriptor features and the positioning information of the visual point cloud are concatenated and stored as a dense visual point cloud descriptor subset. The similarity of descriptors within a predetermined range in the visual point cloud descriptor subset is calculated, and features with similarity below a threshold are retained and the latitude and longitude of the corresponding point cloud are stored.

2. The method according to claim 1, characterized in that, The global map descriptor library consists of key positioning points and visual descriptors of key positioning points. Based on the key positioning points and visual descriptors of key positioning points in the global descriptor library where the vehicle is located, the mutual similarity of descriptors within a predetermined range is calculated. Features with similarity below a threshold are retained, and the latitude and longitude information of the corresponding key positioning points is stored to form the global map descriptor library. The predetermined area contains at least one key positioning point.

3. The method according to claim 1, characterized in that, The extraction of the local map description sub-library includes: when the vehicle is in the initial start-up state or the positioning is lost state, obtaining its coarse positioning information based on the vehicle's driving trajectory history and the last satellite positioning information, determining a certain area around it as a local map based on the coarse positioning information, and extracting the local map description sub-library from the global map description sub-library based on the location information of the local map.

4. The method according to claim 1, characterized in that, Visual sensing information is input into the PointNet point cloud segmentation network, which outputs an initial vector. The capsule network is then trained to obtain dynamic paths and construct the descriptor extraction network. Descriptors are then obtained through the descriptor extraction network.

5. The method according to claim 4, characterized in that, It also includes building a 3D point cloud segmentation network based on the PointNet structure. After pre-training on a radar point cloud segmentation dataset, it is fine-tuned on a visual semantic point cloud. The output features of the highest-dimensional intermediate layer of the network are taken as the initial vector of the capsule network. The semantic information of the point cloud is used as the input label, and the basis vectors are composed of network features. The basis vectors include scale basis vectors used to describe the target scale transformation and geometric basis vectors used to describe the target geometric appearance. In the training of the capsule network, when the input labels are of the same category, the scale basis vector is optimized. When the input labels are of different categories, all basis vectors are optimized at the same time.

6. The method according to claim 4, characterized in that, The training capsule network further includes the following steps: the input and output of each layer of the capsule network are in vector form. If the actual number of points in the input point cloud does not meet the number of points n in the input PointNet, upsampling or downsampling is performed to make it contain exactly n points. The features output by the capsule network are compressed to 18*1 to obtain descriptors. A fully connected layer is connected in series at the last layer of the capsule network for classification. The capsule network is trained until convergence using a classification loss function. The capsule network compresses the features to 18*1 to extract 18-dimensional features, of which 6-dimensional features are used to describe the scale transformation of the input data, and 12-dimensional features are used to describe the geometric representation of the input data. The 12-dimensional features used to describe the geometric representation of the input data are used as descriptors.

7. The method according to claim 5 or 6, characterized in that, A capsule network is trained using point clouds of basic road surface markers with semantic information in a road scene to obtain a weight matrix. Multiply by the input vector, according to the formula: ,get The initial vector input is denoted as... For i=1,2,...32, call the formula: Obtain the output vector j=1,2,...24, where, The coupling coefficient is denoted by , which constrains the sum of all j connected to i to be 1.

8. An electronic device, comprising: processor; And a memory for storing a program, characterized in that the program includes instructions that, when executed by the processor, cause the processor to perform the vehicle visual relocation method based on a capsule network according to any one of claims 1-7.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, in, The computer instructions are used to cause the computer to execute the vehicle visual relocation method based on capsule networks according to any one of claims 1-7.

Citation Information

Patent Citations

  • Global repositioning method and device, equipment and storage medium

    CN114283397A

  • Rapid repositioning method and device based on visual semantic information

    CN110189373A

  • Visual positioning method based on capsule network

    CN112348038A