A Global Localization Method for Structured Road Autonomous Commercial Vehicles Based on Transformer
By adopting a structured road global positioning method based on Transformer in autonomous driving commercial vehicles, combined with deep learning and self-attention mechanism, the problem of high-precision positioning in structured road scenarios is solved, and the global positioning effect with high precision and low cost is achieved.
Patent Information
- Application Number
- CN202410828172.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-06-25
AI Technical Summary
The existing technology is difficult to achieve high-precision global positioning in structured road scenarios, the GPS positioning accuracy is insufficient, the GNSS/RTK method is costly and the signal is easily affected by occlusion, and traditional high-precision inertial navigation has cumulative error problems.
The global positioning method of structured road autonomous driving commercial vehicles based on Transformer is adopted, and road structure information is extracted through the DeepLabV3Plus semantic segmentation network, and combined with the Transformer self-attention mechanism and LSTM network, high-precision self-positioning of vehicles on structured roads is achieved.
It improves the accuracy and robustness of vehicle positioning, reduces interference from dynamic objects, enhances the universality and timeliness of the algorithm, and reduces positioning costs.
Smart Images

Figure CN118691779B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a global positioning method for a structured-road autonomous commercial vehicle based on Transformer, belonging to the technical field of vehicle driving positioning. Background Art
[0002] With the continuous progress of technology, the performance and stability of autonomous driving systems have been greatly improved. Commercial vehicles have taken the lead in using full-series autonomous driving systems to serve logistics transportation, public transportation, and urban services, such as automatic distribution of park logistics, intelligent urban buses, and autonomous driving taxis. The application of commercial vehicle autonomous driving technology has an obvious effect on reducing labor costs, improving transportation efficiency, and enhancing driving safety. With the continuous progress of autonomous driving technology, the commercial application of commercial vehicle autonomous driving will be further accelerated. And high-precision positioning information is the prerequisite for vehicle autonomous navigation and path planning, and is an important part of the autonomous driving strategy. Moreover, for commercial vehicle autonomous driving, accurate positioning information is even more important because they need to more accurately identify the position of each key point. For commercial vehicles driving on urban structured roads, the Global Positioning System (GPS) is the mainstream method for vehicle positioning, but the positioning accuracy is at the meter level, which cannot meet the positioning accuracy requirements of intelligent vehicles. The Global Navigation Satellite System (GNSS) / Real-Time Kinematics (RTK) is a vehicle positioning method that can achieve centimeter-level positioning accuracy. However, satellite signals are sensitive to the occlusion of urban buildings and trees, the positioning signal is prone to loss, and the cost of signal base station construction is relatively high, which greatly increases the cost of commercial vehicle autonomous driving and cannot be widely used for global positioning on structured roads.
[0003] Currently, in the structured road scenario, the mainstream high-precision positioning methods are based on high-precision maps, SLAM, or high-precision inertial navigation, but they have the disadvantages of high cost or cumulative errors. With the continuous development of computer vision and deep learning, low-cost camera sensors have become the main research object for realizing high-precision positioning technology. However, the information collected by cameras often has unstable positioning features, such as pedestrians, vehicles, and other variable features, and objects with stable and reliable positioning features, such as road edges, lane lines, traffic signs, etc., are often blocked by pedestrians or passing vehicles. This affects the positioning accuracy. How to extract stable and reliable features and utilize these features has become the key to visual positioning technology. With the development of deep learning, methods for realizing the absolute pose of cameras using deep learning have been widely studied. The powerful feature extraction ability and learning ability of deep learning can effectively solve the problems of vehicle and pedestrian occlusion and lighting changes. The method for estimating the absolute pose of a camera based on deep learning can directly realize the end-to-end output of the absolute pose. Through the learning of a large number of training samples, accurate positioning results can be obtained. However, when the complete real-time collected image is input into the network, it will not only be interfered by dynamic objects, reducing the positioning robustness. And structured roads often have complete structured information such as lane lines, road edges, and traffic signs. Utilizing the structural information can not only improve the positioning robustness but also improve the universality of the algorithm. Semantic segmentation is used to capture this structural information, and the Transformer self-attention mechanism is used to macroscopically extract this structural information for positioning. Finally, the information is input into the LSTM. The LSTM can learn time series data, which can effectively make up for the disadvantage of inaccurate positioning when the structured road is blocked, and further improve the positioning accuracy of the vehicle. Summary of the Invention
[0004] Object of the Invention: Aiming at the deficiencies in the prior art, the present invention provides a global positioning method for structured road autonomous commercial vehicles based on Transformer. Through the pre-collected road real-scene images and real pose information, the constructed deep learning network is trained to achieve high-precision self-positioning of vehicles in the structured road scenario.
[0005] Technical Solution: A global positioning method for structured road autonomous commercial vehicles based on Transformer includes the following steps:
[0006] Step 1: Image acquisition, image processing, and dataset production:
[0007] Collect the image features of the target city roads, including roads, lane lines, road markings, traffic lights, traffic signs, poles, sidewalks, buildings, and trees, and record the location information of each image at the same time; perform pixel preprocessing on the images, that is, reshape the original images to a size of 513×513, and form a dataset for training and testing the network with the reshaped images.
[0008] Step 2, Establish a network model, which includes the following three parts:
[0009] The first part of the network model is a model for perceiving road structure information. Structured roads contain a large amount of invariant structure information, such as roads, lane lines, road markings, traffic lights, traffic signs, poles, sidewalks, buildings, and trees. Perform semantic segmentation on the original images after reshaping through the DeepLabV3Plus semantic segmentation network to obtain road structure information. On the basis of the original DeepLabV3Plus semantic segmentation network model, introduce an attention module; in the encoder of the original DeepLabV3+ model, the high-level features generated by the backbone network are directly sent to ASPP. The ASPP module uses multiple parallel atrous convolution layers with different sampling rates, and the features extracted for each sampling rate are further processed in separate branches and fused to generate the final result. The ASPP method obtains multi-scale features from the image and obtains image context information. However, due to the too large dilation rate of the ASPP structure, spatial image information will be lost and it is impossible to extract image edge features well. That is, add a PSA module after the ASPP module. The PSA module consists of channel attention and spatial attention, which increases the extraction of important information and improves the utilization rate of information by the model;
[0010] The second part of the network model is an image spatial feature extraction model, which extracts robust position feature information from the image through the Transformer variable self-attention mechanism and encodes these features;
[0011] When facing continuous similar scenes, the image feature vectors of the similar scenes are mutually calculated by cosine. When the cosine value is between 0.4 and 0.6, the effectiveness of the feature extraction model can be proved;
[0012] Adopt the overall architecture based on the sliding window of Swin Transformer, remove the last fully connected layer, and output a 768-dimensional feature vector, which is regarded as the feature vector of the image to be located;
[0013] The third part of the network model is the global pose prediction model. The 768-dimensional feature vector output by the feature encoder is used as the input of the temporal network. The LSTM (Long Short-Term Memory) network has strong learning ability in temporal data. LSTM controls the transmission of information by introducing the concept of "gates", enabling it to capture and remember the dependencies of image feature vectors at long time intervals, and thus more accurately predict the positioning information of the vehicle;
[0014] Step 3: Training of the network model:
[0015] According to the perception road structure information model, image spatial feature extraction model, and global pose prediction model in Step 2, conduct separate training respectively; divide the image data collected in Step 1 into a training set, a validation set, and a test set according to a ratio of 3:1:1;
[0016] Step 4: Positioning of the vehicle pose, including:
[0017] Use the Extended Kalman Filter to integrate the IMU and the network output positioning results. First, obtain the state to be estimated and the system noise equation, then construct the measurement equation and further calculate the system noise variance matrix. When the network output positioning result and the IMU data are updated, update the state of the fusion system by calculating the Kalman filter gain, and finally output the fusion positioning result to obtain the vehicle pose state variables as where x i, y i, are the vehicle position and heading angle in the two-dimensional plane respectively.
[0018] The attention module in Step 2 includes two parts: channel attention and spatial attention. The calculation of the channel attention weight is as follows:
[0019]
[0020] In the formula, σ1 and σ2 represent tensor integer type operators, F SM represents the softmax function, represents a 1×1 convolution and LN layer, which increases the channel dimension from C / 2 to C, F SG represents the sigmoid function;
[0021] The calculation of the spatial attention weight is as follows:
[0022] A sp (X) = F SG [σ 3 (F SM (σ 1 (F GP (W q (X)))) × σ 2 (W v(X)))] (2)
[0023] In the formula, σ1, σ2, and σ3 represent the tensor integer type operator, F SM represents the softmax function, F GP represents global pooling, F SG represents the sigmoid function, and a concatenated structure is adopted, that is, the feature map X first performs channel attention calculation and then spatial attention calculation.
[0024] In step 2, the Swin Transformer uses the self-attention mechanism entirely, that is:
[0025] q = xW q, k = xW k, v = xWv (3)
[0026]
[0027] z = Concat(z (1) ,...,z (m) )W o (5)
[0028] In the formula, x is the feature vector matrix obtained by flattening the feature map generated by the convolution of the original input image, m is the number of heads of the multi-head attention mechanism, σ(·) is the softmax function, d is the vector length, z (m) represents the output matrix of the m-th attention head, q (m) , k (m) and v (m) respectively represent the query matrix, the matrix to be queried, and the actual feature matrix generated from the feature map; W q , W k , W v are the weight matrices for generating q (m) , k (m) , v (m) respectively, and W o is the multi-attention weight matrix, and z is the final output feature vector matrix;
[0029] The deformable attention mechanism is used to replace the self-attention mechanism to improve the operation speed. The image feature map consists of at least 2401 tokens, that is, a 49*49 rectangular token array. By learning a set of offsets, the tokens that are beneficial for attention calculation among the current token and the remaining tokens of the image are determined, and attention calculation is performed on the important points; among them, the offset is a set of two-dimensional vectors, pointing from the current token to the beneficial token;
[0030] Given the input feature map x ∈ R generated from the original image with height, width, and number of channels H, W, C H×W×C, Generate several position points
[0031] First, assign position information to each point, that is, {(0,0),...(H G-1, W G-1 )}, and then normalize all coordinates according to the grid shape H G ×W G , determine the relative position, that is, (-1,-1) represents the upper left corner, (+1,+1) represents the lower right corner; offset:
[0032] △p = θ offset (q) (6)
[0033] In the formula, q is the query vector of the feature point where the attention mechanism is to be executed currently, and θ offset (·) is a lightweight convolutional neural network. The neural network is two convolutional modules with non-linear activation. The input is the query vector q generated from the feature map. First, a 5×5 depth convolution is used to capture local features, and then the GELU activation function and 1×1 convolution are used to obtain the two-dimensional offset, and then the deformable attention mechanism is obtained:
[0034]
[0035] In the formula, φ(;) is the bilinear interpolation function, that is, the feature points important to the current point are obtained through the bilinear interpolation function are the query vector and the actual feature vector of the important feature points;
[0036] Perform multi-head attention calculation on q, k, v to obtain the final result
[0037]
[0038] Reshape the feature vector into a 4×192 matrix, use four LSTM networks, pass it to a fully connected layer, and perform positioning prediction.
[0039] For the DeepLabV3Plus network, 9 categories are defined, namely road, lane line, road marking, traffic light, traffic sign, pole, sidewalk, building, and tree; use the Labelme software to annotate the original image, and set the learning rate l r to be 0.01, and use the GPU to train it;
[0040] For the Transformer feature encoder, use the pre-trained model for the classification task on the ImageNet dataset as the pre-trained model of this application;
[0041] For the LSTM network, the true pose l, define the loss function:
[0042] loss(I) = (l - l*) 2 (10)
[0043] Where l is the true pose and l* is the predicted pose.
[0044] The specific steps of step 4 are as follows:
[0045] Let the estimated state X i at time t i be driven by the system noise G i and is expressed in equation form as:
[0046] X i = A i,i-1 X i + W i-1 G i-1 (11)
[0047] where A i,i-1 is the transition matrix from time t i-1 to time t i and W i-1 is the system noise driving matrix.
[0048] Let L i be the GPS data, whose value satisfies a linear relationship, and the measurement equation is:
[0049] Z i = H i L i + V i (12)
[0050] where H i is the measurement matrix and V i is the measurement noise sequence. In this system, there is:
[0051]
[0052] Let Q i be the system noise sequence variance matrix and R i be the measurement noise sequence variance matrix, which is a non - negative definite matrix in this system;
[0053] When initializing the filter, take the first - frame GPS position and coordinates, that is, the positioning information of the network prediction model as the initial value of the filter state
[0054] When the speed information and the angular velocity information of the gyroscope are updated, the predicted solution is obtained by dead reckoning, and the system noise variance matrix Q i is:
[0055]
[0056] The transition matrix A from time i-1 to time i i,i-1 is as follows:
[0057]
[0058] For W i-1 , when the vehicle steering angle between the current and previous moments is <α, where α is the vehicle steering angle, there is:
[0059]
[0060] When the vehicle pose steering angle between the current and previous moments >α, there is:
[0061]
[0062] where td is the distance the vehicle travels during the time period, and tr is the angle the vehicle turns during the time period, that is:
[0063] td = v * dt, tr = w * dt (18)
[0064] Furthermore, the predicted mean square error matrix is obtained:
[0065] P i / i-1 = A i,i-1 * P i-1 * A T i,i-1 + W i * Q i-1 * W T i-1 (19)
[0066] When the positioning information output by the network model is updated, the positioning result is fused with the current dead reckoning result; according to the current network model, the measurement noise variance matrix is obtained:
[0067]
[0068] where ed is the distance estimation noise;
[0069] Calculate the filtering gain K i :
[0070] K i = P i / i-1 * H T i *(H i * P i / i-1 * H T i + R i ) (21)
[0071] Then the final state estimate is:
[0072] Y i = Z i - H i * X i / i-1 (22)
[0073] X i = X i / i-1 + K i * Y i (23)
[0074] Finally, update the mean square error matrix P, where E is the 3×3 identity matrix:
[0075] P i = (E - K i * H i ) * P i / i-1 (24)
[0076] Final output That is, the vehicle position and heading angle.
[0077] Beneficial effects:
[0078] After adding the PSA attention module, semantic segmentation captures the detailed information of structured roads, removes dynamic objects, retains the image space features, ensures the robustness of positioning, and also improves the generality
[0079] The Transformer can better extract features strongly related to positioning through its powerful self-attention mechanism. Replacing the self-attention mechanism with the deformable attention mechanism can better adapt to the image after semantic segmentation, increasing the timeliness and accuracy.
[0080] When there are situations such as blurred occlusion and ghosting in a few frames, the LSTM network can effectively improve the above situations and ensure the robustness of positioning information. Brief description of the drawings
[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0082] Figure 1 It is the overall flowchart of the present invention.
[0083] Figure 2 It is the structural diagram of the improved DeepLabV3Plus of the present invention.
[0084] Figure 3 It is the structural diagram of the PSA of the present invention.
[0085] Figure 4 This is the structural diagram of the Transformer encoder of the present invention.
[0086] Figure 5 This is the network model diagram of the present invention. Specific implementation manners
[0087] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0088] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0089] In the present invention, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may include direct contact between the first and second features, or may include the first and second features not being in direct contact but being in contact through other features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes the first feature being directly above and obliquely above the second feature, or simply indicating that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "below" and "beneath" the second feature includes the first feature being directly below and obliquely below the second feature, or simply indicating that the horizontal height of the first feature is lower than that of the second feature.
[0090] As Figure 1 shown, a global positioning method for a structured road autonomous commercial vehicle based on Transformer includes the following steps:
[0091] Step 1: Image acquisition, image processing, and dataset production:
[0092] Collect the image features of the target urban road, including roads, lane lines, road markings, traffic lights, traffic signs, poles, sidewalks, buildings, and trees, and record the position information of each image at the same time; perform pixel preprocessing on the images, that is, reshape the original images to a size of 513×513, and form a dataset for training and testing the network with the reshaped images.
[0093] Step 2: Establish a network model, including the following three parts:
[0094] The first part of the network model is a model for perceiving road structure information. Structured roads contain a large amount of invariant structure information, such as roads, lane lines, road markings, traffic lights, traffic signs, poles, sidewalks, buildings, and trees. The original image after reshaping is semantically segmented through the DeepLabV3Plus semantic segmentation network to obtain road structure information.
[0095] Based on the original DeepLabV3Plus semantic segmentation network model, an attention module is introduced; in the encoder of the original DeepLabV3+ model, the high-level features generated by the backbone network are directly fed into ASPP. The ASPP module uses multiple parallel atrous convolution layers with different sampling rates, and the features extracted for each sampling rate are further processed in separate branches and fused to generate the final result. The ASPP method obtains multi-scale features from the image and acquires image context information.
[0096] However, due to the too-large dilation rate of the ASPP structure, spatial image information will be lost, and it is impossible to extract image edge features well. Therefore, a PSA module is added after the ASPP module. The PSA module consists of channel attention and spatial attention, which increases the extraction of important information and improves the utilization rate of information by the model;
[0097] The attention module in Step 2 includes two parts: channel attention and spatial attention. The calculation of the channel attention weight is as follows:
[0098]
[0099] In the formula, σ1 and σ2 represent tensor integer type operators, F SM represents the softmax function, represents a 1×1 convolution and an LN layer, which increases the channel dimension from C / 2 to C, F SG represents the sigmoid function;
[0100] The calculation of the spatial attention weight is as follows:
[0101] A sp (X) = F SG [σ 3 (F SM (σ 1 (F GP (W q (X)))) × σ 2 (W v (X)))] (2)
[0102] In the formula, σ1, σ2, and σ3 represent tensor integer type operators, FSM represents the softmax function, F GP represents global pooling, F SG represents the sigmoid function, adopting a concatenated structure, that is, the feature map X first performs channel attention calculation and then spatial attention calculation.
[0103] The second part of the network model is an image spatial feature extraction model, which extracts robust position feature information from the image through the Transformer variable self-attention mechanism and encodes these features;
[0104] When facing continuous similar scenes, the image feature vectors of the similar scenes are mutually cosine calculated. When the cosine value is between 0.4 and 0.6, the effectiveness of the feature extraction model can be proved;
[0105] Adopt the overall architecture based on the sliding window of Swin Transformer, remove the last fully connected layer, and output a 768-dimensional feature vector, which is regarded as the feature vector of the image to be located;
[0106] In step 2, all self-attention mechanisms are used in Swin Transformer, that is:
[0107] q = xW q, k = xW k, v = xWv (3)
[0108]
[0109] z = Concat(z (1) ,..., z (m) )W o (5)
[0110] In the formula, x is the feature vector matrix obtained by flattening the feature map generated by the original input image through convolution, m is the number of heads of the multi-head attention mechanism, σ(·) is the softmax function, d is the vector length, z (m) represents the output matrix of the m-th attention head, q (m) , k (m) and v (m) respectively represent the query matrix, the matrix to be queried, and the actual feature matrix generated from the feature map; W q , W k , W v are the weight matrices for generating q (m) , k (m) , v (m) respectively, W o is the multi-attention weight matrix, and z is the final output feature vector matrix;
[0111] Replace the self-attention mechanism with a deformable attention mechanism to improve the operation rate. The image feature map consists of at least 2401 tokens, that is, a 49*49 rectangular token array. By learning a set of offsets, the tokens that are beneficial for attention calculation among the current token and the remaining tokens in the image are determined, and attention calculation is performed on the key points; where the offsets are a set of two-dimensional vectors pointing from the current token to the beneficial tokens.
[0112] Given an input feature map \(x\in\mathbb{R}\) generated from an original image with height, width, and number of channels \(H, W, C\) H×W×C , generate a number of position points
[0113] First, assign position information to each point, that is, \(\{(0,0),...(H\) G-1, W G-1 \()\}\). Then, normalize all the coordinates according to the grid shape \(H\) G× W G , determine the relative positions, that is, \((-1,-1)\) represents the upper left corner, and \((+1,+1)\) represents the lower right corner; offsets:
[0114] \(\Delta p=\theta\) offset (q)(6)
[0115] In the formula, \(q\) is the query vector of the feature point where the attention mechanism is to be executed currently, and \(\theta\) offset (·) is a lightweight convolutional neural network. The neural network consists of two convolutional modules with non-linear activation. The input is the query vector \(q\) generated from the feature map. First, a 5×5 depth convolution is used to capture local features, and then the GELU activation function and a 1×1 convolution are adopted to obtain two-dimensional offsets, and then the deformable attention mechanism is obtained:
[0116]
[0117] In the formula, \(\varphi(;)\) is a bilinear interpolation function, that is, the feature points important for the current point are obtained through the bilinear interpolation function are the query vector and the actual feature vector of the important feature points;
[0118] Perform multi-head attention calculation on \(q, k, v\) to obtain the final result
[0119]
[0120] The third part of the network model is the global pose prediction model. The 768-dimensional feature vector output by the feature encoder is used as the input of the temporal network. The LSTM (Long Short-Term Memory) network has strong learning ability in temporal data. By introducing the concept of "gates" to control the transmission of information, LSTM can capture and remember the dependencies of image feature vectors at long time intervals, so as to more accurately predict the positioning information of the vehicle;
[0121] Reshape the feature vector into a 4×192 matrix, use four LSTM networks, and pass it to a fully connected layer for positioning prediction.
[0122] Step 3: Training of the network model:
[0123] According to the perceived road structure information model, image spatial feature extraction model, and global pose prediction model in Step 2, conduct separate training respectively; divide the image data collected in Step 1 into a training set, a validation set, and a test set according to the ratio of 3:1:1;
[0124] For the DeepLabV3Plus network, define 9 categories, namely road, lane line, road marking, traffic light, traffic sign, pole, sidewalk, building, and tree; use the Labelme software to annotate the original images, and set the learning rate l r to be 0.01, and use the GPU to train it;
[0125] For the Transformer feature encoder, use the pre-trained model for the classification task on the ImageNet dataset as the pre-trained model of this application;
[0126] For the LSTM network, for the true pose l, define the loss function:
[0127] loss(I) = (l - l*) 2 (10)
[0128] In the formula, l is the true pose, and l* is the predicted pose.
[0129] Step 4: Positioning of the vehicle pose, including:
[0130] Use the Extended Kalman Filter to integrate the IMU and the network output positioning results to obtain the vehicle pose state variable as where x i, y i, are the vehicle position and heading angle in the two-dimensional plane respectively.
[0131] The specific content of Step 4 is:
[0132] Let the estimated state X at time t i bei Driven by system noise G i and expressed in equation form as:
[0133] X i = A i,i-1 X i + W i-1 G i-1 (11)
[0134] where A i,i-1 is the transition matrix from time t i-1 to time t i and W i-1 is the system noise drive matrix.
[0135] L i is the GPS data, whose value satisfies a linear relationship, and the measurement equation is:
[0136] Z i = H i L i + V i (12)
[0137] where H i is the measurement matrix and V i is the measurement noise sequence. In this system, we have:
[0138]
[0139] Let Q i be the system noise sequence variance matrix and R i be the measurement noise sequence variance matrix, which is a non - negative definite matrix in this system;
[0140] When initializing the filter, take the first - frame GPS position and coordinates, i.e., the positioning information of the network prediction model, as the initial value of the filter state Then the initial estimated mean - square error matrix is:
[0141]
[0142] When the speed information and the gyroscope angular velocity information are updated, the predicted solution is obtained by dead reckoning, and the system noise variance matrix Q i is:
[0143]
[0144] The transition matrix A i,i-1 from time i - 1 to time i is:
[0145]
[0146] For W i-1, when the vehicle steering angle at the front and rear moments < α, where α is the vehicle steering angle, there is:
[0147]
[0148] When the vehicle pose steering angle at the front and rear moments > α, there is:
[0149]
[0150] Where td is the distance the vehicle advances within the time period, and tr is the angle the vehicle turns within the time period, that is:
[0151] td = v * dt, tr = w * dt (19)
[0152] Furthermore, the predicted mean square error matrix is obtained:
[0153] P i / i-1 = A i,i-1 * P i-1 * A T i,i-1 + W i * Q i-1 * W T i-1 (20)
[0154] When the positioning information output by the network model is updated, the positioning result is fused with the current dead reckoning result; according to the current network model, the measurement noise variance matrix is obtained:
[0155]
[0156] Where ed is the distance estimation noise;
[0157] Calculate the filtering gain K i :
[0158] K i = P i / i-1 * H T i *(H i * P i / i-1 * H T i + R i ) (22)
[0159] Then the final state estimate is:
[0160] Y i = Z i - H i * X i / i-1 (23)
[0161] X i = X i / u-1 + Ki *Y i (24)
[0162] Finally, update the mean squared error matrix P, where E is the 3×3 identity matrix:
[0163] P i =(E - K i *H i )*P i / i-1 (25)
[0164] Final output That is, the vehicle position and heading angle.
[0165] The specific working methods carried out according to the above steps include:
[0166] As described in step 1: Collect data to make a dataset. The vehicle is equipped with a monocular camera and GNSS-RTK. First, manually calibrate the timestamps of the monocular camera and GNSS-RTK to ensure the synchronization of high-precision position information and image data. When the vehicle enters the target section and collects data, the vehicle maintains a normal driving state to restore the real driving state to the greatest extent.
[0167] As described in step 2: Build a network model. Figure 5 For the built network model. For the semantic segmentation network, use the DeepLabV3Plus network with the PSA module added. As Figure 2 shown, input the original road image, select MobileNetV2 as the image feature extraction network, and finally obtain the semantically segmented image. For the Transformer encoder, the structure is as Figure 3 shown. Similar to the Swin Transformer, the Transformer encoder is divided into four stages, and each stage includes a window attention module and a shifted window-based attention module. Drawing on the characteristics of the DAT network and the segmented image, too many deformable attentions greatly increase the computational overhead and timeliness. Therefore, only introduce the deformable attention mechanism in stage 4, and only use it in the shifted window-based attention mechanism, while preserving the global self-attention mechanism inside the window. This enables it to better learn global features and reduce the computational amount and training difficulty as much as possible. After the segmented image enters the Transformer encoder and outputs a 768-dimensional row vector, it is input into the LSTM network, and thus the network model is established.
[0168] As described in step 3: Network model training. For the DeepLabV3Plus network, nine categories are defined, namely road, lane line, road marking, traffic light, traffic sign, pole, sidewalk, building, and tree. Use the Cityscapes software to annotate them, and perform rotation, cropping, and occlusion to increase the data volume as data augmentation. Use the pre-trained model on the Cityscapes dataset for training. For the training of the Transformer encoder, first resize the semantically segmented image to 224×224×3, use the network for classification tasks on the ImageNet dataset as the backbone network, and use its pre-trained model for training, freezing the training weights of the first three stages. For the LSTM, to train this model, we use the Stochastic Gradient Descent (SGD) algorithm with an initial learning rate of 0.001, a weight decay of 0.0005, and a batch size of 8, using the mean squared error as the loss function to be minimized.
[0169] As described in step 4: Location implementation. When the vehicle enters the target urban road section, the vehicle camera captures a structured road image and feeds it into the trained DeepLabV3Plus network. MobileNetV2 is selected as the backbone network for image feature extraction. The image enters the network and performs a series of convolutional operations. The feature maps obtained with four different sampling rates are fused, and then further fused with the shallow features obtained from MobileNetV2. Finally, semantic segmentation is completed through a decoding operation to obtain an image with dynamic objects such as vehicles, pedestrians, and cyclists removed. After the size of the output semantic segmentation image is adjusted to 224×224×3, it enters the Transformer encoder, and the image size is 224×224×3. Similar to the Swin Transformer, the input image first enters the convolutional layer, and the entire image is evenly divided into several small feature maps, specifically 56×56×96. That is, 3136 small feature maps are obtained, and each small feature map is flattened into a one-dimensional vector with a length of 96. Reshape the size and divide the feature map into small windows. Each window size is 7, that is, every 49 small feature maps form a window, and there are a total of 64 windows. Enter stage one. First, each window performs window self-attention calculation. After 64 windows perform self-attention calculation simultaneously, reshape the size to 56×56×96. After the calculation is completed, enter the shifted window, that is, rearrange the form of the divided windows and re-divide the windows to correlate the global image features. After re-dividing the windows, due to the offset window size limit, the number of windows and the window size will change. Reshape the feature map so that the number of windows is still 64, and the windows with changed window sizes are filled with zero vectors and reshaped to 7×7, and then perform self-attention calculation for each window. After the calculation is completed, reshape it to 56×56×96 again, and thus enter stage two. First, perform downsampling operation, sample the 56×56×96 feature map equidistantly up, down, left, and right, reshape it to 28×28×384, and then perform a convolution operation to get 28×28×192, that is, 784 small feature maps, and each feature map is flattened into a one-dimensional vector with a length of 192. Similar to stage one, perform the window division operation, each window size is 7, and there are a total of 16 windows. Each window first performs window self-attention, after reshaping the size, enter the shifted window attention calculation, and after reshaping the size, enter stage three. Enter stage three. First, continue to perform the downsampling operation, continue to sample the 28×28×192 feature map equidistantly. After the sampling is completed, reshape the size to 14×14×384. That is, 196 small feature maps, and each small feature map is flattened into a one-dimensional vector with a length of 384. Divide the windows, the window size is 7, and there are a total of 4 windows. After each window performs the window self-attention mechanism calculation, reshape the size and enter the shifted window attention calculation. After reshaping the size, enter stage four. Perform the downsampling operation, sample the 14×14×384 equidistantly, and after sampling, reshape the size to 7×7×768. There are 49 small feature maps, and each small feature map is flattened into a one-dimensional vector with a length of 768.The window size remains 7. At this time, there is one window. The sliding window attention calculation is no longer performed, and the window self-attention mechanism is no longer executed either. Instead, the window deformable attention mechanism is executed, namely formulas (6), (7), (8), and (9). The final output is 7×7×768. Average pooling is performed and flattened, and finally a one-dimensional image feature vector with a length of 768 is obtained. It is fed into the LSTM network and finally enters the fully connected layer to output the global positioning information. The first positioning information output will be used as the initial value in the multi-sensor stage. The IMU information will be loosely coupled with the network positioning information, and the final trajectory information is output after that.
[0170] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, please refer to the description in the method section.
[0171] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A Transformer-based global positioning method for commercial vehicles on structured roads, characterized in that: The following steps are involved: Step 1: Image acquisition, image processing, and data set preparation: Collect image features of target city roads, including roads, lane lines, road markings, traffic lights, traffic signs, poles, sidewalks, buildings, and trees, and record the location information of each image; perform pixel preprocessing on the images, that is, reshape the original images into 513×513 size, and form the reshaped images into the data set for training and testing networks; Step 2: Network model establishment, including the following three parts: The first part of the network model is a model for perceiving road structure information. The DeepLabV3Plus semantic segmentation network is used to perform semantic segmentation on the resized original image. The defined categories include roads, lane lines, road markings, traffic lights, traffic signs, poles, sidewalks, buildings, and trees. Based on the original DeepLabV3Plus semantic segmentation network model, the attention module is introduced; That is, a PSA module is added after the ASPP module. The PSA module consists of channel attention and spatial attention, which increases the extraction of important information and improves the model's utilization of information. The second part of the network model is an image spatial feature extraction model, which extracts robust position feature information from the image and encodes these features through the Transformer variable self-attention mechanism; When faced with continuous similar scenes, the feature vectors of similar scene images perform cosine calculations on each other. When the cosine value is close to 0.5, the effectiveness of the feature extraction model can be proved; The overall architecture based on the Swin Transformer sliding window is adopted, the last fully connected layer is removed, and a 768-dimensional feature vector is output, which is regarded as the feature vector of the image to be located; The third network model is a global posture prediction model, which uses the 768-dimensional feature vector output by the feature encoder as the input of the temporal network and predicts the vehicle's positioning information through the LSTM long short-term memory network. Step 3: Network model training: According to the key information model of the road detection, the image space feature extraction model and the global posture prediction model in step 2, separate training is performed respectively; the image data collected in step 1 is divided into a training set, a verification set and a test set in a ratio of 3:1:1; Step 4: Positioning of vehicle posture, including: The extended Kalman filter is used to integrate the IMU and network output positioning results. First, the estimated state and system noise equations are obtained, and then the measurement equation is constructed to infer the system noise variance matrix. When the network output positioning result and IMU data are updated, the state of the fusion system is updated by calculating the Kalman filter gain, and finally the fusion positioning result is output to obtain the vehicle posture state variable: where x i, y i, They are the vehicle position and heading angle in the two-dimensional plane respectively.
2. The Transformer-based global positioning method for structured road autonomous driving commercial vehicles according to claim 1, characterized in that: The attention module in step 2 includes two parts: channel attention and spatial attention, where the channel attention weight is calculated as: A ch (X)=F SG [W zθ1 ((σ1(W v (X))×F SM (σ2(W q (X))))] (1) Where σ1 and σ2 represent tensor integer operators, F SM represents the softmax function, W zθ1 represents 1×1 convolution and LN layer, increasing the channel dimension from C / 2 to C, F SG Represents the sigmoid function; The spatial attention weight is calculated as: A sp (X)=F SG [σ3(F SM (σ1(F GP (W q (X))))×σ2(W v (X)))] (2) Where σ1, σ2, σ3 represent tensor integer operators, F SM represents the softmax function, F GP represents global pooling, F SG Represents the sigmoid function, which adopts a serial structure, that is, the feature map X first performs channel attention calculation and then spatial attention calculation.
3. The Transformer-based global positioning method for structured road autonomous driving commercial vehicles according to claim 2, characterized in that: In step 2, Swin Transformer uses the self-attention mechanism, namely: q=xW q ,k=xW k ,v=xWv (3) from=Concat(from (1) ,...,With (m) )IN o (5) Where x is the feature vector matrix obtained by flattening the feature map generated by convolution of the original input image, m is the number of heads of the multi-head attention mechanism, σ(·) is the softmax function, d is the vector length, and z is (m) represents the output matrix of the mth attention head, q (m) , k (m) and v (m) They represent the query matrix, the matrix to be queried and the actual feature matrix generated by the feature graph respectively; W q , W k , W v Generate q (m) , k (m) , v (m) The weight matrix, W o is the multi-attention weight matrix, z is the final output feature vector matrix; The deformable attention mechanism is used to replace the self-attention mechanism to improve the operation rate. The image feature map consists of 2401 tokens, that is, a 49*49 rectangular token array. A set of offsets is learned to determine the tokens with favorable attention calculations between the current token and the remaining tokens of the image, and attention calculations are performed on important points. The offsets are a set of two-dimensional vectors, pointing from the current token to the favorable token. Given an input feature map x∈R generated by an original image with length, width and channels H, W, C H×W×C , generate several position points p∈R HG×WG×2 ; First, give each point position information, that is, {(0,0),...(H G-1 ,W G-1 )}, then according to the grid shape H G ×W G , normalize all coordinates and determine the relative position, that is, (-1,-1) represents the upper left corner, and (+1,+1) represents the lower right corner; Offset: △p=θ offset (q) (6) where q is the query vector of the feature point to which the attention mechanism is currently to be executed, θ offset (·) is a lightweight convolutional neural network. The neural network consists of two convolutional modules with nonlinear activation. The query vector q generated by the feature map is input. The local features are first captured by a 5×5 deep convolution. Then, the GELU activation function and 1×1 convolution are used to obtain the two-dimensional offset, and then the deformable attention mechanism is obtained: In the formula, φ(;) is a bilinear interpolation function, which is used to obtain the feature points important to the current point. The query vector and actual feature vector of important feature points; Perform multi-head attention calculation on q, k, v to get the final result 4. The Transformer-based global positioning method for structured road autonomous driving commercial vehicles according to claim 1, characterized in that: The feature vector is reshaped into a 4×192 matrix, and four LSTM networks are used to pass it to a fully connected layer for positioning prediction.
5. The Transformer-based global positioning method for structured road autonomous driving commercial vehicles according to claim 1, characterized in that: For the DeepLabV3Plus network, 9 categories are defined, namely roads, lane lines, road markings, traffic lights, traffic signs, poles, sidewalks, buildings and trees; the original images are annotated using Labelme software, and the learning rate is set to l r is 0.01 and trained using GPU; For the Transformer feature encoder, a pre-trained model for classification tasks on the ImageNet dataset is used as the pre-trained model for this application; For the LSTM network, the true pose l, the loss function is defined as: loss(I)=(l-l*) 2 (10) Where l is the true pose and l* is the predicted pose.
6. The Transformer-based global positioning method for structured road autonomous driving commercial vehicles according to claim 1, characterized in that: The step 4 is specifically as follows: Assume t i The estimated state X at time i Affected by system noise G i Drive, expressed in equation form: X i =A i,i-1 X i +W i-1 G i-1 (11) Among them A i,i-1 t i-1 Time to t i The transfer matrix at time, W i-1 is the system noise driving matrix; L i is GPS data, its value satisfies the linear relationship, and the measurement equation is: Z i =H i L i +V i (12) Among them, H i is the measurement matrix, V i To measure the noise sequence, in this system we have: Let Q i is the system noise sequence variance matrix, R i is the variance matrix of the measurement noise sequence, which is a non-negative definite matrix in this system; When the filter is initialized, the first frame of GPS position and coordinates, that is, the positioning information of the network prediction model, is taken as the initial value of the filter state When the velocity information and gyroscope angular velocity information are updated, the prediction solution is obtained by dead reckoning, and the system noise variance matrix Q i for: The transfer matrix A from time i-1 to time i i,i-1 for: For W i-1 , when the vehicle turning angle <α at the previous and next moments, α is the vehicle turning angle, we have: When the vehicle posture angle > α at the current and next moments, we have: Where td is the distance the vehicle travels in the time period, and tr is the angle the vehicle turns in the time period, that is: td=v*dt,tr=w*dt (18) Then we get the prediction mean square error matrix: P i / i-1 =A i,i-1 *P i-1 *A T i,i-1 +W i *Q i-1 *W T i-1 (19) When the network model outputs the updated positioning information, the positioning result is fused with the current dead reckoning result; the measurement noise variance matrix is obtained according to the current network model: Where ed is the distance estimation noise; Calculate the filter gain K i : K i =P i / i-1 *H T i *(H i *P i / i-1 *H T i +R i ) (21) The final state is estimated to be: Y i =Z i -H i *X i / i-1 (22) X i =X i / i-1 +K i *Y i (23) Finally, update the mean square error matrix P, E is a 3×3 identity matrix: P i =(E-K i *H i )*P i / i-1 (24) Final Output That is, the vehicle position and heading angle.
Citation Information
Patent Citations
Automatic driving method based on convolutional neural network and attention mechanism
CN117115770A
Vehicle track prediction method independent of map information
CN117493424A