Lane positioning method and device, electronic equipment and storage medium
By acquiring and fusing image features and relative pose features and inputting lane serial number identification model, the accuracy of existing lane positioning methods under occlusion and environmental interference is solved, and the accuracy of positioning and anti-interference ability are improved.
Patent Information
- Application Number
- CN202311809171.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing lane positioning methods are inaccurate or unable to output due to occlusion or environmental interference during driving. The high-precision map production cost is high, the iteration period is long, and the GPS dependence is strong, and they are easily disturbed.
By acquiring image features and relative pose features at multiple moments, splicing them into fusion features, and entering the lane serial number identification model, the lane serial number and confidence of the current lane in which the bicycle is located is obtained.
It improves the accuracy and anti-interference ability of lane-level positioning, is suitable for complex and changeable urban open road environments, and reduces dependence on high-precision maps and GPS.
Smart Images

Figure CN120220096A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of vehicle engineering, and particularly to a method, device, electronic device and storage medium for lane positioning. Background Art
[0002] Autonomous vehicles use artificial intelligence, signal processing algorithms, etc. to convert the multi-sensor results of devices such as lidar, global positioning system, and cameras into the required planning control information, enabling the vehicle to automatically perform complex driving behaviors without active human control. And for the vehicle to follow a predetermined trajectory during autonomous driving, accurate lane-level positioning is crucial.
[0003] Current lane-level positioning methods usually use the centimeter-level longitude and latitude results of GPS positioning devices and the image information output by cameras. By making judgments based on the lanes, railings, etc. included in the image and the high-precision map information obtained through longitude and latitude, the lane number where the vehicle is located is determined. However, during driving, due to occlusion by the vehicle or other obstacles, the lane lines included in the road image are unclear or not distinct, resulting in inaccurate or directly unobtainable lane numbers determined by such methods. At the same time, the production cost of high-precision maps is high and the iteration cycle is long. Also, devices such as GPS rely on the status of satellites. When there is occlusion above the vehicle (such as under an overpass) or there are metal objects nearby, the longitude and latitude results will be abnormal, which is not conducive to simple and accurate lane positioning tasks. Moreover, obtaining the vehicle body motion information through IMU, GPS, and vehicle control information requires constructing a relatively complex physical model and adjusting a large number of threshold parameters during use, and it cannot adapt to all real scenarios. Summary of the Invention
[0004] In view of this, the present application provides a method, device, electronic device and storage medium for lane positioning, aiming to improve the accuracy and anti-interference ability of lane-level positioning.
[0005] A method for lane positioning provided by the present application, the method includes:
[0006] Obtain image features and relative pose features at multiple moments;
[0007] Stitch the image features and relative pose features at the same moment to obtain a fused feature;
[0008] By inputting the fused features at multiple consecutive moments into a lane number recognition model, obtain the lane number of the current lane where the vehicle is located and the confidence of the lane number.
[0009] Optionally, obtaining image features at multiple moments includes:
[0010] Collect image data in front of the host vehicle at a preset time interval;
[0011] Obtain multiple target frame image data by performing downsampling and cropping on the collected image data;
[0012] Extract features from the multiple target frame image data to obtain image features of the target frame image data corresponding to multiple moments respectively.
[0013] Optionally, obtain relative pose features at multiple moments, including:
[0014] Obtain the IMU data and wheel speed data of the host vehicle at a preset time interval;
[0015] Input the IMU data and wheel speed data at two adjacent moments within a preset time interval into a relative pose recognition model for processing to obtain relative pose features of the host vehicle at multiple moments.
[0016] Optionally, the training process of the relative pose recognition model includes:
[0017] Collect the IMU data and wheel speed data at two adjacent moments of the first preset quantity as sample data for training the initial relative pose recognition model;
[0018] Determine the relative pose results of each sample data through multi-sensor and Kalman filter fusion;
[0019] Use the relative pose results of each sample data obtained by solving as the sample labels of the corresponding sample data to label the corresponding sample data, and obtain the first target sample data;
[0020] Train the initial relative pose recognition model with the first target sample data of the first preset quantity;
[0021] When the value of the loss function of the model satisfies the first set condition, obtain the relative pose recognition model.
[0022] Optionally, the training process of the lane number recognition model includes:
[0023] Obtain the second preset quantity of fusion features as sample data for training the initial lane number recognition model;
[0024] Obtain the true lane number labels according to the image data corresponding to the image features in the sample data, and label the corresponding sample data with the true lane number labels to obtain the second target sample data;
[0025] Perform backpropagation training on the initial lane number recognition model by inputting the second target sample data at multiple consecutive moments;
[0026] When the value of the loss function of the model satisfies the second set condition, the lane number recognition model is obtained.
[0027] Optionally, obtaining the lane number of the current lane where the vehicle is located and the confidence of the lane number by inputting the fusion features of multiple consecutive moments into the lane number recognition model includes:
[0028] By inputting the fusion features of multiple moments into the lane number recognition model, the lane numbers of the current lane where the vehicle is located sorted from left to right and the confidence of the lane number are obtained, and the lane numbers of the current lane where the vehicle is located sorted from right to left and the confidence of the lane number are obtained.
[0029] Regarding the prior art, the present application has the following advantages:
[0030] A lane positioning method provided by the present application obtains image data, vehicle speed, and vehicle acceleration and angular velocity information by acquiring sensors (such as wheel encoders, IMUs, and cameras) bound to the vehicle. Image features are obtained based on the image data, and the relative pose features of the vehicle itself are obtained based on the IMU data and wheel speed data. By fusing the image features and relative pose features at the same moment, fusion features are obtained. By inputting the fusion features of multiple consecutive moments into the lane number recognition model for processing, the lane number of the current lane where the vehicle is located and the confidence of the lane number are obtained. Since the IMU and wheel speedometer are sensors that are not easily disturbed by environmental changes (such as light changes, rain, snow, fog, dust, etc.), and the picture information of the vision camera is rich, by using the sensor information of multiple moments at the same time, compared with using only the sensor data of a single moment for lane positioning, the accuracy and anti-interference ability have been significantly improved, and at the same time, the robustness is better, which is suitable for the urban open road environment with complex and changeable road conditions.
[0031] The second aspect of the present application provides a lane positioning device, aiming to improve the accuracy and anti-interference ability of lane-level positioning.
[0032] A lane positioning device provided by the present application, the device includes:
[0033] A feature acquisition module for acquiring image features and relative pose features of multiple moments;
[0034] A feature fusion module for splicing the image features and relative pose features at the same moment to obtain fusion features;
[0035] A lane number recognition module for obtaining the lane number of the current lane where the vehicle is located and the confidence of the lane number by inputting the fusion features of multiple consecutive moments into the lane number recognition model.
[0036] A third aspect of the present application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0037] The memory is used to store computer programs;
[0038] The processor is used to implement the steps in the method for lane positioning described in the first aspect when executing the program stored on the memory.
[0039] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for lane positioning in the above-mentioned first aspect are implemented. Description of the Drawings
[0040] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0041] Figure 1 is a flowchart of a method for lane positioning provided by an embodiment of the present application;
[0042] Figure 2 is a schematic structural diagram of a deep learning neural network for extracting image features in a method for lane positioning provided by an embodiment of the present application;
[0043] Figure 3 is a schematic structural diagram of a deep learning neural network for predicting lane numbers in a method for lane positioning provided by an embodiment of the present application;
[0044] Figure 4 is a schematic diagram of the overall technical route in a method for lane positioning provided by an embodiment of the present application;
[0045] Figure 5 is a schematic diagram of a device for lane positioning provided by an embodiment of the present application. Detailed Embodiments
[0046] Hereinafter, the exemplary embodiments of the present application will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be completely conveyed to those skilled in the art.
[0047] Figure 1It is a flowchart of a lane positioning method provided by an embodiment of the present application. As Figure 1 shown, the method includes:
[0048] Step S101: Obtain image features and relative pose features at multiple moments;
[0049] Step S102: Concatenate the image features and relative pose features at the same moment to obtain a fused feature;
[0050] Step S103: Input the fused features at multiple consecutive moments into a lane number recognition model to obtain the lane number of the current lane where the vehicle is located and the confidence of the lane number.
[0051] In this embodiment, the image data in front of the vehicle collected is processed to obtain the image features corresponding to the image data at each moment. At the same time, the IMU data and wheel speed data obtained from the vehicle's wheel encoder and IMU sensor (IMU (Inertial Measurement Unit), that is, the inertial measurement unit) are processed to obtain the relative pose features of the vehicle relative to the previous moment at each moment. Thus, the image features and relative pose features at multiple moments are obtained.
[0052] Exemplarily, the image data in front of the vehicle collected at moment a1 is processed to obtain the image features corresponding to the image data of the vehicle at moment a1; the image data in front of the vehicle collected at moment a2, which is the previous moment of a1, is processed to obtain the image features corresponding to the image data of the vehicle at moment a2; the image data in front of the vehicle collected at moment a3, which is the previous moment of a2, is processed to obtain the image features corresponding to the image data of the vehicle at moment a3, and so on.
[0053] The IMU data and wheel speed data obtained from the vehicle's wheel encoder and IMU sensor at moment a1, and the IMU data and wheel speed data obtained from the vehicle's wheel encoder and IMU sensor at moment a2, which is the previous moment of a1, are processed simultaneously to obtain the relative pose feature of the vehicle at moment a1 relative to moment a2, that is, the relative pose feature of the vehicle at moment a1; the IMU data and wheel speed data obtained from the vehicle's wheel encoder and IMU sensor at moment a2, and the IMU data and wheel speed data obtained from the vehicle's wheel encoder and IMU sensor at moment a3, which is the previous moment of a2, are processed simultaneously to obtain the relative pose feature of the vehicle at moment a3 relative to moment a2, that is, the relative pose feature of the vehicle at moment a2.
[0054] In this embodiment, the fused feature is obtained by concatenating the image features and relative pose features at the same moment.
[0055] Exemplarily, continuing with the above example, the image features of the host vehicle at time a1 are combined with the relative pose features of the host vehicle at time a1 to obtain the fusion features of the host vehicle at time a1; the image features of the host vehicle at time a2 are combined with the relative pose features of the host vehicle at time a2 to obtain the fusion features of the host vehicle at time a2.
[0056] In this embodiment, after obtaining the fusion features at multiple moments, the fusion features at multiple consecutive moments are input into the lane number recognition model to predict the lane number and the corresponding confidence level of the lane where the host vehicle is currently located. This result is the recognition result with the highest confidence directly output by the lane number recognition model. Among them, the fusion features at multiple consecutive moments at least include the fusion features at the current moment, and may also include the fusion features at multiple consecutive moments before the current moment. For example, if the number of fusion features at multiple consecutive moments input into the lane number recognition model is x, then the x fusion features at consecutive moments include the fusion features at the current moment and the fusion features at x - 1 historical moments consecutive to the current moment.
[0057] Among them, the confidence level of the lane number of the lane where the host vehicle is currently located represents the probability of the lane number result output by the lane number recognition model. Exemplarily, the host vehicle numbers the lanes on the road where it is currently traveling, such as numbering them as lane 1, lane 2, and lane 3 from left to right; by inputting the fusion features at multiple consecutive moments into the lane number recognition model, it is obtained that the lane number of the lane where the host vehicle is currently located is lane 1, and at the same time, the confidence level that the host vehicle is in lane 1 is output as 0.9, that is, the probability that the host vehicle is currently in lane 1 is 90%.
[0058] In this embodiment, in another implementation manner of this application, the fusion features at multiple consecutive moments are input into the lane number recognition model for processing to predict the possible lane numbers and the corresponding confidence levels of the lanes where the host vehicle may be located on the road it is currently traveling. In subsequent applications, based on all the obtained results, the lane number corresponding to the result with the highest confidence can be selected, and at this time, it is determined that the host vehicle is currently in the lane corresponding to this lane number. For example, the road where the host vehicle is currently traveling includes three lanes with lane numbers 1, 2, and 3 respectively. By inputting the fusion features at multiple consecutive moments into the lane number recognition model, three prediction results are obtained, including that the host vehicle is currently in the lane corresponding to lane number 1, and the corresponding confidence level is 0.7; the host vehicle is currently in the lane corresponding to lane number 2, and the corresponding confidence level is 0.1; the host vehicle is currently in the lane corresponding to lane number 3, and the corresponding confidence level is 0.2. Among them, the confidence level 0.7 is the highest, so it is determined through screening that the host vehicle is currently in the lane corresponding to lane number 1.
[0059] A lane positioning method provided by the present application obtains image data, vehicle speed, and vehicle acceleration and angular velocity information by acquiring sensors (such as wheel encoders, IMUs, and cameras) bound to the vehicle. Image features are obtained based on the image data acquired by the camera, and relative pose features of the host vehicle are obtained from the pose data acquired by the IMU and the wheel speedometer. The fused features are obtained by fusing the image features and the relative pose features at the same moment. By inputting the fused features at multiple consecutive moments into a lane number recognition model for processing, the lane number of the lane where the host vehicle is currently located and the confidence level of the lane number are obtained. Since the IMU and the wheel speedometer are sensors that are not easily interfered by environmental changes (such as changes in light, rain, snow, fog, sand, etc.), and the image information of the vision camera is rich, the accuracy and anti-interference ability of lane positioning are significantly improved by using the sensor information at multiple moments compared to using only the sensor data at a single moment. At the same time, it has better robustness and is applicable to the complex and changeable urban open road environment.
[0060] At the same time, the combination of image features and relative pose features will make the applicable scenarios of lane-level positioning more extensive, thus achieving further optimization of lane-level positioning.
[0061] For example, when the vehicle reaches a traffic light at the current moment and the image data obtained at this time does not involve relevant lane information, the image features obtained by processing the image data at the current moment will not involve implicit lane information data either. Or there is a large amount of vehicle flow passing in front of the vehicle at the current moment, and the image data obtained by the vehicle at the current moment does not involve relevant lane information. Or the camera of the vehicle is blocked, resulting in the inability to obtain the image data at the current moment. At this time, since the input to the lane number recognition model is the fusion feature obtained by fusing the image feature and the relative pose feature, even if the image feature at the current moment does not involve implicit lane information data, the lane number recognition model can analyze this fusion feature obtained by fusing the image feature and the relative pose feature to achieve the lane positioning of the vehicle. For example, the image feature at the current moment does not involve implicit lane information, and the lane number recognition model cannot perform lane positioning on the vehicle by processing the image feature at the current moment. However, since the fusion feature also includes the relative pose feature between the current moment and the previous moment of the current moment, at this time, the lane number recognition model can process the relative pose feature between the current moment and the previous moment of the current moment input to it and the lane positioning result of the vehicle at the previous moment of the current moment to obtain the lane positioning result of the vehicle at the current moment and the corresponding confidence level. For example, although the image feature in the fusion feature at the current moment input to the lane number recognition model does not involve implicit lane information data, due to the lane positioning result of the vehicle at the previous moment of the current moment, and the relative pose change amount between the current moment and the previous moment of the current moment of the vehicle is very small at this time, the lane number recognition model can output that the lane positioning result of the vehicle at the current moment is the lane positioning result of the vehicle at the previous moment of the current moment and the corresponding confidence level based on the lane positioning result of the vehicle at the previous moment of the current moment and the relative pose feature between the current moment and the previous moment of the current moment of the vehicle.
[0062] In this embodiment, the lane positioning result at the previous moment of the current moment can be obtained by processing the fusion feature at the previous moment of the current moment through a lane number recognition model. Since the input to the lane number recognition model is the fusion features of multiple consecutive moments, the previous moment of the current moment mentioned in this application is any one moment or any multiple moments other than the current moment among the multiple consecutive moments. In this way, even when the image features in the fusion feature at the previous moment adjacent to the current moment do not involve implicit lane information data, the lane number recognition model can still obtain the lane positioning result at the previous moment of the current moment based on the fusion features at other moments other than the current moment among the multiple consecutive moments. Here, the image features involving implicit lane information data mean that the image features include all the feature information of the image data. For example, the lane widths of each lane on the road where the vehicle is currently traveling, the line types and colors of each lane line, etc. This application does not need to perform specific recognition processing on them, that is, it does not need to specifically recognize the lane widths of each lane and the line types and colors of each lane line, but only needs to convert the image data into an abstract array form for representation and put it into the image features. Subsequently, the lane number recognition model can predict the lane number where the vehicle is currently located and the confidence level of the lane number based on the image features represented in this abstract array form.
[0063] In this application, obtaining the image features of multiple moments includes: collecting image data of the front of the vehicle at a preset time interval; obtaining multiple target frame image data by performing downsampling and cropping processing on the collected image data; and obtaining the image features of the target frame image data corresponding to multiple moments respectively by performing feature extraction on the multiple target frame image data.
[0064] In this embodiment, image data of the front of the vehicle is collected at a preset time interval by a camera sensor. The camera sensor can be a front-view camera or a rear-view camera, and the installation position of the camera is preferably such that it can clearly cover the lane lines (the coverage degree is similar to the front scene visible to the driver's eyes). The camera can be selected from a 60° wide-angle to 150° wide-angle camera. The collected image data at the preset time interval is subjected to downsampling and cropping processing. Since the lane information data implicit in the image data used for lane positioning of the vehicle itself is involved, in order to ensure that the image data only contains lane information data as much as possible, the image data is subjected to downsampling and cropping processing to process the image data into a set size while increasing the proportion of lane information data in the image data. Therefore, by performing downsampling and cropping processing on the collected image data, respective target frame image data corresponding to the respective image data are obtained. Among them, the downsampling method can be selected from methods such as bilinear interpolation and nearest neighbor interpolation, which are not specifically limited herein. Among them, the respective target frame image data are in the form of a pixel matrix in the range of 0-255.
[0065] By sending the target frame image data to a feature extraction model, an array of a specific size is obtained through processing. Among them, the array of a specific size is preferably an array composed of 512 floating-point numbers. This is because 512 is a multiple of 2, and the computing performance is better for a computer. Moreover, on the one hand, the length of this array can avoid, for example, setting it to a length of 1024 which is too large, resulting in an increase in the number of model parameters and the amount of calculation. On the other hand, 512 can ensure sufficient description of abstract visual feature information.
[0066] Specifically, a deep learning neural network is set in the feature extraction module. By inputting the target frame image data into the deep learning neural network in the feature extraction module for processing, an array of a specific size is obtained, and this array of a specific size is the image feature corresponding to the target frame image data. An optional deep learning neural network can be a series such as ResNet (Residual Neural Network) and VGGNet (Visual Geometry Group Network). It will extract the information abstracted through multiple layers of functions from the target frame image data and save this information through an array.
[0067] As Figure 2 shown, the figure shows the structure of the deep learning neural network for extracting the image features in the target frame image data. From left to right, they are the overall structure of the feature extraction module, the layer structure, and the block structure. Among them, the layer is composed of a combination of several block layers with the same structure. In this application, the number of block blocks in layer1, 2, 3, and 4 is 3, 3, 9, and 3 in sequence.
[0068] In this application, obtaining relative pose features at multiple moments includes: obtaining IMU data and wheel speed data of the host vehicle at a preset time interval; by inputting the IMU data and wheel speed data of two adjacent moments within the preset time interval into a relative pose recognition model for processing, relative pose features of the host vehicle at multiple moments are obtained.
[0069] In this embodiment, compared with the prior art of obtaining vehicle body motion information by calling a large amount of data information (such as IMU, GPS, and vehicle control information, etc.) and constructing a complex physical model, and at the same time, a large number of threshold parameters need to be adjusted when used, resulting in very complex lane positioning of the host vehicle and being easily affected by the environment (for example, GPS will be affected by a poor external satellite signal environment, such as being blocked by trees or under overpasses, resulting in a decrease in the accuracy of the host vehicle's lane positioning). However, in this application, the data is processed through a deep learning model to achieve lane positioning of the host vehicle. The processing process will be simpler. At the same time, since relative pose information is introduced on the basis of image data, the lane positioning of the host vehicle is less affected by the environment (because GPS data is no longer needed and will not be affected by a poor external satellite signal environment, such as being blocked by trees or under overpasses, resulting in a decrease in the accuracy of the host vehicle's lane positioning), and the applicable scenarios are also more extensive (for example, even when the image features of the current moment cannot be obtained, the lane positioning of the host vehicle at the current moment can be performed through the lane positioning result of the previous moment and the relative pose information of the current moment relative to the previous moment).
[0070] In this embodiment, the common method for obtaining the vehicle motion attitude change by using the angular velocity obtained by the IMU, the acceleration, and the vehicle speed obtained by the wheel speedometer is generally to construct a complex equation (such as Kalman filtering), which requires knowing the noise of various sensors in advance and modeling it with clear physical meaning. This will make the relative pose information of the host vehicle more accurate, and at the same time, this processing process will be very complex. When determining which lane the host vehicle is currently in based on the image features (i.e., a 512-bit floating-point number array) and relative pose information of the host vehicle at two moments in this application, the relative pose information of the host vehicle required does not need to be particularly accurate. Determining the relative pose information of the host vehicle through the existing extremely complex implementation methods will cause waste of computing resources and reduce the processing efficiency. Therefore, this application uses an LSTM deep learning model to learn the motion characteristics of the vehicle. By inputting the IMU data and vehicle speed data of the host vehicle at two moments into the deep learning model, a relative pose feature with not particularly high accuracy is predicted through the deep learning model, and at the same time, it does not require prior knowledge of a large number of sensor noise models as in the previous implementation methods for determining the relative pose information of the host vehicle. This can effectively avoid waste of computing resources and improve the processing efficiency at the same time.
[0071] Specifically, by sending the IMU data and wheel speed data at two adjacent moments with a preset time interval to the relative pose recognition module for processing, an array with a specific length is obtained. Among them, the array with a specific length is a 6-digit array, including the longitudinal displacement, lateral displacement, vertical displacement of the host vehicle at two adjacent moments, and yaw, pitch, and roll. A deep learning neural network is set in the relative pose recognition module. By inputting the IMU data and wheel speed data at two adjacent moments with a preset time interval into the deep learning neural network in the relative pose recognition module for processing, an array with a specific length is obtained, and this array with a specific length is the relative pose feature of the latter moment relative to the former moment among these two adjacent moments. An optional deep learning neural network can be LSTM (Long Short-Term Memory).
[0072] In this embodiment, when the image feature is a 512-bit floating-point number array and the relative pose feature is an array with a length of 6, splicing the image feature and the relative pose feature at the same moment to obtain a fused feature means splicing the 512-bit floating-point number array and the array with a length of 6 at the same moment into an array with a length of 518 bits, and this 518-bit array is the fused feature. By inputting the fused features of multiple consecutive moments into the lane number recognition model for processing, the lane number recognition model will output the lane number of the lane where the host vehicle is located at the current moment and the prediction result of the confidence level of this lane number.
[0073] In this application, the training process of the relative pose recognition model includes: collecting IMU data and wheel speed data at two adjacent moments with a first preset quantity as sample data for training the initial relative pose recognition model; determining the relative pose results of each sample data through multi-sensor and Kalman filter fusion; using the obtained relative pose results of each sample data as the sample labels for the corresponding sample data to perform annotation on the corresponding sample data, obtaining first target sample data; training the initial relative pose recognition model with the first preset quantity of first target sample data; and obtaining the relative pose recognition model when the value of the loss function of the model satisfies a first set condition.
[0074] In this embodiment, IMU data and wheel speed data at two adjacent moments with a preset time interval are collected as sample data for training the initial relative pose recognition model, and at the same time, the quantity of the sample data needs to be the number of the first preset quantity. Among them, the first preset quantity can be valued according to the actual application scenario, and no specific limitation is made here, but the quantity needs to be as large as possible to ensure that a qualified relative pose recognition model is trained.
[0075] After obtaining the first preset quantity of sample data, the relative pose results of each sample data are determined through multi-sensor and Kalman filter fusion, that is, the relative pose change results of the vehicle in the sample data at two adjacent moments are obtained through multi-sensor and Kalman filter fusion processing, including the longitudinal displacement, lateral displacement, vertical displacement of the vehicle between the two adjacent moments, as well as yaw (yaw angle), pitch (pitch angle), and roll (roll angle).
[0076] Using the obtained relative pose results as sample labels to perform annotation on the corresponding sample data, first target sample data is obtained, where the first target sample data is the sample data with sample label annotation.
[0077] The initial relative pose recognition model is trained with the first preset quantity of first target sample data until the value of the loss function of the initial relative pose recognition model satisfies the first set condition, and it is determined that the initial relative pose recognition model is trained qualified and can be used to process the IMU data and wheel speed data at two adjacent moments to obtain the relative pose characteristics of the vehicle at the two adjacent moments. Among them, when the value of the loss function of the initial relative pose recognition model is lower than the first set threshold, it is determined that the value of the loss function of the initial relative pose recognition model satisfies the first set condition, and the value of the first set threshold can be set according to the actual application scenario, and no specific limitation is made here.
[0078] In this application, the training process of the lane number identification model includes: obtaining a second preset number of fused features as sample data for training the initial lane number identification model; obtaining the true value label of the lane number according to the image data corresponding to the image features in the sample data, and labeling the corresponding sample data with the true value label of the lane number to obtain the second target sample data; performing backpropagation training by inputting the second target sample data at multiple consecutive time instants into the initial lane number identification model; and obtaining the lane number identification model when the value of the loss function of the model satisfies the second set condition.
[0079] In this embodiment, the sample data used for training the initial lane number identification model is the fused feature, and the fused feature is a feature obtained by fusing the image feature and the relative pose feature. Obtain a second preset number of fused features as sample data for training the initial lane number identification model. Among them, the second preset number can be set according to the actual application scenario and is not specifically limited here, but the number needs to be as large as possible to ensure that a qualified lane number identification model is trained. Since the image data corresponding to the image features in the sample data is real image,, by processing the image data corresponding to the image features in the sample data (for example, through manual observation and processing to obtain the lane number of the lane where the vehicle is currently located and manually labeling the sample data), the true value label of the lane number of the sample data will be obtained, and the sample data will be labeled with the true value label of the lane number to obtain the second target sample data labeled with the true value label of the lane number.
[0080] After obtaining the second preset number of second target sample data, perform backpropagation training by inputting the second target sample data at multiple consecutive time instants into the initial lane number identification model. When the value of the loss function of the initial lane number identification model satisfies the second set condition, it is determined that the initial lane number identification model is trained qualified, and the lane number identification model is obtained, which can be used to process the fused features at multiple consecutive time instants to obtain the lane number of the lane where the vehicle is currently located and the confidence of the lane number. Among them, when the value of the loss function of the initial lane number identification model is lower than the second set threshold, it is determined that the value of the loss function of the initial lane number identification model satisfies the second set condition. The value of the second set threshold can be set according to the actual application scenario and is not specifically limited here.
[0081] As Figure 3 shown, the figure shows the LSTM structure diagram of the lane number identification module. The information transfer of the lane number in the time dimension is realized by the flow of the memory state C and the intermediate state h in the network. Among them, the C state contains some information related to the lane number for a long time, and the intermediate state information h(t) at the current moment can obtain the result of the lane number output after passing through a fully connected layer.
[0082] In this embodiment, if Figure 4 As shown, the figure shows a schematic diagram of the overall technical route of the present application, and the image data, IMU data and wheel speed data at each moment are obtained through the data acquisition module. The image data is downsampled and cropped by the image preprocessing module to obtain the target frame image data, and the target frame image data is sent to the feature extraction module for feature extraction to obtain the image features corresponding to the target frame image data, and the IMU data and wheel speed data are sent to the relative posture recognition module for processing to obtain the relative posture features, and the image features and relative posture features at the same moment are spliced to obtain the fusion features. By sending the fusion features of multiple moments to the lane number recognition module for processing, the lane number and its probability from left to right and from right to left of the lane where the vehicle is located are obtained.
[0083] In the present application, the method of inputting fused features of multiple consecutive moments into a lane number recognition model to obtain the lane number of the lane in which the vehicle is currently located and the confidence of the lane number includes: inputting fused features of multiple moments into a lane number recognition model to obtain the lane number of the lane in which the vehicle is currently located sorted from left to right and the confidence of the lane number, and obtaining the lane number of the lane in which the vehicle is currently located sorted from right to left and the confidence of the lane number.
[0084] In this embodiment, in order to facilitate the control system of the vehicle to perform corresponding control actions based on the output lane number and the confidence level of the lane number, the final output result includes the lane number and the corresponding confidence level of the lane where the vehicle is currently located, sorted from left to right, and the lane number and the corresponding confidence level of the lane where the vehicle is currently located, sorted from right to left.
[0085] The present application provides a method for lane positioning, which does not require complex motion modeling of IMU and wheel speed, and sets a large number of threshold parameters that need to be adjusted. GPS is not used in real-time vehicle-side user applications, and will not be affected by the poor external satellite signal environment (such as tree cover, under the viaduct) to cause serious deviations in the relative posture estimation of the vehicle. At the same time, no map prior information is required, and the cost of the sensor is low, which is suitable for large-scale mass-produced cars. In addition, the later iteration update is convenient. It only needs to collect more relevant data on lane-level positioning errors of user vehicles when they are in use and bring them into the lane number recognition module for neural network training. No other modules need to be modified. Since the vehicle movement is continuous, by referring to some vehicle information at historical moments, the lane positioning output can be made smoother, and the accuracy and recall rate will also increase (for example, the lane number output of the vehicle should be consistent with the historical lane number when it is in a very short displacement or parking).
[0086] The embodiment of the present application also provides a lane positioning device 500, such asFigure 5 As shown in the figure, the device 500 includes:
[0087] A feature acquisition module 501, configured to acquire image features and relative pose features at multiple moments;
[0088] A feature fusion module 502, configured to splice the image features and relative pose features at the same moment to obtain fused features;
[0089] A lane number recognition module 503, configured to input the fused features at multiple consecutive moments into a lane number recognition model to obtain the lane number of the lane where the vehicle is currently located and the confidence level of the lane number.
[0090] Optionally, the feature acquisition module includes:
[0091] A first data acquisition module, configured to acquire image data in front of the vehicle at a preset time interval;
[0092] An image preprocessing module, configured to obtain multiple target frame image data by performing downsampling and cropping processing on the acquired image data;
[0093] A feature extraction module, configured to extract features from the multiple target frame image data to obtain image features of the target frame image data corresponding to multiple moments respectively.
[0094] Optionally, the feature acquisition module includes:
[0095] A second data acquisition module, configured to acquire IMU data and wheel speed data of the vehicle at a preset time interval;
[0096] A relative pose recognition module, configured to input the IMU data and wheel speed data at two adjacent moments within a preset time interval into a relative pose recognition model for processing to obtain relative pose features of the vehicle at multiple moments.
[0097] Optionally, the training process of the relative pose recognition model in the relative pose recognition module includes: collecting IMU data and wheel speed data at two adjacent moments in a first preset quantity as sample data for training an initial relative pose recognition model; determining the relative pose results of each sample data through multi-sensor and Kalman filter fusion; using the obtained relative pose results of each sample data as sample labels for the corresponding sample data to perform annotation on the corresponding sample data to obtain first target sample data; training the initial relative pose recognition model with the first preset quantity of first target sample data; and obtaining the relative pose recognition model when the value of the loss function of the model satisfies a first set condition.
[0098] Optionally, the training process of the lane number identification model in the lane number identification module includes: obtaining a second preset number of fusion features as sample data for training the initial lane number identification model; obtaining a true label of the lane number according to the image data corresponding to the image features in the sample data, and labeling the corresponding sample data with the true label of the lane number to obtain second target sample data; performing backpropagation training by inputting the second target sample data at multiple consecutive moments into the initial lane number identification model; and obtaining the lane number identification model when the value of the loss function of the model satisfies a second set condition.
[0099] Optionally, the results output by the lane number identification model in the lane number identification module include the lane number and confidence level sorted from left to right of the lane where the vehicle is currently located, and the lane number and confidence level sorted from right to left of the lane where the vehicle is currently located.
[0100] The embodiment of the present application further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used to store a computer program; when the processor executes the program stored on the memory, it implements the steps in the method for lane positioning.
[0101] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method for lane positioning.
[0102] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0103] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the related parts, refer to the partial description of the method embodiment.
[0104] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A method for lane positioning, characterized in that, The method includes: Obtaining image features and relative pose features at multiple moments; Stitching the image features and relative pose features at the same moment to obtain fused features; Inputting the fused features at multiple consecutive moments into a lane number recognition model to obtain the lane number of the vehicle's current position and the confidence level of the lane number.
2. The method according to claim 1, wherein Obtaining image features at multiple moments, including: Collecting image data in front of the vehicle at a preset time interval; Obtaining multiple target frame image data through downsampling and cropping the collected image data; Obtaining image features of the target frame image data corresponding to multiple moments through feature extraction of the multiple target frame image data.
3. The method according to claim 1, wherein Obtaining relative pose features at multiple moments, including: Obtaining the IMU data and wheel speed data of the vehicle at a preset time interval; Inputting the IMU data and wheel speed data at two adjacent moments within a preset time interval into a relative pose recognition model for processing to obtain the relative pose features of the vehicle at multiple moments.
4. The method according to claim 3, characterized in that, The training process of the relative pose recognition model includes: Collecting the IMU data and wheel speed data at two adjacent moments of a first preset quantity as sample data for training an initial relative pose recognition model; Determining the relative pose results of each sample data through multi-sensor and Kalman filter fusion; Using the obtained relative pose results of each sample data as the sample labels of the corresponding sample data to label the corresponding sample data, obtaining first target sample data; Training the initial relative pose recognition model with the first preset quantity of first target sample data; Obtaining the relative pose recognition model when the value of the loss function of the model satisfies a first set condition.
5. The method according to claim 1, wherein The training process of the lane number recognition model includes: Obtaining a second preset quantity of fused features as sample data for training an initial lane number recognition model; Obtaining the true lane number label based on the image data corresponding to the image features in the sample data, and labeling the corresponding sample data with the true lane number label to obtain second target sample data; Performing backpropagation training on the initial lane number recognition model by inputting the second target sample data at multiple consecutive moments; Obtaining the lane number recognition model when the value of the loss function of the model satisfies a second set condition.
6. The method according to claim 1, characterized in that, The step of inputting the fused features at multiple consecutive moments into a lane number recognition model to obtain the lane number of the vehicle's current position and the confidence level of the lane number includes: Inputting the fused features at multiple moments into a lane number recognition model to obtain the lane numbers sorted from left to right of the vehicle's current position and the confidence levels of the lane numbers, and obtaining the lane numbers sorted from right to left of the vehicle's current position and the confidence levels of the lane numbers.
7. A device for lane positioning, characterized in that, The device includes: A feature acquisition module for obtaining image features and relative pose features at multiple moments; A feature fusion module for stitching the image features and relative pose features at the same moment to obtain fused features; The lane number identification module is used to input the fusion features at multiple consecutive moments into the lane number identification model to obtain the lane number of the current lane where the vehicle is located and the confidence level of the lane number.
8. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is used to implement the steps in the lane positioning method according to any one of claims 1-6 when executing the programs stored on the memory.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps in the lane positioning method according to any one of claims 1-6.