Model training method, lane line detection method, electronic device, storage medium and program product
By using the method of space-time fusion module and quantization tools in the lane line detection model training, the problem of large amount of model calculation and difficulty in detecting blocked lane lines in the prior art is solved, and more efficient and accurate lane line detection is achieved.
Patent Information
- Application Number
- CN202510101595.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-02
AI Technical Summary
The existing lane line detection method based on deep learning is used in the model training stage. Since the model stores a large number of floating-point parameters, the model runs through the model, and it is difficult to detect some of the blocked lane lines.
A model training method is adopted, including obtaining training data sets, creating neural network models, training and optimization, and quantizing the model through quantization tools to reduce the number of bits of floating-point parameters and improve detection efficiency. At the same time, the space-time fusion module is used to perform timing fusion on the images of continuous frames, thereby improving the detection ability of obscured lane lines.
The calculation amount of the model during operation is reduced, the detection ability of obscured lane lines in continuous frame images is improved, and the path planning and safety of the assisted driving system is enhanced.
Smart Images

Figure CN119919907A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle technology, and in particular to a model training method, a lane line detection method, an electronic device, a storage medium and a program product. Background Art
[0002] With the rapid development of assisted driving technology, autonomous navigation and driving safety of vehicles have become key areas of research and development. Lane detection is an important part of the assisted driving system. It provides vehicles with road boundary information to help implement path planning, lane centering, lane changing and other operations. The rapid development of deep learning has provided new ideas for lane detection in traffic scenes. The current mainstream method of lane detection based on deep learning has a large amount of computational complexity during model training because the model stores a large number of floating-point parameters. In addition, the model relies on single-frame image processing, making it difficult to detect partially obscured lane lines. Summary of the invention
[0003] In view of this, the purpose of the embodiments of the present application is to provide a model training method, a lane line detection method, an electronic device, a storage medium and a program product, which can improve the problem of large amount of calculation when the model is running and difficulty in detecting partially obscured lane lines.
[0004] In order to achieve the above technical objectives, the technical solutions adopted in this application are as follows:
[0005] In a first aspect, an embodiment of the present application provides a model training method, the method comprising:
[0006] Acquire a training data set, and create a neural network model based on a creation instruction, wherein the neural network model includes an image feature extraction module, a spatiotemporal fusion module, a deep feature extraction module, and an output module, wherein the training data set includes N image groups, each of which includes a plurality of first images acquired in a continuous time, wherein the first images are annotated with a true value label representing a lane line, and N is an integer greater than or equal to 1;
[0007] Based on the training data set, the image feature extraction module, the spatiotemporal fusion module, the deep feature extraction module and the output module, the neural network model is trained to obtain a trained neural network model;
[0008] Based on a preset optimization strategy, the trained neural network model is optimized to obtain an optimized neural network model;
[0009] The optimized neural network model is quantized according to a quantization tool to obtain a quantized neural network model as a detection model for lane line detection.
[0010] In combination with the first aspect, in some optional implementations, obtaining the training data set includes:
[0011] Acquire time-continuous multi-view images collected by the vehicle camera and point cloud data collected by the radar;
[0012] Segmenting the point cloud data according to a preset time interval to obtain point cloud segments;
[0013] Reconstructing roads for each of the point cloud segments to obtain a bird's-eye view map corresponding to each of the point cloud segments;
[0014] Using a labeling tool, labeling true value labels representing lane lines in the overhead map;
[0015] Based on the odometer and posture information of the vehicle, each multi-view image corresponding to the point cloud segment is matched to the bird's-eye view map to obtain a true value label representing a lane line in the multi-view image;
[0016] All the multi-view images with the true value labels and the parameters corresponding to the cameras are converted into data in a specified format to form the training data set.
[0017] In combination with the first aspect, in some optional implementations, the neural network model is trained based on the training data set, the image feature extraction module, the spatiotemporal fusion module, the deep feature extraction module and the output module to obtain a trained neural network model, including:
[0018] Extracting features of the first image in the training data set by the image feature extraction module to obtain first image features;
[0019] According to the parameters corresponding to the camera at each viewing angle, determine the homography transformation matrix from the corresponding viewing angle to the top-view viewing angle, and based on the homography transformation matrix, convert the first image feature of the corresponding viewing angle into the second image feature of the top-view viewing angle;
[0020] Aligning the temporal fusion features of the first image of the previous frame in the image group with the second image features corresponding to the first image of the current frame through the temporal and spatial fusion module, so as to convert the temporal fusion features of the first image of the previous frame into third image features unified in the vertical coordinate system of the current frame, wherein the temporal fusion features of the first image of the previous frame are image features in the cache;
[0021] In each image group, the second image feature of the first image of the current frame and the third image feature corresponding to the previous frame are convolved and spliced, and the fourth image feature obtained after the splicing is convolved with a convolution kernel of a specified size to convert the number of channels of the fourth image feature to the number of channels representing the spatial fusion feature, to obtain a fifth image feature representing the temporal fusion, and the fifth image feature is cached as the temporal fusion feature of the first image of the current frame;
[0022] Performing deep feature extraction on the fifth image feature by the deep feature extraction module to obtain a sixth image feature;
[0023] The output module outputs the lane line coordinates and categories in the current first image corresponding to the sixth image feature based on the sixth image feature or based on the sixth image feature and the temporal fusion feature corresponding to the previous frame of the first image in the image group, and obtains the trained neural network model.
[0024] In combination with the first aspect, in some optional implementations, performing deep feature extraction on the fifth image feature by the deep feature extraction module to obtain a sixth image feature includes:
[0025] cropping a region of interest from the fifth image feature based on a target detection algorithm;
[0026] The deep feature extraction module performs deep feature extraction on the region of interest to obtain the sixth image feature.
[0027] In combination with the first aspect, in some optional implementations, the preset optimization strategy includes an adaptive learning rate optimization strategy;
[0028] The method of optimizing the trained neural network model based on a preset optimization strategy to obtain an optimized neural network model includes:
[0029] The adaptive learning rate optimization strategy is used to adjust the model parameters of the trained neural network model, and training is performed based on the neural network model after adjusting the model parameters, and the cross entropy loss function is used as the objective function to calculate the loss value of the neural network model after each training until the loss value converges to obtain the optimized neural network model, wherein the model parameters include the learning rate and the momentum decay coefficient.
[0030] In conjunction with the first aspect, in some optional implementations, the quantization tool includes a TensorRT tool;
[0031] The step of quantizing the optimized neural network model according to the quantization tool to obtain the quantized neural network model includes:
[0032] Convert the optimized neural network model into a model in a preset format supported by the TensorRT tool;
[0033] The model in the preset format is input into the TensorRT tool, and a quantization operation is performed to convert the number of bits of the floating-point parameters in the optimized neural network model from initial bits to specified bits, so as to obtain the quantized neural network model, wherein the number of bits of the initial bits is higher than the specified bits.
[0034] In a second aspect, an embodiment of the present application further provides a lane line detection method, which is applied to the detection model obtained by the above-mentioned model training method, and the method includes:
[0035] Acquire a second image, where the second image is an image obtained by photographing a road surface in a driving environment where the vehicle is located;
[0036] Input the second image into the quantized detection model, wherein the detection model stores temporal fusion features of historical images within a preset time before the second image is taken, wherein the temporal fusion features are image features representing temporal fusion obtained by extracting features from the historical images by the detection model;
[0037] Based on the temporal fusion features, the second image is detected by the detection model to obtain a detection result indicating whether a lane line exists in the second image.
[0038] In conjunction with the second aspect, in some optional implementations, the method further includes:
[0039] During the detection of the second image by the detection model, the temporal fusion features obtained by the detection model performing image feature extraction on the second image are cached, and the cached temporal fusion features are used for temporal fusion with the next frame of the second image.
[0040] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a processor and a memory coupled to each other, wherein the memory stores a computer program, and when the computer program is executed by the processor, the electronic device executes the model training method of the first aspect described above, or executes the lane line detection method of the second aspect described above.
[0041] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, the computer executes the model training method of the first aspect or the lane line detection method of the second aspect.
[0042] In a fifth aspect, an embodiment of the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the model training method of the first aspect, or implements the lane line detection method of the second aspect.
[0043] The invention adopting the above technical solution has the following advantages:
[0044] In the technical solution provided in the present application, the neural network model includes an image feature extraction module, a spatiotemporal fusion module, a deep feature extraction module and an output module. During the model training process, the spatiotemporal fusion module can perform temporal fusion on the image features of two temporally adjacent images, and utilize the temporal fusion between images of consecutive frames so that the trained neural network model can have the ability to perform temporal fusion processing on images of consecutive frames, thereby facilitating the ability to perform lane line recognition on images of lane lines that are obscured in consecutive frame images. In addition, by quantizing the model, the number of bits of floating-point parameters in the model can be reduced, thereby facilitating the reduction of the amount of calculation of the detection model during operation and accelerating the reasoning process of lane line detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The present application may be further described by the non-limiting embodiments given in the accompanying drawings. It should be understood that the following drawings only illustrate certain embodiments of the present application and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings may be obtained based on these drawings without creative effort.
[0046] Figure 1 A flowchart of the model training method provided in an embodiment of the present application.
[0047] Figure 2 A schematic diagram of a scenario for collecting and processing training data provided in an embodiment of the present application.
[0048] Figure 3 A schematic diagram of the network structure of a neural network model provided in an embodiment of the present application.
[0049] Figure 4 A schematic diagram of the time-series fusion of multiple image frames provided in an embodiment of the present application.
[0050] Figure 5 A schematic diagram of the fusion of temporal features and spatial features provided in an embodiment of the present application.
[0051] Figure 6 A flowchart of a lane detection method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] The present application will be described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that in the drawings or descriptions, similar or identical parts use the same figure numbers, and the implementation methods not shown or described in the drawings are forms known to ordinary technicians in the relevant technical field. In the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0053] Example 1
[0054] The present application provides an electronic device, which may include a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, the electronic device can execute the corresponding steps in the following model training method, or execute the steps in the lane line detection method in the following second embodiment.
[0055] In this embodiment, the processor may be an integrated circuit chip having the ability to process signals. For example, the processor may be, but is not limited to, a central processing unit (CPU) or a graphics processing unit (GPU), and may implement or execute the methods, steps, and logic diagrams disclosed in the embodiments of the present application.
[0056] The memory may be, but is not limited to, a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc. In this embodiment, the memory may be used to store detection models, detection results, time series fusion features of historical images, etc. Of course, the memory may also be used to store programs, and the processor executes the program after receiving the execution instruction.
[0057] It should be noted that the electronic device may be a graphics card server, or other device that can be used for model training. The memory and the processor may be integrated or exist independently, which is not specifically limited here.
[0058] Please refer to Figures 1 to 4 The present application provides a model training method, which can be applied to the above-mentioned electronic device, and each step of the method can be executed or implemented by the electronic device. The model training method can include the following steps:
[0059] Step 110, obtaining a training data set, and creating a neural network model based on a creation instruction, wherein the neural network model includes an image feature extraction module, a spatiotemporal fusion module, a deep feature extraction module, and an output module, wherein the training data set includes N image groups, each of which includes a plurality of first images acquired in a continuous time, wherein the first images are annotated with a true value label representing a lane line, and N is an integer greater than or equal to 1;
[0060] Step 120, training the neural network model based on the training data set, the image feature extraction module, the spatiotemporal fusion module, the deep feature extraction module and the output module to obtain a trained neural network model;
[0061] Step 130, optimizing the trained neural network model based on a preset optimization strategy to obtain an optimized neural network model;
[0062] Step 140, quantizing the optimized neural network model according to a quantization tool to obtain a quantized neural network model as a detection model for lane line detection.
[0063] The following is a detailed description of each step of the model training method, as follows:
[0064] In step 110, obtaining a training data set may include:
[0065] Acquire time-continuous multi-view images collected by the vehicle camera and point cloud data collected by the radar;
[0066] The point cloud data is segmented according to a preset interval to obtain point cloud segments. The preset interval can be flexibly set according to actual conditions, for example, 1 minute;
[0067] Reconstructing the road for each of the point cloud segments to obtain a bird's eye view map corresponding to each of the point cloud segments, the bird's eye view map being a BEV (Bird's Eye View) map;
[0068] Using a labeling tool, labeling true value labels representing lane lines in the overhead map, and the labeling tool can be flexibly selected according to actual conditions;
[0069] Based on the odometer and posture information of the vehicle, each multi-view image corresponding to the point cloud segment is matched to the bird's-eye view map to obtain a true value label representing a lane line in the multi-view image;
[0070] Convert all the multi-view images with the true value labels and the parameters corresponding to the cameras into data in a specified format to form the training data set, wherein the parameters corresponding to the cameras may include but are not limited to the installation position and orientation of the cameras on the vehicle, the focal length of the cameras, etc. The specified format may be flexibly selected according to actual conditions, for example, it may be an LMDB format, which may reduce disk I / O operations, thereby facilitating improved data reading speed.
[0071] As an example, the training dataset can be obtained by:
[0072] Please refer to Figure 2 First, the multi-view images collected by the cameras with multiple viewpoints on the vehicle can be adapted according to the number of cameras on different vehicles. For example, a vehicle equipped with 7 cameras and three laser radars is used to collect data during driving to obtain multi-view images and point cloud data. Among them, the resolution of the front-view camera and the front-view narrow-angle camera can be 2160×3840, and the resolution of the remaining 5 surround-view cameras can be 1536×1920. Before lane line marking, the collected point cloud data will be divided into Clips (i.e., point cloud fragments) by minute, and then each Clip will be reconstructed to obtain the BEV map of the section. Road reconstruction is a conventional implementation method, which is used to convert the road under the perspective of the laser radar into the road under the top-down perspective, so as to obtain the BEV map. Next, the lane lines are annotated on the BEV map using the key point annotation tool as true value labels. In addition, according to the odometer and posture information of the vehicle, the multi-view image corresponding to each Clip can be matched to the true value of the BEV map, and the true value label of each point cloud fragment at each moment can be obtained, wherein the matching method can be implemented by an image registration algorithm. After image registration, if the image area / pixel in the multi-view image matches the image area / pixel where the true value label is located on the BEV map, the matching image area / pixel in the multi-view image has a true value label. A Clip can be used to annotate an image group. By annotating a Clip, engineers can obtain the annotations of all images during the acquisition of the Clip, without manually annotating each image, which is conducive to improving the efficiency of image annotation. After the annotation is completed, the true value label, image data and camera parameters can be packaged into LMDB format according to the timestamp when the data was collected, thereby forming a training data set. When the training data set is used for training later, the true value label and camera parameters at any time can be obtained according to the timestamp.
[0073] It should be noted that the first image can be understood as an image during the training of the neural network model; the second image described below can be understood as an image that needs to be used for lane line detection using the model after the neural network model has completed training.
[0074] During the creation of the neural network model, engineers can send creation instructions to the graphics card server through a local device (such as a personal computer) for detection purposes, thereby controlling the graphics card server to create a neural network model with an image feature extraction module, a spatiotemporal fusion module, a deep feature extraction module, and an output module. The structure of the neural network model can be referred to Figure 3 .
[0075] In step 120, the neural network model is trained based on the training data set, the image feature extraction module, the spatiotemporal fusion module, the deep feature extraction module and the output module to obtain a trained neural network model, which may include:
[0076] Step 121, extracting features of the first image in the training data set by the image feature extraction module to obtain first image features;
[0077] Step 121, determining a homography transformation matrix from a corresponding perspective to a top-down perspective according to parameters corresponding to the camera at each perspective, and converting a first image feature of the corresponding perspective into a second image feature of the top-down perspective based on the homography transformation matrix;
[0078] Step 122, aligning the temporal fusion features of the first image of the previous frame in the image group with the second image features corresponding to the first image of the current frame through the temporal and spatial fusion module, so as to convert the temporal fusion features of the first image of the previous frame into third image features unified in the vertical coordinate system of the current frame / current moment, wherein the temporal fusion features of the first image of the previous frame are the fifth image features of the cached first image of the previous frame;
[0079] Step 123, in each image group, convolve and splice the second image feature of the first image of the current frame and the third image feature corresponding to the previous frame, and convolve the fourth image feature obtained after splicing with a convolution kernel of a specified size to convert the number of channels of the fourth image feature to the number of channels representing the spatial fusion feature, obtain the fifth image feature representing the time series fusion, and cache the fifth image feature as the time series fusion feature of the first image of the current frame, wherein the cached time series fusion feature of the current frame is used for time series fusion processing with the next frame of image;
[0080] Step 124, performing deep feature extraction on the fifth image feature by the deep feature extraction module to obtain a sixth image feature;
[0081] Step 125, the output module outputs the lane line coordinates and category in the current first image corresponding to the sixth image feature based on the sixth image feature or based on the sixth image feature and the temporal fusion feature corresponding to the previous frame of the first image in the image group, and obtains the trained neural network model.
[0082] In this embodiment, step 122 and step 123 can be understood as a process of performing temporal fusion on the first image of the current frame and the temporal fusion features of the first image of the previous frame. The temporal fusion features in the cache can be used as historical features for temporal fusion with the first image of the next frame. The process of temporal fusion can refer to the above steps 122 and 123, which will not be repeated here. In addition, the cached temporal fusion features can update the temporal fusion features originally stored in the cache module, that is, only the temporal fusion features of the first image of the most recent frame can be cached each time to reduce the amount of cached data.
[0083] If the first image of the current frame is the first image in the image group, the fifth image feature of the last image in the previous image group (the last image and the first image of the current frame are images collected in the same batch and have the characteristics of continuity in time and space) can be used as the time series fusion feature of the first image of the previous frame. If the first image of the current frame is the first image in the image group, and the images of the previous image group and the current image group are not continuous, then at this time, it is not necessary to align the features of the first image of the current frame, and the image stitching process can be omitted. At this time, the second image feature of the first image of the current frame is directly subjected to a two-layer convolution, and a convolution of a convolution kernel of a specified size is performed. The obtained image feature is used as the fifth image feature and cached as a historical feature. The historical feature is used for time series fusion with the next frame of the image. In this way, the images in the image group can be sequentially fused one by one.
[0084] The step of performing deep feature extraction on the fifth image feature by the deep feature extraction module to obtain the sixth image feature may include:
[0085] cropping a region of interest from the fifth image feature based on a target detection algorithm;
[0086] The deep feature extraction module performs deep feature extraction on the region of interest to obtain the sixth image feature.
[0087] In this embodiment, by cutting out the region of interest and performing deep feature extraction on the region of interest, it is beneficial to reduce the amount of calculation and improve the training efficiency. Among them, the target detection algorithm can be flexibly selected according to the actual situation, and is not specifically limited here.
[0088] Please refer to Figure 4and Figure 5 , as an example, the training process of the neural network model can be as follows:
[0089] In the first step, a lightweight backbone network such as RegNet is used as an image feature extraction module to extract the shallow features (i.e., the first image features) of each view image (i.e., the first image). RegNet is a convolutional neural network for image recognition that can be used to improve the performance and efficiency of the model. The network structure improves the performance of the model by increasing the network depth and reducing the network width.
[0090] In the second step, the camera parameters (including the camera's installation position, orientation and focal length on the vehicle) are used to calculate the homography transformation matrix from each perspective to the BEV perspective. The first image features of each perspective are then projected to the BEV perspective to obtain the BEV features. The BEV features of each perspective are preliminarily fused by element-by-element addition to obtain the second image features.
[0091] The third step is to perform temporal fusion on the second image features of each image in each image group. Figure 4 , Figure 5As shown in the figure, the homography transformation matrix from the VCS-XY plane at the current moment to the VCS-XY plane where the BEV feature at the historical moment is located is calculated through the inter-frame motion odometer information of the vehicle, and the inverse mapping relationship of the image coordinates between the two is generated. The warping operator can be used to align the corresponding features of two adjacent moments, such as aligning the BEV spatial features at time t (the second image features of the first image at the current frame / time t) with the temporal fusion features at time t-1 (the historical features of the first image at the previous frame / time t-1), and converting the temporal fusion features at the previous moment into BEV features unified in the VCS coordinate system at the current moment as the third image feature. The warping operator can be used for image alignment and image stitching. Then, the BEV spatial features at time t and the temporal fusion features at time t-1 are respectively subjected to two layers of convolution and then spliced to obtain the fourth image features. The fourth image features are then sent to the 1x1 convolution module for convolution to reduce the number of channels to the number of channels representing the spatial fusion features, thereby obtaining the BEV features after temporal fusion as the fifth image feature, which is the temporal fusion feature of the first image at time t. The fifth image feature can be input into the deep feature extraction module on the one hand, and cached on the other hand for temporal fusion with the first image at time t+1. In the form of Recurren, the system can cache only one frame of historical features. Among them, VCS-XY can be understood as the vertical coordinate system 0-XY, which can be a coordinate system established on the plane where the bird's-eye view map is located. The historical features are used to correct the output data of the output module and identify the lane lines in the occluded area, which is conducive to smoothing the output of the model. For example, the lane lines output by the model in the first image of the current frame are crooked due to various reasons, but the lane lines in history are all straight. Then, the neural network model can make certain corrections to the output of the current frame based on the historical features to improve the problem of crooked lane lines.
[0092] In the fourth step, the target detection algorithm is used to cut out the valuable area (i.e., the region of interest) from the first-stage BEV features after spatiotemporal fusion, and then the backbone network is used to extract the second-stage deep features of the region of interest. The backbone network can be the same as or different from the backbone network used to extract the first-stage shallow features. For example, the backbone network used to extract the second-stage deep features can be, but is not limited to, AlexNet and ResNet.
[0093] In the fifth step, the output module can divide the first image used for training into multiple BEV grids, and output the lane line coordinates and categories in each BEV grid. If there is no lane line in the grid, the result of no lane line is output. Among them, the output of the category is a binary classification vector, that is, the grid is divided into foreground or background, and the foreground is the lane line; the lane line coordinate output format is [R, Sin, Cos], and the offset dx = R × Sin, dy = R × Cos relative to the center of the grid can regress the accurate lane line coordinates. By applying a certain threshold to the model "ClassMap", a foreground point set can be obtained, and then various attribute values corresponding to these foreground points can be obtained. These foreground points belong to multiple lane line instances, so they need to be clustered by a clustering algorithm (such as the DBSCAN clustering algorithm) to distinguish these instances, split a large point set into multiple sub-point sets, and finally output the lane line instance, and obtain a neural network model that completes one training. By repeating the first to fifth steps, different first images are trained to obtain a neural network model with higher detection accuracy.
[0094] In this embodiment, the preset optimization strategy includes an Adam (Adaptive Moment Estimation) strategy. In step 130, based on the preset optimization strategy, the trained neural network model is optimized to obtain an optimized neural network model, which may include:
[0095] The adaptive learning rate optimization strategy is used to adjust the model parameters of the trained neural network model, and training is performed based on the neural network model after adjusting the model parameters, and the cross entropy loss function is used as the objective function to calculate the loss value of the neural network model after each training until the loss value converges to obtain the optimized neural network model, wherein the model parameters include the learning rate and the momentum decay coefficient.
[0096] As an example, the initial learning rate is set to 0.005, the first-order momentum attenuation coefficient is set to 0.9, and the second-order momentum attenuation coefficient is set to 0.999. The BEV sensing range is set to 102.6 meters in front, 51 meters in the back, and 57.6 meters on the left and right. The BEV sensing range is the image area range in the first image to be detected. If the image area range exceeds the BEV sensing range, no detection is required.
[0097] During each training process, after each prediction, the model will compare the result with the true value label and calculate a loss value, which represents the difference between the predicted result and the true value label. The goal of model training is to reduce the loss value. In each round of training, the model will reversely derive each neuron from back to front based on the loss value, calculate the gradient vector, and then determine the adjustment direction of the neuron value based on the gradient vector to make the model converge.
[0098] In this embodiment, the quantization tool includes a TensorRT tool. In step 140, the optimized neural network model is quantized according to the quantization tool to obtain a quantized neural network model, including:
[0099] Convert the optimized neural network model into a model in a preset format supported by the TensorRT tool, where the preset format may be an ONNX format. The ONNX (Open Neural Network Exchange) format is an open model file format designed to promote interoperability between different deep learning frameworks;
[0100] The model in the preset format is input into the TensorRT tool, and a quantization operation is performed to convert the number of bits of the floating-point parameters in the optimized neural network model from initial bits to specified bits, so as to obtain the quantized neural network model, wherein the number of bits of the initial bits is higher than the specified bits.
[0101] In this embodiment, the trained model is converted into a format supported by TensorRT and optimized, including layer fusion, precision calibration, quantization, etc. Before the neural network model is quantized, the model stores a large number of 32-bit or 16-bit floating-point parameters, which requires a large amount of calculation and slows down the reasoning speed. The quantization operation can convert 32-bit or 16-bit floating-point parameters into low-order floating-point parameters, for example, into 8-bit floating-point parameters, thereby reducing the amount of calculation and improving the reasoning speed of the model.
[0102] The quantized neural network model is a detection model with lane line detection capability, which can be deployed in the vehicle's Orin-X chip. Orin-X is a system chip for autonomous driving, which enables the vehicle to automatically detect lane lines.
[0103] Example 2
[0104] Please refer to Figure 6 , the present application also provides a lane line detection method, which can be applied to electronic devices. It is understandable that when an electronic device is used to implement the steps of the lane line detection method, the electronic device can be an electronic control system deployed in a vehicle. The processor of the electronic device can be an Orin-X chip. Among them, the detection model that the lane line detection method relies on can be obtained by the model training method provided in Example 1, and the detection model can be deployed in the Orin-X chip of the electronic device. The lane line detection method can include the following steps:
[0105] Step 210, acquiring a second image, where the second image is an image obtained by photographing a road surface in a driving environment where the vehicle is located;
[0106] Step 220, inputting the second image into the quantized detection model, wherein the detection model stores temporal fusion features of historical images within a preset time before the second image is captured, wherein the temporal fusion features are image features representing temporal fusion obtained by extracting features from the historical images by the detection model;
[0107] Step 230: Based on the temporal fusion feature, the second image is detected by the detection model to obtain a detection result indicating whether a lane line exists in the second image.
[0108] In this embodiment, the second image refers to the image to be identified after the model completes quantization. The first image in the first embodiment is an image during training. If the second image is the first image after the vehicle turns on the lane recognition function, there is no historical feature when executing step 230. At this time, the second image can be directly detected by the detection model to obtain the corresponding detection result. In addition, during the detection of the first second image, the image feature representing time series fusion obtained by the detection model performing image feature extraction on the second image (that is, the fifth image feature in Example 1) will be cached as the historical feature of the next frame of the second image.
[0109] Since the changes in lane lines are often time-series, they can also be understood as continuity. For example, in the second images of several historical frames, the lane lines are not blocked, but the second image of the current frame is blocked. Then the detection model can identify the lane lines in the blocked area based on historical features. In addition, historical features are conducive to smoothing the output of the model. For example, the lane lines output by the model in the second image of the current frame are very crooked due to various reasons, but the lane lines in history are all straight. Then, the model can make certain corrections to the output of the current frame based on historical information.
[0110] After step 230, the method may further include:
[0111] During the detection of the second image by the detection model, the temporal fusion features obtained by the detection model through image feature extraction of the second image are cached, and the cached temporal fusion features are used for temporal fusion with the second image of the next frame, wherein the temporal fusion features refer to the fifth image features of the second image of the current frame.
[0112] Based on the above design, the model training method and the lane line detection method are not coupled with any specific control system and can be used in the lane line detection module of any assisted driving system. This solution not only utilizes the spatial features in a single-frame image, but also combines the temporal information of continuous multi-frame images. The detection model effectively integrates the temporal and spatial features to achieve more stable and accurate lane line detection. The present invention can maintain a high detection accuracy in a complex driving environment, effectively solve the problem that the continuous lane line is partially blocked (or partially unclear) and the blocked (unclear) lane line cannot be detected, and provide strong protection for the path planning and safety of the assisted driving system. The stability of the detection is enhanced by the continuity in the multi-frame image, which helps to eliminate noise and errors in a single frame; in addition, the forward reasoning time of the model is less than 30ms, which meets the real-time reasoning requirements of the vehicle-mounted system.
[0113] The lane line detection method can perform real-time lane line detection on a real vehicle, and the detection results can be transmitted to the control module for decision-making. This lane line detection method based on a convolutional neural network with time series information can not only solve the problem of limited computing power of the assisted driving vehicle embedded platform, but also solve the problem that traditional algorithms cannot detect partially blocked (or unclear) lane lines when continuous lane lines are partially blocked (or partially unclear).
[0114] In actual scenarios, lane lines are often blocked by vehicles or part of the lane lines are unclear. The present invention uses the timing information between consecutive frames to improve the incomplete or unclear lane lines in the current frame due to occlusion or blur, thereby improving the detection accuracy. According to experiments, the forward reasoning speed of the present invention is only 23 milliseconds, which not only improves the accuracy relative to the single-task model, but also greatly improves the efficiency compared to multiple single tasks.
[0115] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed on a computer, the computer executes the model training method or lane line detection method described in the above embodiment.
[0116] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements each step in the above-mentioned model training method or lane line detection method.
[0117] Through the description of the above implementation methods, technical personnel in this field can clearly understand that the present application can be implemented by hardware, and can also be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including a number of instructions for enabling a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute the methods described in each implementation scenario of the present application.
[0118] In the embodiments provided in the present application, it should be understood that the disclosed equipment and methods can also be implemented in other ways. The above-described device and method embodiments are merely schematic, for example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the equipment, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a part of a module, a program segment or a code, and a part of the module, program segment or code contains one or more executable instructions for implementing the specified logical function. It should also be noted that each box in the block diagram and / or the flowchart, and the combination of the boxes in the block diagram and / or the flowchart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. In addition, each functional module in each embodiment of the present application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0119] The above description is only an embodiment of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A model training method, characterized in that: The method comprises: Acquire a training data set, and create a neural network model based on a creation instruction, wherein the neural network model includes an image feature extraction module, a spatiotemporal fusion module, a deep feature extraction module, and an output module, wherein the training data set includes N image groups, each of which includes a plurality of first images acquired in a continuous time, wherein the first images are annotated with a true value label representing a lane line, and N is an integer greater than or equal to 1; Based on the training data set, the image feature extraction module, the spatiotemporal fusion module, the deep feature extraction module and the output module, the neural network model is trained to obtain a trained neural network model; Based on a preset optimization strategy, the trained neural network model is optimized to obtain an optimized neural network model; The optimized neural network model is quantized according to a quantization tool to obtain a quantized neural network model as a detection model for lane line detection.
2. The method according to claim 1, characterized in that The step of obtaining a training data set includes: Obtain time-continuous multi-view images collected by the vehicle camera and point cloud data collected by the radar; Segmenting the point cloud data according to a preset time interval to obtain point cloud segments; Reconstructing roads for each of the point cloud segments to obtain a bird's-eye view map corresponding to each of the point cloud segments; Using a labeling tool, labeling true value labels representing lane lines in the overhead map; Based on the odometer and posture information of the vehicle, each multi-view image corresponding to the point cloud segment is matched to the bird's-eye view map to obtain a true value label representing a lane line in the multi-view image; All the multi-view images with the true value labels and the parameters corresponding to the cameras are converted into data in a specified format to form the training data set.
3. The method according to claim 1, characterized in that: The neural network model is trained based on the training data set, the image feature extraction module, the spatiotemporal fusion module, the deep feature extraction module and the output module to obtain a trained neural network model, including: Extracting features of the first image in the training data set by the image feature extraction module to obtain first image features; According to the parameters corresponding to the camera at each viewing angle, determine the homography transformation matrix from the corresponding viewing angle to the top-view viewing angle, and based on the homography transformation matrix, convert the first image feature of the corresponding viewing angle into the second image feature of the top-view viewing angle; Aligning the temporal fusion features of the first image of the previous frame in the image group with the second image features corresponding to the first image of the current frame through the temporal and spatial fusion module, so as to convert the temporal fusion features of the first image of the previous frame into third image features unified in the vertical coordinate system of the current frame, wherein the temporal fusion features of the first image of the previous frame are image features in the cache; In each image group, the second image feature of the first image of the current frame and the third image feature corresponding to the previous frame are convolved and spliced, and the fourth image feature obtained after the splicing is convolved with a convolution kernel of a specified size to convert the number of channels of the fourth image feature to the number of channels representing the spatial fusion feature, to obtain a fifth image feature representing the temporal fusion, and the fifth image feature is cached as the temporal fusion feature of the first image of the current frame; Performing deep feature extraction on the fifth image feature by the deep feature extraction module to obtain a sixth image feature; The output module outputs the lane line coordinates and categories in the current first image corresponding to the sixth image feature based on the sixth image feature or based on the sixth image feature and the temporal fusion feature corresponding to the previous frame of the first image in the image group, and obtains the trained neural network model.
4. The method according to claim 3, characterized in that: The step of performing deep feature extraction on the fifth image feature by the deep feature extraction module to obtain a sixth image feature includes: cropping a region of interest from the fifth image feature based on a target detection algorithm; The deep feature extraction module performs deep feature extraction on the region of interest to obtain the sixth image feature.
5. The method according to claim 1, characterized in that The preset optimization strategy includes an adaptive learning rate optimization strategy; The method of optimizing the trained neural network model based on a preset optimization strategy to obtain an optimized neural network model includes: The adaptive learning rate optimization strategy is used to adjust the model parameters of the trained neural network model, and training is performed based on the neural network model after adjusting the model parameters, and the cross entropy loss function is used as the objective function to calculate the loss value of the neural network model after each training until the loss value converges to obtain the optimized neural network model, wherein the model parameters include the learning rate and the momentum decay coefficient.
6. The method according to claim 1, characterized in that The quantization tool includes a TensorRT tool; The step of quantizing the optimized neural network model according to the quantization tool to obtain the quantized neural network model includes: Convert the optimized neural network model into a model in a preset format supported by the TensorRT tool; The model in the preset format is input into the TensorRT tool, and a quantization operation is performed to convert the number of bits of the floating-point parameters in the optimized neural network model from initial bits to specified bits, so as to obtain the quantized neural network model, wherein the number of bits of the initial bits is higher than the specified bits.
7. A lane line detection method, characterized in that: Applied to the detection model obtained by the method according to any one of claims 1 to 6, the method comprising: Acquire a second image, where the second image is an image obtained by photographing a road surface in a driving environment where the vehicle is located; Input the second image into the quantized detection model, wherein the detection model stores temporal fusion features of historical images within a preset time before the second image is taken, wherein the temporal fusion features are image features representing temporal fusion obtained by extracting features from the historical images by the detection model; Based on the temporal fusion features, the second image is detected by the detection model to obtain a detection result indicating whether a lane line exists in the second image.
8. The method according to claim 7, characterized in that The method further comprises: During the detection of the second image by the detection model, the temporal fusion features obtained by the detection model performing image feature extraction on the second image are cached, and the cached temporal fusion features are used for temporal fusion with the next frame of the second image.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory coupled to each other, wherein the memory stores a computer program. When the computer program is executed by the processor, the electronic device executes the method as claimed in any one of claims 1 to 6, or executes the method as claimed in claim 7 or 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 6, or the method according to claim 7 or 8.
11. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6, or performs the method according to claim 7 or 8.