Training method and device of lane line recognition model and lane line recognition method
Through the online deep mutual learning and feature fusion technology of package coding networks and sample coding networks, the problems of complex lane line recognition and difficulty in model training are solved, and high-precision and high-efficiency lane line detection are achieved.
Patent Information
- Application Number
- CN202510134278.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, lane line recognition is complex, model training is difficult, detection accuracy and speed are low, and prior knowledge of lane line cannot be fully utilized.
Packet encoding network and sample encoding network are used for online deep mutual learning. By fusing global feature maps and local feature maps, a fusion feature map is generated as a supervised learning signal, network parameters are adjusted, model complexity is reduced, and detection accuracy and speed are improved.
The detection accuracy and speed of the lane line recognition model is improved, the training difficulty is reduced, the prior knowledge of lane line is fully utilized, and the practical value of the model is enhanced.
Smart Images

Figure CN120014581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method and device for training a lane line recognition model, and a method for recognizing lane lines. Background Art
[0002] Globally, with the rapid increase in the number of vehicles, the safety of cars during driving has become particularly important, and lane recognition is an important part. In order to ensure the safety of cars during driving, this requires accurate perception of lanes. The purpose of lane detection is to obtain the accurate shape of each lane on the road; that is, it is not only required to obtain the direction and shape of the lane, but also to distinguish each lane instance. Lane recognition faces many challenges, such as complex road conditions, occlusion, lighting changes, and semantic ambiguity. When applied to vehicle systems, the algorithm needs to meet high real-time requirements and run efficiently on limited hardware resources, which increases the difficulty of the task. Due to the advancement of deep learning technology, researchers have developed many strategies to greatly simplify, speed up, and enhance the task of lane recognition.
[0003] In the existing technology, lane line areas are segmented through edge detection, filtering and other technologies, which requires manual adjustment of parameters and filters, which is a lot of work and makes lane line recognition complicated. The detection accuracy is improved by increasing the network depth and width of a single model, which makes the network structure of the model more and more complex, and the model is prone to overfitting and difficult to train, and usually requires more computing resources, which seriously restricts the practical value of the detection model. The existing network model cannot pay attention to the lines in the lane line image information in a differentiated manner, feature information cannot be fully extracted, and the performance is poor under severe occlusion. The prior knowledge of the lane line is not fully utilized, which will lead to a decrease in the detection accuracy and speed of the network model. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a lane line recognition model training method and device, and a lane line recognition method to achieve the purpose of solving the complex problem of lane line recognition, reducing the difficulty of training, and improving the model detection accuracy and speed.
[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0006] A first aspect of an embodiment of the present invention discloses a method for training a lane recognition model, the method comprising:
[0007] Obtain a data set including multiple lane line images;
[0008] Preprocessing each lane line image in the data set, and dividing the preprocessed data set into a training set and a test set;
[0009] Constructing a lane line recognition model to be trained; the lane line recognition model to be trained includes: a packet encoding network and an example encoding network;
[0010] Inputting the lane line images in the training set into the packet coding network and the example coding network for online deep mutual learning, respectively, to obtain a global feature map output by the packet coding network and a local feature map output by the example coding network;
[0011] Fusing the global feature map and the local feature map to obtain a fused feature map;
[0012] Inputting the fused feature map as a supervised learning signal into the packet coding network and the example coding network respectively, adjusting the packet coding network and the example coding network, and obtaining a trained lane line recognition model when the packet coding network and the example coding network converge;
[0013] The trained lane line recognition model is evaluated using the test set, and if the evaluation passes, a trained lane line recognition model is obtained.
[0014] Preferably, the preprocessing of each lane line image in the data set includes:
[0015] Removing interference information from each lane line image; the interference information includes: non-lane line information;
[0016] and / or,
[0017] For each lane line image, data enhancement processing is performed to obtain multiple new lane line images; the data enhancement processing includes: rotation processing, upper and lower mirror processing and shearing processing;
[0018] and / or,
[0019] For each lane line image, image enhancement processing is performed to obtain a clear lane line image; the image enhancement processing includes: noise reduction processing.
[0020] Preferably, the step of constructing a lane line recognition model to be trained includes:
[0021] Constructing a packet coding network with a multi-view attention mechanism, so that when the packet coding network receives the lane line image, a first weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a global feature map is extracted from the first weighted feature map;
[0022] Constructing an example encoding network with the multi-view attention mechanism, so that when the example encoding network receives the lane line image, a second weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a local feature map is extracted from the second weighted feature map;
[0023] Based on the packet encoding network and the example encoding network, a lane line recognition model to be trained is constructed.
[0024] Preferably, the step of inputting the lane line images in the training set into the packet encoding network and the example encoding network for online deep mutual learning comprises:
[0025] Inputting the lane line images in the training set into the packet coding network and the example coding network for training respectively, and using a cross entropy loss function to adjust the training direction and training speed of the packet coding network and the example coding network;
[0026] Calculating the KL divergence value of the packet encoding network output and the KL divergence value of the example encoding network output respectively;
[0027] Based on the KL divergence value output by the packet coding network, the KL divergence value output by the example coding network and a Softmax activation function, mutual learning between output probabilities of the packet coding network and the example coding network is performed.
[0028] Preferably, the output parameters of the last pooling layer of the packet coding network and the last pooling layer of the example coding network are consistent, so that the size of the global feature map output by the packet coding network and the local feature map output by the example coding network are consistent. Accordingly, the global feature map and the local feature map are fused to obtain a fused feature map, including:
[0029] Performing a serial operation on the global feature map and the local feature map to obtain a serial feature map;
[0030] A 1×1 point-by-point convolution operation is performed on the series feature maps to obtain a fused feature map.
[0031] Preferably, the evaluating the trained lane line recognition model using the test set includes:
[0032] For each lane line image in the test set, input it into the trained lane line recognition model to obtain a global feature map and a local feature map output by the trained lane line recognition model;
[0033] The global feature map and the local feature map are fused to obtain a fused feature map to be identified;
[0034] Inputting the fused feature map to be identified into the trained classifier to obtain a lane line recognition result;
[0035] If the lane line recognition result is correct, it is determined that the trained lane line recognition model has passed the evaluation.
[0036] A second aspect of an embodiment of the present invention discloses a training device for a lane recognition model, the device comprising:
[0037] An acquisition unit, used for acquiring a data set including multiple lane line images;
[0038] A preprocessing unit, used to preprocess each lane line image in the data set, and divide the preprocessed data set into a training set and a test set;
[0039] A construction unit, used to construct a lane line recognition model to be trained; the lane line recognition model to be trained includes: a packet encoding network and an example encoding network;
[0040] A mutual learning unit, used to input the lane line images in the training set into the packet coding network and the example coding network for online deep mutual learning, so as to obtain a global feature map output by the packet coding network and a local feature map output by the example coding network;
[0041] A fusion unit, used for fusing the global feature map and the local feature map to obtain a fused feature map;
[0042] a reverse supervision unit, configured to input the fused feature map as a supervised learning signal into the packet coding network and the example coding network respectively, adjust the packet coding network and the example coding network, and obtain a trained lane line recognition model when the packet coding network and the example coding network converge;
[0043] An evaluation unit is used to evaluate the trained lane line recognition model using the test set, and if the evaluation passes, a trained lane line recognition model is obtained.
[0044] Preferably, the preprocessing unit for preprocessing each lane line image in the data set is specifically used to:
[0045] Removing interference information from each lane line image; the interference information includes: non-lane line information;
[0046] and / or,
[0047] For each lane line image, data enhancement processing is performed to obtain multiple new lane line images; the data enhancement processing includes: rotation processing, upper and lower mirror processing and shearing processing;
[0048] and / or,
[0049] For each lane line image, image enhancement processing is performed to obtain a clear lane line image; the image enhancement processing includes: noise reduction processing.
[0050] Preferably, the construction unit is specifically used for:
[0051] Constructing a packet coding network with a multi-view attention mechanism, so that when the packet coding network receives the lane line image, a first weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a global feature map is extracted from the first weighted feature map;
[0052] Constructing an example encoding network with the multi-view attention mechanism, so that when the example encoding network receives the lane line image, a second weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a local feature map is extracted from the second weighted feature map;
[0053] Based on the packet encoding network and the example encoding network, a lane line recognition model to be trained is constructed.
[0054] A third aspect of an embodiment of the present invention discloses a lane line recognition method, the method comprising:
[0055] Acquire a target image to be identified;
[0056] Inputting the target image into a pre-built lane line recognition model to obtain a target local feature map and a target global feature map output by the lane line recognition model; the lane line recognition model is obtained by online deep mutual learning between a packet coding network and an example coding network;
[0057] Fusing the target local feature map and the target global feature map to obtain a target fused feature map;
[0058] The target fusion feature map is input into a pre-trained classifier to obtain a lane line recognition result of the target image.
[0059] Based on the training method and device of a lane line recognition model and a lane line recognition method provided by the above-mentioned embodiment of the present invention, a data set including multiple lane line images is collected; each lane line image in the data set is preprocessed, and the preprocessed data set is divided into a training set and a test set; a lane line recognition model to be trained is constructed; the lane line recognition model to be trained includes: a packet coding network and an example coding network; the lane line images in the training set are respectively input into the packet coding network and the example coding network for online deep mutual learning to obtain a global feature map output by the packet coding network and a local feature map output by the example coding network; the global feature map and the local feature map are fused to obtain a fused feature map; the fused feature map is respectively input into the packet coding network and the example coding network as a supervised learning signal, and the packet coding network and the example coding network are adjusted. When the packet coding network and the example coding network converge, a trained lane line recognition model is obtained; the trained lane line recognition model is evaluated using the test set, and if the evaluation passes, a trained lane line recognition model is obtained. In this scheme, the packet encoding network and the example encoding network are subjected to online deep mutual learning during the training process to improve the representation capability without manually adjusting parameters and solve the complex problem of lane line recognition; the complementarity between the two networks is fully utilized to reduce the complexity of the model and the difficulty of training; the fusion feature map is used in the online deep mutual learning to learn unknown prior knowledge, which is conducive to improving the detection accuracy and speed of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0061] Figure 1 An architecture diagram of a lane recognition model training system disclosed in an embodiment of the present invention;
[0062] Figure 2 A flowchart of a method for training a lane recognition model disclosed in an embodiment of the present invention;
[0063] Figure 3 A flowchart of a lane line recognition method disclosed in an embodiment of the present invention;
[0064] Figure 4 A structural diagram of a training device for a lane recognition model disclosed in an embodiment of the present invention;
[0065] Figure 5 A structural diagram of a lane line recognition device disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0066] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0067] In this application, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.
[0068] As can be seen from the background technology, in the prior art, lane line areas are segmented through edge detection, filtering and other technologies, which requires manual adjustment of parameters and filters, which is a large workload and leads to complex lane line recognition; the detection accuracy is improved by increasing the network depth and width of a single model, which makes the network structure of the model more and more complex, and the model is prone to overfitting and difficult to train, and usually requires more computing resources, which seriously restricts the practical value of the detection model; the existing network model cannot pay attention to the lines in the lane line image information in a differentiated manner, feature information cannot be fully extracted, and the performance is poor under severe occlusion conditions, and the prior knowledge of the lane line is not fully utilized, all of which will lead to a decrease in the detection accuracy and speed of the network model.
[0069] Therefore, the embodiments of the present invention disclose a training method and device for a lane line recognition model, and a lane line recognition method. In this scheme, the packet coding network and the example coding network are subjected to online deep mutual learning during the training process to improve the representation capability without manually adjusting parameters, thereby solving the complex problem of lane line recognition; the complementarity between the two networks is fully utilized to reduce the complexity of the model and the difficulty of training; the fused feature map is used in the online deep mutual learning to learn unknown prior knowledge, which is beneficial to improving the detection accuracy and speed of the model.
[0070] like Figure 1 As shown, it is an architecture diagram of a training system for a lane recognition model disclosed in an embodiment of the present invention. The training system includes: a packet coding network, an example coding network, an automatic fusion module and a Softmax classifier.
[0071] It should be noted that the packet encoding network and the example encoding network constitute the lane line recognition model.
[0072] During the training process, the lane line images are input into the packet coding network and the example coding network respectively for online deep mutual learning to obtain the global feature map output by the packet coding network and the local feature map output by the example coding network.
[0073] Among them, the packet encoding network and the example encoding network first perform linear projection processing on the lane line image.
[0074] It should be noted that linear projection is a concept in machine learning and mathematics. It refers to the process of mapping data from one space to another space through linear transformation. The information obtained after linear projection is the same thing in different dimensions as the previous input information. In mathematics, linear projection is a linear transformation that maps one vector to another vector, so that the target vector is the "shadow" or "projection" of the original vector in a specific direction. Linear projection is also an important component in deep learning design, which affects the performance and efficiency of the model in many ways. Through carefully designed linear projections, neural networks can better capture and utilize information in the data.
[0075] Then, the bag encoder network and the example encoder network enhance the feature representation through the attention mechanism to obtain the corresponding weighted feature maps, which involves paying attention to different parts of the image from different angles or scales.
[0076] In the process of online deep mutual learning, the complementarity between the two networks is fully utilized to mine the implicit lane line information, laying the foundation for improving the accuracy of lane line images; then, the lane line information implicit in the two networks is transferred to the automatic fusion module, feature fusion is performed, and the fused features are fed back to the packet coding network and the example coding network as supervised learning signals; then, an online deep mutual learning relationship is established between the packet coding network, the example coding network and the automatic fusion module. Through online deep mutual learning, the recognition performance after automatic fusion can be improved, and the detection performance of the packet coding network and the example coding network can also be improved.
[0077] The specific process of online deep mutual learning includes: calculating the KL divergence values of the packet coding network and the example coding network outputs respectively, obtaining the difference between the probability distributions of the two network outputs, and based on the difference between the probability distributions, using the Softmax activation function with temperature T to promote mutual learning between the output probabilities of the two networks, learning the spatial road information implicit in the packet coding network and the example coding network, fusing the global feature map output by the packet coding network and the local feature map output by the example coding network, using the fused feature map as a supervised learning signal, and using the cross entropy loss function to continuously adjust the training direction and speed of the packet coding network and the example coding network, so that the knowledge learned by the network model is richer and the ability of the network model to recognize lane lines is more accurate, thereby reversely supervising the online deep mutual learning of the packet coding network and the example coding network to form a closed-loop and effective learning process.
[0078] Whenever a fused feature map is obtained, the fused feature map is input into the Softmax classifier, the Softmax classifier is trained to generate lane line recognition results, and a trained Softmax classifier is obtained.
[0079] Based on the training system of a lane recognition model disclosed in the above embodiment of the present invention, Figure 2 FIG. 1 is a flowchart of a method for training a lane line recognition model disclosed in an embodiment of the present invention, which mainly includes the following steps:
[0080] Step S101: acquiring a data set including multiple lane line images.
[0081] In step S101, lane line images are collected from open source datasets on the Internet, including many complex scenes, such as S roads, Y lanes, as well as nighttime and multi-lane scenes.
[0082] Step S102: preprocess each lane line image in the data set, and divide the preprocessed data set into a training set and a test set.
[0083] In the specific implementation process of step S102, interference information in each lane line image is removed; and / or, data enhancement processing is performed on each lane line image to obtain multiple new lane line images; and / or, image enhancement processing is performed on each lane line image to obtain a clear lane line image.
[0084] Specifically, the dataset is processed to remove interference information, non-lane line information in the image is filtered out, and important and valuable information is retained; due to the small number of samples, it is necessary to use methods such as rotation processing, upper and lower mirror processing, and shearing processing to enhance the data set; the blurred lane line image is denoised to improve the quality of the lane line image.
[0085] Finally, the preprocessed data set is randomly divided into a training set and a test set in a ratio of 7:3. The details are as follows:
[0086]
[0087]
[0088] Among them, x i represents the i-th lane line image, y i represents the lane line label corresponding to the i-th lane line image, and M represents the training set I train The number of lane line images in the test set I test The number of lane line images.
[0089] Step S103: construct a lane line recognition model to be trained.
[0090] The lane line recognition model to be trained is a multi-view coding network including: a packet coding network and an example coding network.
[0091] In the specific implementation process of step S103, a packet coding network with a multi-view attention mechanism is constructed, so that when the packet coding network receives a lane line image, a first weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a global feature map is extracted from the first weighted feature map.
[0092] It should be noted that the packet coding network with multi-view attention mechanism can divide the pre-processed lane line image into examples and form a package, assign the image label to the package, assign weights to each example from two perspectives, and aggregate the example-level features into the global features at the package level through the pooling strategy guided by the multi-view attention algorithm to obtain the first weighted feature map, so as to fully utilize the spatial information of the image to complete the lane line image recognition. The packet coding network focuses on capturing the long-range dependency of the lane line and global feature modeling. Long-range dependency refers to the ability of the packet coding network to capture the long-range dependency between different regions in the image.
[0093] The image label refers to the label that is pre-labeled for each lane line image in the training set according to the lane line type (it may be a double yellow line, a single yellow line, a dashed or real line, etc.).
[0094] The two perspectives are: Perspective 1 is the probability that an example belongs to a certain class, and Perspective 2 is the contribution of an example to the classification of a packet into a certain class. The formula is as follows:
[0095]
[0096]
[0097] Among them, X c For package-level images, X d is the example level image, , represents the collection, S is the number of categories, T is the number of examples, represents the probability that the jth example belongs to the i-th category, represents the contribution of the j-th example to the package classified as the i-th category, and σ is the probability calculation symbol, that is, the probability value that a package-level or example-level image may belong to a certain lane line category.
[0098] It should be noted that aggregating all the examples together gives the maximum probability that the lane line image is classified as a certain lane line.
[0099] An example encoding network with a multi-view attention mechanism is constructed, so that when the example encoding network receives a lane line image, it obtains a second weighted feature map based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and extracts a local feature map from the second weighted feature map.
[0100] It should be noted that the example encoding network with a multi-view attention mechanism forms example-level features after the preprocessed lane line image passes through the feature extractor, and obtains the second weighted feature map containing example-level features through the pooling strategy guided by the multi-view attention algorithm, thereby extracting the attention-weighted detail features (i.e., local features), enhancing the characterization of local and global information of the road lane line image, mining the important image features hidden in the middle of the lane line image, and accurately describing the lane line information.
[0101] Among them, the packet encoding network focuses on capturing the long-range dependencies and global feature modeling of lane lines, while the example encoding network focuses on the detail extraction and local feature modeling of lane lines. Therefore, the packet encoding network and the example encoding network complement each other.
[0102] In one embodiment, the local features obtained by the example coding network and the packet-level features generated during the training of the packet coding network (i.e., the intermediate features before the global features are formed) are subjected to Hadamard calculation to obtain an integral matrix, based on which a local feature map can be decomposed, and finally the example coding network outputs a local feature map to obtain a better local feature performance. The formula is as follows:
[0103]
[0104] in, The symbol for the Hadamard integral product, X p represents the integral matrix obtained after the product.
[0105] It should be noted that the integral matrix obtained here is used to decompose the eigenvectors and eigenvalues, which can be understood as decomposing the integral matrix through linear transformation and other methods to obtain a local feature map.
[0106] Finally, based on the packet encoding network and the example encoding network, the lane line recognition model to be trained is constructed.
[0107] Step S104: input the lane line images in the training set into the packet coding network and the example coding network for online deep mutual learning, and obtain the global feature map output by the packet coding network and the local feature map output by the example coding network.
[0108] In step S104, online deep mutual learning refers to online knowledge distillation of the packet encoding network and the example encoding network.
[0109] First, the training set is input into the packet coding network and the example coding network for mutual learning, making full use of the complementarity between the two networks to mine implicit lane line knowledge and increase the prior knowledge of lane lines, which can lay the foundation for improving the accuracy of lane line image classification in the case of poor line of sight; then, the lane line knowledge implicit in the two networks is transferred to the automatic fusion module, and the adaptive feature fusion operation is performed. After the fusion feature map is obtained, it is fed back to the packet coding network and the example coding network as a supervisory signal. Then, an online deep mutual learning relationship is established between the packet coding network and the example coding network. By performing online deep mutual learning, the classification performance after automatic fusion can be improved, and the classification performance of the packet coding network and the example coding network can also be improved.
[0110] In online deep mutual learning, the cross entropy loss of the packet encoding network and the example encoding network is first calculated, where the cross entropy loss calculation formula of the packet encoding network is as follows:
[0111]
[0112] Among them, y i represents the label of the lane line image, h represents the lane line category corresponding to the lane line image, and h i represents the i-th block marker, represents the probability output of the packet encoding network, N refers to the number of samples, and H refers to the number of lane line categories.
[0113] The cross entropy loss calculation formula for the example encoding network is as follows:
[0114]
[0115] Among them, y i represents the label of the lane line image, h represents the lane line category corresponding to the lane line image, and hi represents the i-th block marker, represents the probability output of the example encoding network, N refers to the number of samples, and H refers to the number of lane line categories.
[0116] It should be noted that the online deep mutual learning uses the cross entropy loss function to measure the difference between the prediction and the reality. The online deep mutual learning process is mainly as follows: first, the data set is input into the packet coding network and the example coding network for mutual learning, and the complementarity between the two networks is fully utilized to mine the implicit lane line knowledge. The two models learn from each other and increase the prior knowledge of the lane line, which can lay the foundation for improving the accuracy of lane line image classification in the case of poor vision; then, the lane line knowledge implicit in the two networks is transferred to the automatic fusion module, and the adaptive feature fusion operation is performed. The advantages of the two networks are integrated in the final fusion feature map, and the cross entropy loss function is used to continuously adjust the training direction and training speed of the packet coding network and the example coding network, so that the knowledge learned by the packet coding network and the example coding network is richer, and the ability of the packet coding network and the example coding network to identify lane lines is more accurate.
[0117] Then, the KL divergence (Kullback-Leiblerdivergence), also known as information gain or relative entropy, of the packet coding network and the example coding network outputs is calculated respectively. It can reflect the difference between the probability distributions of the two network outputs. Based on the difference between the probability distributions, the Softmax activation function with temperature T is used to promote mutual learning between the output probabilities of the two networks.
[0118] Specifically, the divergence calculation formula from the packet encoding network to the example encoding network is as follows:
[0119]
[0120] The divergence calculation formula from the example encoding network to the packet encoding network is as follows:
[0121]
[0122] Among them, the packet coding network probability distribution p 1 The calculation formula is as follows:
[0123]
[0124] in, represents the logit output of the packet encoding network, and T represents the temperature of knowledge distillation.
[0125] Among them, the example encoding network probability distribution p 2 The calculation formula is as follows:
[0126]
[0127] in, represents the logit output of the example encoding network, and T represents the temperature of knowledge distillation.
[0128] It should be noted that the meanings of the parameters in the formulas in the embodiments of the present invention can refer to each other.
[0129] Step S105: Fusing the global feature map and the local feature map to obtain a fused feature map.
[0130] In step S105, the feature maps of the last layer of feedforward network of the packet coding network and the example coding network are first extracted, and adaptive average pooling is performed on the two feature maps to match their sizes. The specific process is: the length and width of the feature map output by the last pooling layer of the packet coding network and the example coding network are set to 1.
[0131] In a specific implementation, the output parameters of the last pooling layer of the packet coding network and the last pooling layer of the example coding network are made consistent in advance, so that the size of the global feature map output by the packet coding network and the local feature map output by the example coding network are consistent.
[0132] In the specific implementation process of step S105, the global feature map and the local feature map are connected in series to obtain a connected feature map; and a 1×1 point-by-point convolution operation is performed on the connected feature map to obtain a fused feature map.
[0133] The size of the concatenated feature map is [1, 1, c 1 +c 2 ], these three values represent the length, width and number of channels respectively. The total number of channels C can be set adaptively as needed, as shown below:
[0134]
[0135] In the embodiment of the present invention, the rich complementary semantic information from the packet encoding network and the example encoding network is fully utilized, including the shape, length, depth, etc. of the lane line image. The new features generated by fusion can more accurately characterize the lane line image, laying an important foundation for improving the accuracy of lane line image classification.
[0136] Step S106: input the fused feature map as a supervised learning signal into the packet coding network and the example coding network respectively, adjust the packet coding network and the example coding network, and when the packet coding network and the example coding network converge, obtain the trained lane line recognition model.
[0137] In step S106, the feature map is fused as a supervised learning signal, and the cross-entropy loss function is used to continuously adjust the training direction and speed of the packet coding network and the example coding network, so that the knowledge learned by the network model is richer and the ability of the network model to identify lane lines is more accurate, thereby reversely supervising the online deep mutual learning of the packet coding network and the example coding network to form a closed-loop and effective learning process.
[0138] It should be noted that if the packet coding network and the example coding network have not converged, the lane line images in the training set will continue to be input into the packet coding network and the example coding network for online deep mutual learning until the packet coding network and the example coding network converge.
[0139] Step S107: Use the test set to evaluate the trained lane line recognition model. If the evaluation passes, a trained lane line recognition model is obtained.
[0140] In one embodiment, the fused feature map is input into the Softmax classifier to complete the lane line image classification and obtain the lane line recognition result until a trained Softmax classifier is obtained, that is, a trained Softmax classifier is obtained.
[0141] Correspondingly, the specific implementation process of using the test set to evaluate the trained lane recognition model is as follows:
[0142] For each lane line image in the test set, it is input into the trained lane line recognition model to obtain the global feature map and local feature map output by the trained lane line recognition model; the global feature map and the local feature map are fused to obtain the fused feature map to be identified; the fused feature map to be identified is input into the trained classifier to obtain the lane line recognition result; if the lane line recognition result is correct, it is determined that the trained lane line recognition model has passed the evaluation.
[0143] Based on the training method of a lane line recognition model disclosed in the above-mentioned embodiment of the present invention, in this scheme, the packet coding network and the example coding network are subjected to online deep mutual learning during the training process, so as to improve the representation capability without manually adjusting the parameters and solve the complex problem of lane line recognition; make full use of the complementarity between the two networks to reduce the complexity of the model and reduce the difficulty of training; use the fusion feature map in the online deep mutual learning to learn unknown prior knowledge, which is conducive to improving the detection accuracy and speed of the model.
[0144] Based on the training method of a lane line recognition model disclosed in the above embodiment of the present invention, Figure 3 FIG. 1 is a flowchart of a lane line recognition method disclosed in an embodiment of the present invention, comprising the following steps:
[0145] Step S201: Acquire a target image to be identified.
[0146] Step S202: Input the target image into a pre-built lane line recognition model to obtain a target local feature map and a target global feature map output by the lane line recognition model.
[0147] Among them, the lane line recognition model is obtained by online deep mutual learning between the packet encoding network and the example encoding network.
[0148] Step S203: Fusing the target local feature map and the target global feature map to obtain a target fused feature map.
[0149] In step S203, the fusion process of the target local feature map and the target global feature map is consistent with the fusion process of the local feature map and the global feature map during the lane line recognition model training process, and they can refer to each other.
[0150] Step S204: input the target fusion feature map into a pre-trained classifier to obtain a lane line recognition result of the target image.
[0151] In step S204, the pre-trained classifier is a Softmax classifier. For the specific training process, please refer to the above-mentioned embodiment of the present invention.
[0152] Based on a lane line recognition method disclosed in the above-mentioned embodiment of the present invention, in this scheme, a lane line recognition model obtained by online deep mutual learning of a packet coding network and an example coding network is used to perform lane line recognition. On the one hand, the representation ability of the packet coding network and the example coding network is improved to solve the complex problem of lane line recognition; on the other hand, the complementarity of the packet coding network and the example coding network is fully explored and utilized to improve the accuracy and speed of lane line recognition; then, the global features and local features output by the packet coding network and the example coding network are fused, and the fused features are input into a pre-trained classifier to achieve rapid detection and accurate recognition of lane lines.
[0153] Corresponding to the training method of a lane line recognition model disclosed in the above embodiment of the present invention, as Figure 4 As shown, it is a structural diagram of a training device for a lane recognition model disclosed in an embodiment of the present invention, and the device includes: a collection unit 401, a preprocessing unit 402, a construction unit 403, a mutual learning unit 404, a fusion unit 405, a reverse supervision unit 406 and an evaluation unit 407.
[0154] The acquisition unit 401 is used to acquire a data set including multiple lane line images.
[0155] The preprocessing unit 402 is used to preprocess each lane line image in the data set, and divide the preprocessed data set into a training set and a test set.
[0156] In one embodiment, the preprocessing unit 402 for preprocessing each lane line image in the data set is specifically used to:
[0157] Remove interference information in each lane line image; the interference information includes: non-lane line information; and / or, for each lane line image, perform data enhancement processing to obtain multiple new lane line images; the data enhancement processing includes: rotation processing, upper and lower mirror processing and shearing processing; and / or, for each lane line image, perform image enhancement processing to obtain a clear lane line image; the image enhancement processing includes: noise reduction processing.
[0158] The construction unit 403 is used to construct a lane line recognition model to be trained; the lane line recognition model to be trained includes: a packet encoding network and an example encoding network.
[0159] In one embodiment, the construction unit 403 is specifically configured to:
[0160] Construct a packet coding network with a multi-view attention mechanism, so that when the packet coding network receives a lane line image, a first weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a global feature map is extracted from the first weighted feature map;
[0161] Construct an example encoding network with a multi-view attention mechanism, so that when the example encoding network receives a lane line image, a second weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a local feature map is extracted from the second weighted feature map;
[0162] Based on the packet encoding network and the example encoding network, a lane line recognition model to be trained is constructed.
[0163] The mutual learning unit 404 is used to input the lane line images in the training set into the packet coding network and the example coding network for online deep mutual learning, so as to obtain the global feature map output by the packet coding network and the local feature map output by the example coding network.
[0164] The fusion unit 405 is used to fuse the global feature map and the local feature map to obtain a fused feature map.
[0165] In one embodiment, the output parameters of the last pooling layer of the packet coding network and the last pooling layer of the example coding network are consistent, so that the size of the global feature map output by the packet coding network and the local feature map output by the example coding network are consistent. Accordingly, the fusion unit 405 is specifically used to:
[0166] The global feature map and the local feature map are connected in series to obtain a connected feature map; a 1×1 point-by-point convolution operation is performed on the connected feature map to obtain a fused feature map.
[0167] The reverse supervision unit 406 is used to input the fused feature map as a supervised learning signal into the packet coding network and the example coding network respectively, adjust the packet coding network and the example coding network, and obtain the trained lane line recognition model when the packet coding network and the example coding network converge.
[0168] The evaluation unit 407 is used to evaluate the trained lane line recognition model using the test set. If the evaluation passes, a trained lane line recognition model is obtained.
[0169] In one embodiment, the training device for the lane recognition model further includes:
[0170] The classifier training unit is used to input the fused feature map into a preset classifier whenever a fused feature map is obtained, train the preset classifier to generate a lane line recognition result, and obtain a trained classifier.
[0171] Accordingly, the evaluation unit 407 is specifically configured to:
[0172] For each lane line image in the test set, input it into the trained lane line recognition model to obtain the global feature map and local feature map output by the trained lane line recognition model;
[0173] The global feature map and the local feature map are fused to obtain a fused feature map to be identified;
[0174] Input the fused feature map to be identified into the trained classifier to obtain the lane line recognition result;
[0175] If the lane line recognition result is correct, it is determined that the trained lane line recognition model has passed the evaluation.
[0176] Based on the training device of a lane line recognition model disclosed in the above-mentioned embodiment of the present invention, in this scheme, the packet coding network and the example coding network are subjected to online deep mutual learning during the training process, so as to improve the representation capability without manually adjusting the parameters and solve the complex problem of lane line recognition; make full use of the complementarity between the two networks to reduce the complexity of the model and the difficulty of training; use the fused feature map in the online deep mutual learning to learn unknown prior knowledge, which is beneficial to improve the detection accuracy and speed of the model.
[0177] Corresponding to the lane line recognition method disclosed in the above embodiment of the present invention, as Figure 5, which is a structural diagram of a lane line recognition device disclosed in an embodiment of the present invention, the device includes: an acquisition unit 501, an input unit 502, a fusion unit 503 and a classifier recognition unit 504.
[0178] The acquisition unit 501 is used to acquire a target image to be identified.
[0179] The input unit 502 is used to input the target image into the pre-built lane line recognition model to obtain the target local feature map and the target global feature map output by the lane line recognition model.
[0180] Among them, the lane line recognition model is obtained by online deep mutual learning between the packet encoding network and the example encoding network.
[0181] The fusion unit 503 is used to fuse the target local feature map and the target global feature map to obtain a target fused feature map.
[0182] Among them, the fusion process of the target local feature map and the target global feature map is consistent with the fusion process of the local feature map and the global feature map during the lane line recognition model training process, and they can refer to each other.
[0183] The classifier recognition unit 504 is used to input the target fusion feature map into a pre-trained classifier to obtain a lane line recognition result of the target image.
[0184] The pre-trained classifier is a Softmax classifier. For the specific training process, please refer to the above-mentioned embodiment of the present invention.
[0185] Based on a lane line recognition device disclosed in the above-mentioned embodiment of the present invention, in this scheme, a lane line recognition model obtained by online deep mutual learning of a packet coding network and an example coding network is used to perform lane line recognition. On the one hand, the representation ability of the packet coding network and the example coding network is improved to solve the complex problem of lane line recognition; on the other hand, the complementarity of the packet coding network and the example coding network is fully explored and utilized to improve the accuracy and speed of lane line recognition; then, the global features and local features output by the packet coding network and the example coding network are fused, and the fused features are input into a pre-trained classifier to achieve rapid detection and accurate recognition of lane lines.
[0186] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0187] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0188] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A training method for a lane recognition model, characterized in that: The method comprises: Obtain a data set including multiple lane line images; Preprocessing each lane line image in the data set, and dividing the preprocessed data set into a training set and a test set; Constructing a lane line recognition model to be trained; the lane line recognition model to be trained includes: a packet encoding network and an example encoding network; Inputting the lane line images in the training set into the packet coding network and the example coding network for online deep mutual learning, respectively, to obtain a global feature map output by the packet coding network and a local feature map output by the example coding network; Fusing the global feature map and the local feature map to obtain a fused feature map; Inputting the fused feature map as a supervised learning signal into the packet coding network and the example coding network respectively, adjusting the packet coding network and the example coding network, and obtaining a trained lane line recognition model when the packet coding network and the example coding network converge; The trained lane line recognition model is evaluated using the test set, and if the evaluation passes, a trained lane line recognition model is obtained.
2. The method according to claim 1, characterized in that The preprocessing of each lane line image in the data set includes: Removing interference information from each lane line image; the interference information includes: non-lane line information; and / or, For each lane line image, data enhancement processing is performed to obtain multiple new lane line images; the data enhancement processing includes: rotation processing, upper and lower mirror processing and shearing processing; and / or, For each lane line image, image enhancement processing is performed to obtain a clear lane line image; the image enhancement processing includes: noise reduction processing.
3. The method according to claim 1, characterized in that The step of constructing a lane line recognition model to be trained includes: Constructing a packet coding network with a multi-view attention mechanism, so that when the packet coding network receives the lane line image, a first weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a global feature map is extracted from the first weighted feature map; Constructing an example encoding network with the multi-view attention mechanism, so that when the example encoding network receives the lane line image, a second weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a local feature map is extracted from the second weighted feature map; Based on the packet encoding network and the example encoding network, a lane line recognition model to be trained is constructed.
4. The method according to claim 1, characterized in that The step of inputting the lane line images in the training set into the packet encoding network and the example encoding network for online deep mutual learning comprises: Inputting the lane line images in the training set into the packet coding network and the example coding network for training respectively, and using a cross entropy loss function to adjust the training direction and training speed of the packet coding network and the example coding network; Calculating the KL divergence value of the packet encoding network output and the KL divergence value of the example encoding network output respectively; Based on the KL divergence value output by the packet coding network, the KL divergence value output by the example coding network and a Softmax activation function, mutual learning between output probabilities of the packet coding network and the example coding network is performed.
5. The method according to claim 1, characterized in that The output parameters of the last pooling layer of the packet coding network and the last pooling layer of the example coding network are consistent, so that the size of the global feature map output by the packet coding network and the local feature map output by the example coding network are consistent. Accordingly, the global feature map and the local feature map are fused to obtain a fused feature map, including: Performing a serial operation on the global feature map and the local feature map to obtain a serial feature map; A 1×1 point-by-point convolution operation is performed on the series feature maps to obtain a fused feature map.
6. The method according to claim 1, characterized in that The evaluating the trained lane line recognition model by using the test set includes: For each lane line image in the test set, input it into the trained lane line recognition model to obtain a global feature map and a local feature map output by the trained lane line recognition model; The global feature map and the local feature map are fused to obtain a fused feature map to be identified; Inputting the fused feature map to be identified into the trained classifier to obtain a lane line recognition result; If the lane line recognition result is correct, it is determined that the trained lane line recognition model has passed the evaluation.
7. A training device for a lane recognition model, characterized in that: The device comprises: An acquisition unit, used for acquiring a data set including multiple lane line images; A preprocessing unit, used to preprocess each lane line image in the data set, and divide the preprocessed data set into a training set and a test set; A construction unit, used to construct a lane line recognition model to be trained; the lane line recognition model to be trained includes: a packet encoding network and an example encoding network; A mutual learning unit, used to input the lane line images in the training set into the packet coding network and the example coding network for online deep mutual learning, so as to obtain a global feature map output by the packet coding network and a local feature map output by the example coding network; A fusion unit, used for fusing the global feature map and the local feature map to obtain a fused feature map; a reverse supervision unit, configured to input the fused feature map as a supervised learning signal into the packet coding network and the example coding network respectively, adjust the packet coding network and the example coding network, and obtain a trained lane line recognition model when the packet coding network and the example coding network converge; An evaluation unit is used to evaluate the trained lane line recognition model using the test set, and if the evaluation passes, a trained lane line recognition model is obtained.
8. The device according to claim 7, characterized in that The preprocessing unit for preprocessing each lane line image in the data set is specifically used to: Removing interference information from each lane line image; the interference information includes: non-lane line information; and / or, For each lane line image, data enhancement processing is performed to obtain multiple new lane line images; the data enhancement processing includes: rotation processing, upper and lower mirror processing and shearing processing; and / or, For each lane line image, image enhancement processing is performed to obtain a clear lane line image; the image enhancement processing includes: noise reduction processing.
9. The device according to claim 7, characterized in that The building block is specifically used for: Constructing a packet coding network with a multi-view attention mechanism, so that when the packet coding network receives the lane line image, a first weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a global feature map is extracted from the first weighted feature map; Constructing an example encoding network with the multi-view attention mechanism, so that when the example encoding network receives the lane line image, a second weighted feature map is obtained based on the lane line image and a pooling strategy guided by the multi-view attention algorithm, and a local feature map is extracted from the second weighted feature map; Based on the packet encoding network and the example encoding network, a lane line recognition model to be trained is constructed.
10. A method for identifying lane lines, characterized in that: The method comprises: Acquire a target image to be identified; Inputting the target image into a pre-built lane line recognition model to obtain a target local feature map and a target global feature map output by the lane line recognition model; the lane line recognition model is obtained by online deep mutual learning between a packet coding network and an example coding network; Fusing the target local feature map and the target global feature map to obtain a target fused feature map; The target fusion feature map is input into a pre-trained classifier to obtain a lane line recognition result of the target image.