Neural network training method, lane line detection method, device and electronic device
By introducing multiple feature extraction networks and attention map generation networks into the neural network, and through loss calculation and parameter adjustment methods, the detection accuracy and speed problems in lane line detection are solved, achieving more efficient and accurate lane line detection.
Patent Information
- Application Number
- CN201910708803.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-08-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2039-08-01
AI Technical Summary
The prior art is difficult to achieve fast, stable and accurate detection in lane line detection, especially in the development of unmanned driving technology, which has a major demand for efficient lane line detection.
A neural network training method is proposed, including a task detection network and a plurality of first networks for feature extraction. Through steps such as feature extraction, attention map generation, loss calculation and network parameter adjustment, network parameters are optimized to improve detection accuracy.
Through this method, the network detection accuracy can be improved without increasing the training data, so that the neural network can perform more stably and accurately in lane line detection.
Smart Images

Figure CN112307850B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural networks, and in particular, to a neural network training method, a lane line detection method, an apparatus, and an electronic device. Background Art
[0002] Lane line detection has always been one of the core technologies in the field of driverless. Stably, accurately, especially quickly detecting lane lines is of great significance to the development of driverless technology. Summary of the Invention
[0003] Embodiments of the present invention provide a neural network training method, a lane line detection method, an apparatus, and an electronic device.
[0004] The technical solution of the embodiments of the present invention is implemented as follows:
[0005] Embodiments of the present invention provide a neural network training method,
[0006] The neural network includes a task detection network and N first networks for feature extraction, where N is an integer greater than or equal to 2; the method includes:
[0007] Performing feature extraction processing on a first input image through the nth first network to obtain a feature map corresponding to the nth first network; n is an integer greater than or equal to 1 and less than or equal to N; when n = 1, the first input image is a first sample image, and when n is an integer greater than 1 and less than or equal to N, the first input image is a feature map corresponding to the (n - 1)th first network;
[0008] Generating m attention maps respectively based on m feature maps among the N feature maps; m is less than or equal to N; the m attention maps are generated by respectively processing the m feature maps through m generation networks;
[0009] Determining a first loss based on the differences between the m attention maps;
[0010] The task detection network determines a detection result according to the feature map output by the Nth first network;
[0011] Determining a second loss based on the determined detection result and the annotation result in the first sample image;
[0012] Adjusting the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss.
[0013] In the above solution, the determining the first loss based on the differences between the m attention maps includes:
[0014] Determine the difference between the k-th attention map and the j-th attention map among the m attention maps, and determine a first loss based on the difference; j is an integer greater than or equal to 1 and less than m; k is an integer greater than j;
[0015] Adjusting the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss includes:
[0016] Adjust the network parameters of the first network and the generation network corresponding to the j-th attention map according to the first loss;
[0017] Adjust the network parameters of the N first networks and the task detection network according to the second loss.
[0018] In the above solution, determining the difference between the k-th attention map and the j-th attention map among the m attention maps, and determining a first loss based on the difference includes:
[0019] Respectively determine the Euclidean distance between the k-th attention map and the j-th attention map to obtain k - j Euclidean distances;
[0020] Determine a first loss based on the k - j Euclidean distances.
[0021] In the above solution, when the difference between k and j is greater than 1, determining the first loss based on the k - j Euclidean distances includes:
[0022] Perform a specific process on the k - j Euclidean distances, and determine a first loss based on the specific process result; wherein, the specific process includes: average process or weighted average process.
[0023] In the above solution, the task detection network is used for lane line detection, the task detection network includes a second network, and the annotation result of the first sample image includes the annotated lane line;
[0024] The task detection network determines a detection result according to the feature map output by the N-th first network, including:
[0025] The second network determines the lane line in the first sample image according to the feature map output by the N-th first network;
[0026] Determining the second loss based on the determined detection result and the annotation result in the first sample image includes:
[0027] Determine the second loss based on the lane line in the determined first image and the annotated lane line in the first sample image.
[0028] In the above solution, the task detection network further includes a third network;
[0029] The task detection network determines a detection result based on the feature map output by the Nth first network, and further includes:
[0030] The third network determines a feature vector representing the number of detected lane lines based on the feature map output by the Nth first network;
[0031] The method further includes: determining a third loss according to the feature vector and an indication vector of the number of lane lines corresponding to the first sample image; the indication vector of the number of lane lines is determined according to the lane lines marked in the first sample image;
[0032] The adjusting the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss includes:
[0033] Adjusting the network parameters of the N first networks, the second network, the third network, and the m generation networks according to the first loss, the second loss, and the third loss.
[0034] In the above solution, the neural network is obtained by the following steps:
[0035] Processing a second sample image by using an initial neural network to determine a detection result of the second sample image;
[0036] Adjusting the network parameters of the initial neural network according to the determined detection result of the second sample image and the annotation result of the second sample image until the detection accuracy of the initial neural network reaches a first preset threshold to obtain the neural network.
[0037] An embodiment of the present invention further provides a lane line detection method, and the method includes:
[0038] Detecting a road image by using a neural network to determine lane lines in the road image and / or determine a feature vector representing the number of lane lines in the road image, where the neural network is trained by using the neural network training method described in the embodiment of the present invention, and the task detection network in the neural network is used for lane line detection.
[0039] An embodiment of the present invention further provides a neural network training device, where the neural network includes a task detection network and N first networks for feature extraction, and N is an integer greater than or equal to 2;
[0040] The device includes:
[0041] A feature extraction module, configured to perform feature extraction processing on a first input image through an n-th first network to obtain a feature map corresponding to the n-th first network; n is an integer greater than or equal to 1 and less than or equal to N; when n = 1, the first input image is a first sample image, and when n is an integer greater than 1 and less than or equal to N, the first input image is a feature map corresponding to the (n - 1)-th first network;
[0042] A generation module, configured to generate m attention maps respectively based on m feature maps among the N feature maps; m is less than or equal to N; the m attention maps are generated based on the processing of the m feature maps by m generation networks respectively;
[0043] A first loss determination module, configured to determine a first loss based on the differences among the m attention maps;
[0044] A detection module, configured to determine a detection result by the task detection network according to the feature map output by the N-th first network;
[0045] A second loss determination module, configured to determine a second loss based on the determined detection result and the annotation result in the first sample image;
[0046] An adjustment module, configured to adjust the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss.
[0047] In the above solution, the first loss determination module is configured to determine the difference between the k-th attention map and the j-th attention map among the m attention maps, and determine the first loss based on the difference; j is an integer greater than or equal to 1 and less than m; k is an integer greater than j;
[0048] The adjustment module is configured to adjust the network parameters of the first network and the generation network corresponding to the j-th attention map according to the first loss, and adjust the network parameters of the N first networks and the task detection network according to the second loss.
[0049] In the above solution, the first loss determination module is configured to respectively determine the Euclidean distances between the k-th attention map and the j-th attention map to obtain k - j Euclidean distances; and determine the first loss based on the k - j Euclidean distances.
[0050] In the above solution, when the difference between k and j is greater than 1, the first loss determination module is configured to determine the first loss based on the k - j Euclidean distances, including: performing specific processing on the k - j Euclidean distances, and determining the first loss based on the specific processing result; wherein the specific processing includes: average processing or weighted average processing.
[0051] In the above solution, the neural network is applied to lane line detection, and the task detection network further includes a second network; the annotation result of the first sample image includes the annotated lane lines.
[0052] The detection module is configured to enable the second network to determine the lane lines in the first sample image according to the feature map output by the Nth first network.
[0053] The second loss determination module is configured to determine a second loss based on the lane lines in the determined first image and the lane lines annotated in the first sample image.
[0054] In the above solution, the task detection network further includes a third network; the device further includes a third loss determination module.
[0055] The detection module is further configured to enable the third network to determine a feature vector representing the number of detected lane lines according to the feature map output by the Nth first network.
[0056] The third loss determination module is configured to determine a third loss according to the feature vector and the indication vector of the number of lane lines corresponding to the first sample image; the indication vector of the number of lane lines is determined according to the lane lines annotated in the first sample image.
[0057] The adjustment module is configured to adjust the network parameters of the N first networks, the second network, the third network, and the m generation networks according to the first loss, the second loss, and the third loss.
[0058] In the above solution, the device further includes a training module, and the training module is configured to train to obtain the neural network by using the following steps:
[0059] Process a second sample image by using an initial neural network to determine the detection result of the second sample image; adjust the network parameters of the initial neural network according to the determined detection result of the second sample image and the annotation result of the second sample image until the detection accuracy of the initial neural network reaches a first preset threshold to obtain the neural network.
[0060] An embodiment of the present invention further provides a lane line detection device, and the detection device includes: a detection unit and a determination unit; wherein,
[0061] The detection unit is configured to detect a road image by using a neural network.
[0062] The determination unit is configured to determine the lane lines in the road image based on the detection result of the detection unit and / or determine a feature vector representing the number of lane lines in the road image.
[0063] Among them, the neural network is trained by using the neural network training method described in the embodiments of the present invention, and the task detection network in the neural network is used for lane line detection.
[0064] The embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the neural network training method described in the embodiments of the present invention are implemented; or when the program is executed by a processor, the steps of the lane line detection method described in the embodiments of the present invention are implemented.
[0065] The embodiments of the present invention further provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the neural network training method described in the embodiments of the present invention are implemented; or when the processor executes the program, the steps of the lane line detection method described in the embodiments of the present invention are implemented.
[0066] The neural network training method, lane line detection method, device and electronic device provided by the embodiments of the present invention, the neural network includes a task detection network and N first networks for feature extraction, where N is an integer greater than or equal to 2; the method includes: performing feature extraction processing on a first input image through the nth first network to obtain a feature map corresponding to the nth first network; n is an integer greater than or equal to 1 and less than or equal to N; when n = 1, the first input image is a first sample image, and when n is an integer greater than 1 and less than or equal to N, the first input image is the feature map corresponding to the (n - 1)th first network; respectively generating m attention maps based on m feature maps among the N feature maps; m is less than or equal to N; the m attention maps are generated based on the processing of the m feature maps by m generation networks respectively; determining a first loss based on the differences between the m attention maps, and the task detection network determines a detection result according to the feature map output by the Nth first network; determining a second loss based on the determined detection result and the annotation result in the first sample image; adjusting the network parameters of the N first networks, the m generation networks and the task detection network according to the first loss and the second loss. By adopting the technical solution of the embodiments of the present invention, first, an attention map is obtained by processing the feature map processed by the first network through a generation network to obtain more prominent local features; then, a first loss is determined based on the differences between the attention maps, and the network parameters of the neural network and the generation network are adjusted based on the first loss to guide the features learned by different first networks to other first networks, and based on this, the network parameters of the first network are adjusted, so that the features extracted by the first network can imitate each other, thereby improving the network detection accuracy without increasing the training data. Description of the Drawings
[0067] Figure 1 Flow schematic of the neural network training method according to an embodiment of the present invention Figure 1 ;
[0068] Figure 2 Flow schematic of the neural network training method according to an embodiment of the present invention Figure 2 ;
[0069] Figure 3a Data flow schematic diagram of the neural network training method according to an embodiment of the present invention;
[0070] Figure 3b is Figure 3a Schematic diagram of the attention map in
[0071] Figure 4 Flow schematic diagram three of the neural network training method according to an embodiment of the present invention;
[0072] Figure 5 Composition structure schematic of the neural network training device according to an embodiment of the present invention Figure 1 ;
[0073] Figure 6 Composition structure schematic of the neural network training device according to an embodiment of the present invention Figure 2 ;
[0074] Figure 7 Composition structure schematic diagram three of the neural network training device according to an embodiment of the present invention;
[0075] Figure 8 Hardware composition structure schematic diagram of the electronic device according to an embodiment of the present invention. Specific embodiments
[0076] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0077] An embodiment of the present invention provides a neural network training method. The neural network includes a task detection network and N first networks for feature extraction, where N is an integer greater than or equal to 2. Figure 1 Flow schematic of the neural network training method according to an embodiment of the present invention Figure 1 ; As Figure 1 shown, the method includes:
[0078] Step 101: Perform feature extraction processing on the first input image through the nth first network to obtain the feature map corresponding to the nth first network; n is an integer greater than or equal to 1 and less than or equal to N; when n = 1, the first input image is the first sample image, and when n is an integer greater than 1 and less than or equal to N, the first input image is based on the feature map corresponding to the (n - 1)th first network;
[0079] Step 102: Generate m attention maps respectively based on m feature maps among the N feature maps; m is less than or equal to N; the m attention maps are generated based on the processing of the m feature maps by m generation networks respectively;
[0080] Step 103: Determine a first loss based on the differences among the m attention maps;
[0081] Step 104: The task detection network determines a detection result according to the feature map output by the Nth first network; determine a second loss based on the determined detection result and the annotation result in the first sample image;
[0082] Step 105: Adjust the network parameters of the N first networks, the m generation networks and the task detection network according to the first loss and the second loss.
[0083] Among them, there is no chronological order between the execution of Step 103 and Step 104. Step 103 can be executed first and then Step 104, or Step 104 can be executed first and then Step 103, or Step 103 and Step 104 can be executed simultaneously.
[0084] In addition, the number of network parameters of the neural network in this embodiment is much smaller than that of the existing neural network (which can be called a large network). As an implementation manner, compared with the number of network parameters of the existing neural network (which can be called a large network), the number of network parameters of the neural network in this embodiment can be 50%, or even 20% of the number of network parameters of the large network. Among them, as an example, the number of network parameters of the neural network can be reflected by the number of network layers in the neural network; for example, if the number of network layers of the large network is 100 layers, then the number of network layers of the neural network (small network) in this embodiment can be 50 layers, or even 20 layers. That is, the embodiment of the present invention directly trains on a small network with a small number of network parameters, occupies less physical storage space, and greatly improves the calculation speed. It is particularly suitable for the driverless scenario that requires lane line detection, because in the driverless scenario, the lane line detection network is required to have a fast calculation speed, and the physical storage space of in-vehicle hardware is usually limited, and it is not easy to deploy a large network.
[0085] The neural network in this embodiment includes at least a task detection network and N first networks for feature extraction. As an implementation, the first network can be implemented by a convolutional network or by an encoder structure including convolutional layers. The task detection network is related to a set task and is used to obtain a detection result related to the task. As an implementation, the task detection network can perform saliency detection, classification, and matting; further, the task detection network can perform vehicle detection, lane line detection, etc. For example, if the task detection network is used for lane line detection, an image containing the lane line detection result can be output through the task detection network. In practical applications, the task detection network can also be implemented by a convolutional network.
[0086] Among them, the N first networks are connected in sequence, and the output data of the previous first network is used as the input data of the next first network; the task detection network is connected to the Nth first network, that is, the feature map output by the Nth first network is used as the input data of the task detection network. In this embodiment, after feature extraction processing by the first network, feature maps corresponding to each first network are obtained. The output data of each first network is called a feature map. Thus, N feature maps can be obtained through N first networks. The first input data of the first first network (that is, the data processed by the first first network) is called the first sample image, and the first input data of other first networks (that is, the 2nd to the Nth first networks) is the feature map output by the previous first network of this first network; for example, if the current first network is the nth first network, the previous first network is the (n - 1)th first network.
[0087] Assume that the number of first networks is 4. Then the first sample image is input into the first network 1 for feature extraction processing to obtain the feature map of the first network 1. The feature map of the first network 1 is input into the first network 2 for feature extraction, and so on, until the first network 4 outputs a feature map.
[0088] It should be noted that the feature maps in this embodiment refer to the feature maps obtained after feature extraction processing by the first network, and the feature maps obtained through different first networks are different.
[0089] In this embodiment, the attention map is obtained based on the feature map. It can be understood that in this embodiment, there are m generation networks for processing the feature map to generate corresponding attention maps, that is, the generation networks further learn specific and local knowledge (features) in a self-learning manner to obtain the attention maps. Wherein, m is less than or equal to N. As an implementation manner, m is equal to N, that is, the feature maps output by each first network are input to the corresponding generation network to generate attention maps, and a total of N attention maps are generated. As another implementation manner, m is less than N, then the feature maps output by some first networks are input to the corresponding generation network to generate attention maps, and the total number of generated attention maps is less than N.
[0090] Wherein, the generation network may include at least one convolutional layer; by processing the feature map through the at least one convolutional layer, on the one hand, the features in the feature map can be further extracted through the convolutional processing of the at least one convolutional layer; on the other hand, since the sizes and channel numbers of the feature maps obtained by the feature extraction processing of different first networks may be different, therefore, through the processing of the feature map by the at least one convolutional layer of the generation network, the adjustment of the channel (channel) number and the image size is realized, so that the channel numbers and image sizes of the obtained m attention maps are the same, which is convenient for comparing the differences between the attention maps.
[0091] In this embodiment, the local features in the attention map are more prominent than the local features in the feature map. For example, the local features that are easily noticed in the attention map can be lane lines. It can be understood that the process of feature extraction of the feature map by the generation network mainly performs feature extraction processing on the local features in the feature map. As an implementation manner, after the feature map is processed by the generation network to obtain the attention map, the local features corresponding to the lane lines in the attention map are more prominent than the local features corresponding to the lane lines in the feature map.
[0092] As an example, generating an attention map based on a feature map may include: performing at least one convolutional processing on the feature map based on a generation network, and adjusting the channel number and image size of the feature map to obtain processed multi-channel data; processing the corresponding multi-channel data according to pixel points, and the processing method may include one of the following: summation processing, average processing, maximum value processing, etc.; generating an attention map based on the processed data.
[0093] It can be understood that different feature maps are processed by the first network different times, and the attention maps generated according to different feature maps are also different.
[0094] In other embodiments of the present invention, the image obtained by processing the feature map through the generation network may also be a Saliency Map or a Probability Map.
[0095] In an alternative embodiment of the present invention, determining the first loss based on the differences between the m attention maps includes: determining the difference between the k-th attention map and the j-th attention map among the m attention maps, and determining the first loss based on the difference; j is an integer greater than or equal to 1 and less than m; k is an integer greater than j.
[0096] In this embodiment, as an implementation manner, the value of k can be j + 1, that is, by determining the difference between the (j + 1)-th attention map and the j-th attention map, the difference between the (j + 1)-th attention map and the j-th attention map is determined, that is, the knowledge of the (j + 1)-th attention map is transmitted to the j-th attention map to obtain the first loss.
[0097] As another implementation manner, the value of k is greater than j + 1. For example, when the value of j is 1 and the value of k is 4, for the knowledge (or features) transmission of the first attention map, the differences between the second attention map and the first attention map, the third attention map and the first attention map, and the fourth attention map and the first attention map can be determined respectively. That is, the knowledge (or features) of the second attention map is transmitted to the first attention map, the knowledge (or features) of the third attention map is transmitted to the first attention map, and the knowledge (or features) of the fourth attention map is transmitted to the first attention map, so that the first attention map obtains the knowledge (or features) of the second attention map, the third attention map, and the fourth attention map, and the first loss is determined based on the transmitted knowledge (or features).
[0098] It can be seen that in the embodiments of the present invention, the features learned by different first networks are guided to other first networks, and based on this, the network parameters of the first network are adjusted, that is, the knowledge distillation method is adopted so that the features extracted by the first network can imitate each other, and the network detection accuracy is improved without increasing the training data.
[0099] Among them, as an example, the specific form of the first loss can be represented by the following function expression:
[0100]
[0101] Among them, A m and A m+1They are the feature maps output by the m-th first network and the (m + 1)-th first network respectively. Ψ(·) represents the generation network proposed in this embodiment. Taking the example that each feature map is input into the corresponding generation network to generate an attention map, Ψ(A m ) and Ψ(A m+1 ) represent the m-th and (m + 1)-th attention maps respectively. M represents the number of first networks.
[0102] In an alternative embodiment of the present invention, determining the difference between the k-th attention map and the j-th attention map among the m attention maps, and determining a first loss based on the difference includes: respectively determining the Euclidean distance between the k-th attention map and the j-th attention map to obtain k - j Euclidean distances; determining the first loss based on the k - j Euclidean distances.
[0103] Wherein, when the difference between k and j is greater than 1, determining the first loss based on the k - j Euclidean distances includes: performing a specific process on the k - j Euclidean distances, and determining the first loss based on the result of the specific process; wherein, the specific process includes: average process or weighted average process.
[0104] In this embodiment, the first loss between two attention maps can be obtained by calculating the Euclidean distance between the two attention maps. As an implementation manner, when the difference between k and j is 1, that is, in this implementation manner, only the knowledge (or features) of the (j + 1)-th attention map is passed to the j-th attention map, by calculating the Euclidean distance between the (j + 1)-th attention map and the j-th attention map, the first loss between the (j + 1)-th attention map and the j-th attention map is determined based on the Euclidean distance.
[0105] As another implementation manner, when the difference between k and j is greater than 1, that is, when the knowledge (or features) of multiple attention maps after the j-th attention map is passed to the j-th attention map, by calculating the Euclidean distance between the k-th attention map and the j-th attention map respectively to obtain k - j Euclidean distances, processing is performed by summing and averaging the k - j Euclidean distances, or performing a weighted average process, etc., and the first loss is determined based on the processing result.
[0106] In this embodiment, when the neural network has not been trained to a converged state, that is, when the accuracy of the detection result obtained by the task detection network does not meet the preset requirements, the detection result obtained by the task detection network is not accurate. Therefore, it is necessary to determine the second loss based on the detection result determined by the task detection network and the annotation result in the first sample image, and adjust the network parameters of the N first networks, the m generation networks, and the task detection network respectively based on the first loss and the second loss.
[0107] In an alternative embodiment of the present invention, the task detection network is used for lane line detection. The task detection network includes a second network, and the annotation result of the first sample image includes the annotated lane lines. The task detection network determines the detection result according to the feature map output by the Nth first network, including: the second network determines the lane lines in the first sample image according to the feature map output by the Nth first network; the determining the second loss based on the determined detection result and the annotation result in the first sample image includes: determining the second loss based on the determined lane lines in the first image and the annotated lane lines in the first sample image.
[0108] In this embodiment, the lane line annotation result in the first sample image can be represented by a lane line binary image. The pixel points of the lane line part in the lane line binary image are 1, and the pixel points of the remaining background part are 0. Then, the second network obtains an image containing lane line markings according to the feature map output by the Nth first network, and determines the second loss based on the image containing lane line markings and the lane line binary image. Specifically, the Euclidean distance between the lane lines in the two images can be calculated, and the second loss is determined based on the calculated Euclidean distance.
[0109] In this embodiment, the loss of the neural network includes two parts: the first loss and the second loss. As an example, the loss of the neural network can be expressed by the following expression:
[0110]
[0111] where represents the second loss, L distill (A m ,A m+1 ) represents the first loss, and β represents the weight coefficient. Among them, the specific form of the first loss can be referred to as described above and will not be elaborated here.
[0112] In one implementation manner, adjusting the network parameters of the N first networks, the second network, and the m generation networks according to the first loss and the second loss may include: performing a derivative processing on the network parameters related to the second loss to obtain a first derivative processing result of the network parameters corresponding to the second loss, and adjusting the network parameters in the second network based on the first derivative processing result. This implementation manner is applicable to the case where the network parameters are only associated with the second loss and not with the first loss, and then the network parameters are the network parameters in the second network.
[0113] Alternatively, performing a derivative processing on the network parameters related to the first loss to obtain a second derivative processing result of the network parameters corresponding to the first loss, and adjusting the network parameters in the generation network based on the second derivative processing result. This implementation manner is applicable to the case where the network parameters are only associated with the first loss and not with the second loss, and then the network parameters are the network parameters in the generation network.
[0114] In another implementation manner, adjusting the network parameters of the N first networks, the second network, and the m generation networks according to the first loss and the second loss may include: performing a derivative processing on each network parameter related to the first loss and the second loss to obtain a first derivative processing result of the network parameter corresponding to the second loss and a second derivative processing result of the network parameter corresponding to the first loss, performing a weighted summation processing on the first derivative processing result and the second derivative processing result (which can refer to expression (2)), and adjusting the network parameters of the first network according to the processing result. This implementation manner is applicable to the case where the network parameters are associated with both the first loss and the second loss, and then the network parameters are the network parameters in the first network.
[0115] In an alternative embodiment of the present invention, adjusting the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss includes: adjusting the network parameters of the first network and the generation network corresponding to the j-th attention map according to the first loss; adjusting the network parameters of the N first networks and the task detection network according to the second loss.
[0116] In this embodiment, since the attention map is obtained based on the feature map, and the feature map is obtained through the feature extraction process of the first network, that is, the attention map is related to the network parameters of the first network and the generation network. Based on this, the first loss determined based on the differences between the m attention maps in this embodiment is related to the network parameters of the first network and the generation network. Therefore, the network parameters of the first network and the generation network corresponding to the j-th attention map are adjusted according to the first loss. The second loss is related to the network parameters of the N first networks and the task detection network (i.e., the second network). Therefore, the network parameters of the N first networks and the task detection network are adjusted according to the second loss.
[0117] Among them, for example, reference can be made to Figure 3a , if the first loss 1 is determined based on two attention maps output by the generation network 1 and the generation network 2, then the first loss 1 is used to adjust the network parameters of the generation network 1 and the first network 1.
[0118] In step 105 of this embodiment, the convergence condition for adjusting the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss is that the loss converges to a preset value, that is, when the sum of the obtained first loss and the second loss converges to a stable value, the training is terminated. At this time, the trained neural network is obtained, and the neural network reaches a convergent state.
[0119] In this embodiment, whether it is two adjacent attention maps or two non-adjacent attention maps, the subsequent attention map is based on a feature map corresponding to the first network that has been processed more times than the feature map corresponding to the previous attention map. For example Figure 3a as shown in, the attention map output by the generation network 4 is obtained by the generation network 4 processing the feature map output by the first network 4, that is, the feature map output by the first network 4 has undergone the feature extraction processes of the first network 1 to the first network 4, while the attention map output by the generation network 1 is obtained by the generation network 1 processing the feature map output by the first network 1, that is, the feature map output by the first network 1 has only undergone the feature extraction process of the first network 1.
[0120] In the above process, Figure 3a vertically speaking, the local knowledge extracted by the generation network is more than the local knowledge extracted by the corresponding first network; Figure 3aLooking horizontally, the features (or knowledge) extracted by the first network 4 are more than those extracted by the first network 1, that is, the features (or knowledge) extracted by the first network with a later ranking are more than those extracted by the first network with an earlier ranking. The first network with a relatively later ranking (such as the first network 4) can be called a deep network, and the first network with a relatively earlier ranking (such as the first network 1) can be called a front-end network.
[0121] Adopting the technical solution of the embodiment of the present invention, first, the attention map is obtained by processing the feature map processed by the first network through the generation network to obtain more significant local features; then, the first loss is determined based on the difference between the attention maps, and the network parameters of the neural network and the generation network are adjusted based on the first loss to guide the features learned by different first networks to other first networks, and based on this, the network parameters of the first network are adjusted, so that the features extracted by the first network can imitate each other, so that without increasing the training data, the lane line detection accuracy equivalent to that of a large network can be achieved, and the network parameters are small, so it brings a faster calculation speed and a smaller storage space.
[0122] The embodiment of the present invention also provides a neural network training method; the neural network includes a task detection network and N first networks for feature extraction, where N is an integer greater than or equal to 2, and the task detection network is used for lane line detection; the task detection network includes a second network and a third network. Figure 2 Schematic flow of the neural network training method according to the embodiment of the present invention Figure 2 ; Figure 3a Schematic data flow diagram of the neural network training method according to the embodiment of the present invention; Figure 3b is Figure 3a schematic diagram of the attention map in Figure 2 、 Figure 3a and Figure 3b As shown, the method includes:
[0123] Step 202: Perform feature extraction processing on the first input image through the nth first network to obtain the feature map corresponding to the nth first network; n is an integer greater than or equal to 1 and less than or equal to N; when n = 1, the first input image is the first sample image, and when n is an integer greater than 1 and less than or equal to N, the first input image is the feature map corresponding to the (n - 1)th first network;
[0124] Step 203: Generate m attention maps respectively based on m feature maps among the N feature maps; m is less than or equal to N;
[0125] Step 204: Determine the first loss based on the difference between the m attention maps;
[0126] Step 205: The second network determines the lane lines in the first sample image according to the feature map output by the Nth first network;
[0127] Step 206: Determine a second loss based on the lane lines in the determined first image and the lane lines annotated in the first sample image, where the first sample image includes annotated lane lines;
[0128] Step 207: The third network determines a feature vector representing the number of detected lane lines according to the feature map output by the Nth first network;
[0129] Step 208: Determine a third loss according to the feature vector and the indication vector of the number of lane lines corresponding to the first sample image; the indication vector of the number of lane lines is determined according to the lane lines annotated in the first sample image;
[0130] Step 209: Adjust the network parameters of the N first networks, the second network, the third network, and the m generation networks according to the first loss, the second loss, and the third loss.
[0131] For the detailed description of steps 202 to 206 in this embodiment, reference may specifically be made to the detailed description of steps 101 to 104 in the foregoing embodiment, which will not be elaborated here.
[0132] In this embodiment, the neural network is applied to lane line detection. The neural network at least includes a task detection network and N first networks for feature extraction. The task detection network includes a second network and a third network. As an implementation manner, the first network may be implemented by a convolutional network. The second network is used to obtain a first image containing lane line markings. In practical applications, the second network may also be implemented by a convolutional network. The third network is used to obtain a feature vector representing whether lane lines are detected, that is, the second network and the third network are related to the set task. Among them, the first network may be referred to as the backbone network, and the second network and the third network may be referred to as branch networks.
[0133] Among them, the second network and the third network are respectively connected to the Nth first network, that is, the feature maps output by the Nth first network are respectively used as the input data of the second network and the third network. The output data of the second network is a first image containing lane line markings. The output data of the third network is a feature vector representing the number of detected lane lines.
[0134] In the case where the neural network has not been trained to a converged state, that is, when the accuracy rate of the labeled lane lines in the obtained first image does not reach a preset threshold and / or the accuracy rate of the feature vector representing the number of detected lane lines does not reach another preset threshold, it is necessary to determine a second loss based on the first sample image including the labeled lane lines and the first image containing lane line markings, and determine a third loss based on the indication vector of the number of lane lines corresponding to the first sample image and the feature vector representing the number of detected lane lines. Then, based on the first loss, the second loss, and the third loss, the network parameters of the N first networks, the second network, the third network, and the m generation networks are adjusted. That is, in this embodiment, the loss of the neural network includes three parts: the first loss, the second loss, and the third loss. Among them, as an example, the value of each vector component in the feature vector representing the number of detected lane lines can be 0 or 1, and the number of vector components with a value of 1 in the feature vector can represent the number of detected lane lines. For example Figure 3a as shown in, the obtained feature vector includes 3 vector components with a value of 1, which can represent that there are three lane lines in the figure.
[0135] As an example, the loss of the neural network can be expressed by the following expression:
[0136]
[0137] Among them, represents the third loss, represents the second loss, L distill (A m ,A m+1 ) represents the first loss, and both α and β represent weight coefficients. Among them, the specific form of the first loss can refer to that described in the foregoing embodiment and will not be elaborated here.
[0138] In this embodiment, the third loss is a loss related to the network parameters of the N first networks and the third network, the second loss is a loss related to the network parameters of the N first networks and the second network, and the first loss is a loss related to the network parameters of the N first networks and the m generation networks. Then, adjusting the network parameters of the N first networks, the second network, the third network, and the m generation networks according to the first loss, the second loss, and the third loss may include: performing a derivative operation on each network parameter related to the loss to obtain a first derivative operation result of the network parameter corresponding to the second loss, and adjusting the network parameter in the second network based on the first derivative operation result. This implementation is applicable to the case where the network parameter is only associated with the second loss and not with the first loss and the third loss. Then, the network parameter is the network parameter in the second network. Alternatively, performing a derivative operation on the network parameter related to the loss to obtain a second derivative operation result of the network parameter corresponding to the first loss, and adjusting the network parameter in the generation network based on the second derivative operation result. This implementation is applicable to the case where the network parameter is only associated with the first loss and not with the second loss and the third loss. Then, the network parameter is the network parameter in the generation network. Alternatively, performing a derivative operation on the network parameter related to the loss to obtain a third derivative operation result of the network parameter corresponding to the third loss, and adjusting the network parameter in the third network based on the third derivative operation result. This implementation is applicable to the case where the network parameter is only associated with the third loss and not with the second loss and the first loss. Then, the network parameter is the network parameter in the third network.
[0139] Alternatively, adjusting the network parameters of the N first networks, the second network, the third network, and the m generation networks according to the first loss, the second loss, and the third loss may include: performing a derivative operation on each network parameter related to the loss to obtain a third derivative operation result of the network parameter corresponding to the third loss, a first derivative operation result of the network parameter corresponding to the second loss, and a second derivative operation result of the network parameter corresponding to the first loss, performing a weighted summation operation on the third derivative operation result, the first derivative operation result, and the second derivative operation result, and adjusting the network parameter of the first network according to the operation result. This implementation is applicable to the case where the network parameter is associated with the first loss, the second loss, and the third loss. Then, the network parameter is the network parameter in the first network.
[0140] An embodiment of the present invention further provides a neural network training method. Figure 4 It is a schematic flowchart III of the neural network training method according to the embodiment of the present invention; asFigure 4 As shown in the foregoing embodiment, on the basis of the foregoing embodiment, before performing step 202, the method further includes:
[0141] Step 200: Process the second sample image using the initial neural network to determine the detection result of the second sample image;
[0142] Step 201: Adjust the network parameters of the initial neural network according to the determined detection result of the second sample image and the annotation result of the second sample image until the detection accuracy of the initial neural network reaches a first preset threshold to obtain the neural network.
[0143] In this embodiment, the initial neural network has a network architecture including N first networks and a task detection network. Among them, the task detection network may include a second network and / or a third network. Of course, according to the set task content, the task detection network may include other networks corresponding to the task content.
[0144] In an alternative embodiment of the present invention, the process of using the initial neural network to process the second sample image to determine the detection result of the second sample image includes: using the initial neural network to process the second sample image to determine the lane lines in the second sample image, and / or determining a feature vector representing the number of detected lane lines in the second sample image;
[0145] The adjustment of the network parameters of the initial neural network according to the determined detection result of the second sample image and the annotation result of the second sample image includes: adjusting the network parameters of the initial neural network according to the determined lane lines in the second sample image and the annotated lane lines in the second sample image, and / or according to the determined feature vector representing the number of detected lane lines in the second sample image and the indication vector of the number of lane lines corresponding to the second sample image until the detection accuracy of the initial neural network for lane line detection reaches a first preset threshold to obtain the neural network; the indication vector of the number of lane lines corresponding to the second sample image is determined according to the lane lines annotated in the second sample image.
[0146] Among them, the first preset threshold may be 90%, and of course it may also be other values. The fact that the detection accuracy of the initial neural network for lane line detection reaches the first preset threshold indicates that the initial neural network is close to the convergence state, that is, the neural network starting from the execution of step 202 is a neural network close to the convergence state, so that in the subsequent process of executing steps 202 to 209, the features reflected in the generated attention map are more accurate and the knowledge distillation is more efficient.
[0147] Optionally, adjusting the network parameters of the initial neural network according to the lane lines in the determined second sample image and the labeled lane lines in the second sample image, and / or according to the feature vector representing the number of detected lane lines in the determined second sample image and the indication vector of the number of lane lines corresponding to the second sample image, includes: determining a fourth loss according to the lane lines in the determined second sample image and the labeled lane lines in the second sample image, and / or determining a fifth loss according to the feature vector representing the number of detected lane lines in the determined second sample image and the indication vector of the number of lane lines corresponding to the second sample image; adjusting the network parameters of the initial neural network according to the fourth loss and / or the fifth loss.
[0148] Wherein, adjusting the network parameters of the initial neural network according to the fourth loss and / or the fifth loss includes: adjusting the network parameters of the N first networks and the second network in the initial neural network according to the fourth loss, and / or adjusting the network parameters of the N first networks and the third network in the initial neural network according to the fifth loss.
[0149] Wherein, the form of the fourth loss can refer to the description of the second loss in the foregoing embodiments, and the form of the fifth loss can refer to the description of the third loss in the foregoing embodiments, which will not be elaborated here.
[0150] An embodiment of the present invention further provides a lane line detection method, the method includes: using a neural network to detect a road image, determining the lane lines in the road image, and / or determining a feature vector representing the number of lane lines in the road image, wherein, the neural network is trained by using the neural network training method described in the foregoing embodiments of the present invention, and the task detection network in the neural network is used for lane line detection.
[0151] In this embodiment, the neural network has a network architecture including N first networks and a task detection network, wherein, the task detection network may include a second network and / or a third network, of course, according to the set task content, the task detection network may include other networks corresponding to the task content.
[0152] In an optional embodiment of the present invention, the using a neural network to detect a road image, determining the lane lines in the road image, and / or determining a feature vector representing the number of lane lines in the road image, includes: the first network performs feature extraction processing on the road image to obtain a feature map; the second network determines the lane lines in the road image according to the feature map, and / or, the third network determines a feature vector representing the number of detected lane lines according to the feature map.
[0153] In this embodiment, the neural network for detecting lane lines does not include the m generation networks. That is, the neural network includes N first networks and a second network, or includes N first networks and a third network, or includes N first networks, a second network, and a third network. The neural network processes the road image to obtain the lane lines in the road image and / or obtain a feature vector corresponding to the road image representing the detected number of lane lines.
[0154] An embodiment of the present invention also provides a neural network training device. Figure 5 The structural schematic diagram of the neural network training device according to the embodiment of the present invention Figure 1 ; as Figure 5 shown, the device is used to train a neural network; the neural network includes a task detection network and N first networks for feature extraction, where N is an integer greater than or equal to 2; the device includes:
[0155] A feature extraction module 41, configured to perform feature extraction processing on a first input image through the nth first network to obtain a feature map corresponding to the nth first network; n is an integer greater than or equal to 1 and less than or equal to N; when n = 1, the first input image is a first sample image, and when n is an integer greater than 1 and less than or equal to N, the first input image is the feature map corresponding to the (n - 1)th first network;
[0156] A generation module 42, configured to generate m attention maps respectively based on m feature maps among the N feature maps; m is less than or equal to N; the m attention maps are generated based on the processing of the m feature maps by m generation networks respectively;
[0157] A first loss determination module 43, configured to determine a first loss based on the differences between the m attention maps;
[0158] A detection module 44, configured to determine a detection result by the task detection network according to the feature map output by the Nth first network;
[0159] A second loss determination module 45, configured to determine a second loss based on the determined detection result and the annotation result in the first sample image;
[0160] An adjustment module 46, configured to adjust the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss.
[0161] In an optional embodiment of the present invention, the first loss determination module 43 is configured to determine the difference between the kth attention map and the jth attention map among the m attention maps, and determine a first loss based on the difference; j is an integer greater than or equal to 1 and less than m; k is an integer greater than j;
[0162] The adjustment module 46 is configured to adjust the network parameters of the first network and the generation network corresponding to the j-th attention map according to the first loss, and adjust the network parameters of the N first networks and the task detection network according to the second loss.
[0163] In an alternative embodiment of the present invention, the first loss determination module 43 is configured to respectively determine the Euclidean distances between the k-th attention map and the j-th attention map to obtain k-j Euclidean distances; and determine the first loss based on the k-j Euclidean distances.
[0164] Optionally, the first loss determination module 43 is configured to determine the first loss based on the k-j Euclidean distances when the difference between k and j is greater than 1, including: performing a specific process on the k-j Euclidean distances, and determining the first loss based on the result of the specific process; wherein the specific process includes: average process or weighted average process.
[0165] In an alternative embodiment of the present invention, the task detection network is applied to lane line detection, and the task detection network includes a second network 1; the annotation result of the first sample image includes the annotated lane lines.
[0166] The detection module 44 is configured to enable the second network to determine the lane lines in the first sample image according to the feature map output by the N-th first network.
[0167] The second loss determination module 45 is configured to determine the second loss based on the lane lines in the determined first image and the lane lines annotated in the first sample image.
[0168] In an alternative embodiment of the present invention, as Figure 6 shown, the task detection network further includes a third network; the device further includes a third loss determination module 47;
[0169] The detection module 44 is further configured to enable the third network to determine a feature vector representing the number of detected lane lines according to the feature map output by the N-th first network.
[0170] The third loss determination module 47 is configured to determine the third loss according to the feature vector and the indication vector of the number of lane lines corresponding to the first sample image; the indication vector of the number of lane lines is determined according to the lane lines annotated in the first sample image.
[0171] The adjustment module 46 is configured to adjust the network parameters of the N first networks, the second network, the third network, and the m generation networks according to the first loss, the second loss, and the third loss.
[0172] In an alternative embodiment of the present invention, as Figure 7 shown, the device further includes a training module 48, and the training module 48 is used to train and obtain the neural network by the following steps:
[0173] Process the second sample image using the initial neural network to determine the detection result of the second sample image; according to the determined detection result of the second sample image and the annotation result of the second sample image, adjust the network parameters of the initial neural network until the detection accuracy of the initial neural network reaches a first preset threshold to obtain the neural network.
[0174] In the embodiments of the present invention, the feature extraction module 41, the generation module 42, the first loss determination module 43, the detection module 44, the second loss determination module 45, the adjustment module 46, the third loss determination module 47, and the training module 48 in the network training device can all be implemented by a central processing unit (CPU, Central Processing Unit), a digital signal processor (DSP, Digital Signal Processor), a microcontroller unit (MCU, Microcontroller Unit), or a field-programmable gate array (FPGA, Field-Programmable Gate Array) in practical applications.
[0175] It should be noted that when the neural network training device provided in the above embodiment performs neural network training, only the above-mentioned division of each program module is used for illustration. In practical applications, the above-mentioned processing can be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-mentioned processing. In addition, the neural network training device provided in the above embodiment and the embodiment of the neural network training method belong to the same concept. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.
[0176] The embodiments of the present invention also provide a lane line detection device. The detection device includes: a detection unit and a determination unit; wherein,
[0177] The detection unit is used to detect a road image using a neural network;
[0178] The determination unit is used to determine the lane lines in the road image based on the detection result of the detection unit, and / or determine a feature vector representing the number of lane lines in the road image;
[0179] Among them, the neural network is trained by the neural network training method described in the embodiments of the present invention, and the task detection network in the neural network is used for lane line detection.
[0180] In the embodiments of the present invention, the detection unit and the determination unit in the lane line detection device can both be implemented by a CPU, a DSP, an MCU or an FPGA in practical applications.
[0181] It should be noted that when the lane line detection device provided in the above embodiments performs lane line detection, only the division of the above program modules is used for illustration. In practical applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-described processing. In addition, the lane line detection device provided in the above embodiments and the embodiments of the lane line detection method belong to the same concept. For the specific implementation process, please refer to the method embodiments, which will not be elaborated here.
[0182] The embodiments of the present invention also provide an electronic device. Figure 8 is a schematic diagram of the hardware composition structure of the electronic device according to the embodiments of the present invention. As Figure 8 shown, the electronic device includes a memory 52, a processor 51, and a computer program stored on the memory 52 and executable on the processor 51. When the processor 51 executes the program, it implements the steps of the neural network training method according to the embodiments of the present invention; or, when the processor executes the program, it implements the steps of the lane line detection method according to the embodiments of the present invention.
[0183] It can be understood that the various components in the electronic device are coupled together through a bus system 53. The bus system 53 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 53 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 8 all kinds of buses are labeled as the bus system 53.
[0184] It can be understood that the memory 52 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM).The memory 52 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memories.
[0185] The method disclosed in the above embodiments of the present invention can be applied to the processor 51 or implemented by the processor 51. The processor 51 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 51 or instructions in the form of software. The above-mentioned processor 51 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 51 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory 52. The processor 51 reads the information in the memory 52 and combines its hardware to complete the steps of the foregoing method.
[0186] In an exemplary embodiment, the electronic device can be implemented by one or more application-specific integrated circuits (ASICs, Application Specific Integrated Circuit), DSPs, programmable logic devices (PLDs, Programmable Logic Device), complex programmable logic devices (CPLDs, Complex Programmable Logic Device), field-programmable gate arrays (FPGAs, Field-Programmable Gate Array), general-purpose processors, controllers, microcontroller units (MCUs, Micro Controller Unit), microprocessors (Microprocessor), or other electronic components, and is used to execute the foregoing method.
[0187] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the neural network training method described in the embodiments of the present invention; or, when the program is executed by a processor, it implements the steps of the lane line detection method described in the embodiments of the present invention.
[0188] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0189] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0190] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.
[0191] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage media include various media that can store program codes, such as removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0192] Alternatively, if the above-mentioned integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention essentially or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The foregoing storage media include various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
[0193] As described above, it is only the specific implementation manner of the present invention. However, the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claimed rights.
Claims
1. A neural network training method, characterized in that, the neural network includes a task detection network and N first networks for feature extraction, where N is an integer greater than or equal to 2; the method includes: performing feature extraction processing on the first input image through the nth first network to obtain the feature map corresponding to the nth first network; n is an integer greater than or equal to 1 and less than or equal to N; when n = 1, the first input image is the first sample image, and when n is an integer greater than 1 and less than or equal to N, the first input image is the feature map corresponding to the (n - 1)th first network; generating m attention maps respectively based on m feature maps among the N feature maps; m is less than or equal to N; the m attention maps are generated based on the processing of the m feature maps by m generation networks respectively; determining a first loss based on the differences between the m attention maps; the task detection network determines a detection result according to the feature map output by the Nth first network; the task detection network is a detection network related to a set task, and the detection result is a detection result related to the task; determining a second loss based on the determined detection result and the annotation result in the first sample image; the annotation result is annotation information related to the task; adjusting the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss.
2. The method according to claim 1, characterized in that, the determining the first loss based on the differences between the m attention maps includes: determining the difference between the kth attention map and the jth attention map among the m attention maps, and determining the first loss based on the difference; j is an integer greater than or equal to 1 and less than m; k is an integer greater than j; the adjusting the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss includes: adjusting the network parameters of the first network and the generation network corresponding to the jth attention map according to the first loss; adjusting the network parameters of the N first networks and the task detection network according to the second loss.
3. The method according to claim 2, characterized in that, the determining the difference between the kth attention map and the jth attention map among the m attention maps, and determining the first loss based on the difference includes: respectively determining the Euclidean distances between the kth attention map and the jth attention map to obtain k - j Euclidean distances; determining the first loss based on the k - j Euclidean distances.
4. The method according to claim 3, characterized in that, in the case where the difference between k and j is greater than 1, the determining the first loss based on the k - j Euclidean distances includes: performing specific processing on the k - j Euclidean distances, and determining the first loss based on the specific processing result; where the specific processing includes: average processing or weighted average processing.
5. The method according to claim 1, characterized in that, The task detection network is used for lane line detection. The task detection network includes a second network. The annotation result of the first sample image includes the annotated lane lines; The task detection network determines the detection result according to the feature map output by the Nth first network, including: The second network determines the lane lines in the first sample image according to the feature map output by the Nth first network; Determining the second loss based on the determined detection result and the annotation result in the first sample image includes: Determining the second loss based on the determined lane lines in the first image and the annotated lane lines in the first sample image.
6. According to the method described in claim 5, wherein, The task detection network further includes a third network; The task detection network determines the detection result according to the feature map output by the Nth first network, and further includes: The third network determines a feature vector representing the number of detected lane lines according to the feature map output by the Nth first network; The method further includes: Determining a third loss according to the feature vector and the indication vector of the number of lane lines corresponding to the first sample image; the indication vector of the number of lane lines is determined according to the lane lines annotated in the first sample image; Adjusting the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss includes: Adjusting the network parameters of the N first networks, the second network, the third network, and the m generation networks according to the first loss, the second loss, and the third loss.
7. According to the method described in any one of claims 1 to 6, wherein, The neural network is obtained by the following steps: Processing a second sample image by using an initial neural network to determine the detection result of the second sample image; Adjusting the network parameters of the initial neural network according to the determined detection result of the second sample image and the annotation result of the second sample image until the detection accuracy of the initial neural network reaches a first preset threshold to obtain the neural network.
8. A lane line detection method, wherein, The method includes: Detecting a road image by using a neural network to determine the lane lines in the road image and / or determine a feature vector representing the number of lane lines in the road image, wherein the neural network is trained by using the method described in any one of claims 1-7, and the task detection network in the neural network is used for lane line detection.
9. A neural network training device, wherein, The neural network includes a task detection network and N first networks for feature extraction, N being an integer greater than or equal to 2; the device includes: A feature extraction module, configured to perform feature extraction processing on a first input image through an n-th first network to obtain a feature map corresponding to the n-th first network; n is an integer greater than or equal to 1 and less than or equal to N; when n = 1, the first input image is a first sample image, and when n is an integer greater than 1 and less than or equal to N, the first input image is a feature map corresponding to the (n - 1)-th first network; A generation module, configured to generate m attention maps respectively based on m feature maps among the N feature maps; m is less than or equal to N; the m attention maps are generated based on the processing of the m feature maps by m generation networks respectively; A first loss determination module, configured to determine a first loss based on the differences between the m attention maps; A detection module, configured to determine a detection result by the task detection network according to the feature map output by the N-th first network; the task detection network is a detection network related to a set task, and the detection result is a detection result related to the task; A second loss determination module, configured to determine a second loss based on the determined detection result and the annotation result in the first sample image; the annotation result is annotation information related to the task; An adjustment module, configured to adjust the network parameters of the N first networks, the m generation networks, and the task detection network according to the first loss and the second loss.
10. The apparatus according to claim 9, wherein, the first loss determination module is configured to determine the difference between the k-th attention map and the j-th attention map among the m attention maps, and determine the first loss based on the difference; j is an integer greater than or equal to 1 and less than m; k is an integer greater than j; the adjustment module is configured to adjust the network parameters of the first network and the generation network corresponding to the j-th attention map according to the first loss, and adjust the network parameters of the N first networks and the task detection network according to the second loss.
11. The apparatus according to claim 10, wherein, the first loss determination module is configured to respectively determine the Euclidean distances between the k-th attention map and the j-th attention map to obtain k - j Euclidean distances; and determine the first loss based on the k - j Euclidean distances.
12. The apparatus according to claim 11, wherein, when the difference between k and j is greater than 1, the first loss determination module is configured to determine the first loss based on the k - j Euclidean distances, including: performing specific processing on the k - j Euclidean distances, and determining the first loss based on the specific processing result; wherein the specific processing includes: average processing or weighted average processing.
13. The apparatus according to claim 9, wherein, the task detection network is applied to lane line detection, and the task detection network includes a second network; the annotation result of the first sample image includes the annotated lane lines; the detection module is configured to determine the lane lines in the first sample image by the second network according to the feature map output by the N-th first network; The second loss determination module is configured to determine a second loss based on the lane lines in the determined first image and the lane lines annotated in the first sample image.
14. The apparatus according to claim 13, wherein, the task detection network further includes a third network; the apparatus further includes a third loss determination module; the detection module is further configured to enable the third network to determine a feature vector representing the number of detected lane lines according to the feature map output by the Nth first network; the third loss determination module is configured to determine a third loss according to the feature vector and the indication vector of the number of lane lines corresponding to the first sample image; the indication vector of the number of lane lines is determined according to the lane lines annotated in the first sample image; the adjustment module is configured to adjust the network parameters of the N first networks, the second network, the third network, and the m generation networks according to the first loss, the second loss, and the third loss.
15. The apparatus according to any one of claims 9 to 14, wherein, the apparatus further includes a training module, and the training module is configured to train the neural network by using the following steps: Process a second sample image by using an initial neural network to determine the detection result of the second sample image; adjust the network parameters of the initial neural network according to the determined detection result of the second sample image and the annotation result of the second sample image until the detection accuracy of the initial neural network reaches a first preset threshold to obtain the neural network.
16. A lane line detection apparatus, wherein, the detection apparatus includes: a detection unit and a determination unit; wherein, the detection unit is configured to detect a road image by using a neural network; the determination unit is configured to determine the lane lines in the road image based on the detection result of the detection unit, and / or determine a feature vector representing the number of lane lines in the road image; wherein, the neural network is trained by using the method according to any one of claims 1-7, and the task detection network in the neural network is used for lane line detection.
17. A computer-readable storage medium, on which a computer program is stored, wherein, when the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7; or when the program is executed by a processor, it implements the steps of the method according to claim 8.
18. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7; or when the processor executes the program, it implements the steps of the method according to claim 8.
Citation Information
Patent Citations
Key point detection method, neural network training method, devices and electronic apparatus
CN108229490A
Neural network for drawing multi-label identification and related method, media and device
CN109754015A