Road surface identification identification method and road surface identification identification device

The color and shape information of pavement marks are processed through the dual convolutional neural network, which solves the problem of information loss in the prior art and improves the identification accuracy of pavement marks.

CN120359538APending Publication Date: 2025-07-22NISSAN MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280102705.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, when using a convolutional neural network to identify road surface marks, it is easy to lose information of slender areas, resulting in the problem of inconspicuous contours or the misidentification of lane marks of different colors in the distance.

Method used

A dual convolutional neural network is used for machine learning, which processes the color information and shape information of the image respectively, and selects the appropriate convolution operation through the output value of the convolution layer, and combines the color and shape information for road surface identification and recognition.

Benefits of technology

It effectively suppresses information loss in elongated areas on the image, and improves the accuracy of identification of lane markings with inconspicuous contours and distant colors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359538A_ABST
    Figure CN120359538A_ABST
Patent Text Reader

Abstract

A road surface identification method and a road surface identification device according to the present invention identify a road surface identification reflected in a first image of the surroundings of a vehicle captured by an imaging unit mounted on the vehicle. And for the relationship between the color information of the first image and the first label, performing first machine learning by using a first convolutional neural network, and for the relationship between the shape information of the first image and the first label, performing second machine learning by using a second convolutional neural network. While the first machine learning and the second machine learning are being performed, on the basis of an output value for each pixel constituting a convolutional layer of one of the first convolutional neural network and the second convolutional neural network, a convolution calculation for each pixel constituting a convolutional layer of the other is selected. And then identifying the road surface identifier by using at least one of the first convolutional neural network and the second convolutional neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a road sign recognition method and a road sign recognition device. Background Art

[0002] Currently, an intelligent driving control method has been proposed, which detects the driving environment of a vehicle, obtains the detection results of target objects of at least one category in the driving environment, triggers a driving warning based on whether the detection results meet the driving warning regulation conditions and whether the driving speed of the vehicle exceeds the speed threshold, improves the driving safety of the vehicle, and limits the number of warnings of low danger level and / or low urgency level according to the speed threshold (refer to Patent Document 1).

[0003] Prior Art Documents

[0004] Patent Documents

[0005] Patent Document 1: Japanese Patent Application Laid-Open No. 2021-504245

[0006] Problems to be Solved by the Invention

[0007] According to Patent Document 1, based on an image of the surroundings of a vehicle being taken, a convolutional neural network is used to identify lanes of multiple categories. However, due to the need to perform convolutional calculations in the convolutional neural network, the loss of information in narrow and long regions existing in the image is obvious. Therefore, there is a problem that traffic signs such as traffic signs with unclear contours and lanes of different colors in the distance are easily mis-identified. Summary of the Invention

[0008] The present invention has been completed in view of the above problems. An object of the present invention is to provide a road sign recognition method and a road sign recognition device that can suppress information loss in narrow and long regions existing in an image in the recognition of traffic signs reflected in an image using a convolutional neural network, and can suppress mis-identification of traffic signs such as traffic signs with unclear contours or lanes of different colors in the distance.

[0009] In order to solve the above problems, a road surface marking recognition method and a road surface marking recognition device according to an aspect of the present invention recognize road surface markings reflected in a first image of the surroundings of a vehicle captured by a photographing unit mounted on the vehicle. For the relationship between the color information of the first image and the first label, first machine learning is performed using a first convolutional neural network, and for the relationship between the shape information of the first image and the first label, second machine learning is performed using a second convolutional neural network. During the first machine learning and the second machine learning, based on the output value of each pixel of the convolutional layer constituting one of the first convolutional neural network and the second convolutional neural network, the convolution operation of each pixel of the convolutional layer constituting the other is selected. Then, at least one of the first convolutional neural network and the second convolutional neural network is used to recognize the road surface markings.

[0010] Advantages of the Invention

[0011] According to the present invention, in the recognition of traffic signs reflected in an image using a convolutional neural network, information loss in elongated regions present in the image can be suppressed, and misrecognition of traffic signs such as traffic signs with unclear contours and lanes of different colors in the distance can be suppressed. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a block diagram showing the structure of a road surface marking recognition device according to an embodiment of the present invention.

[0013] Figure 2 is a flowchart showing an example of the processing of a road surface marking recognition device according to an embodiment of the present invention.

[0014] Figure 3 is a block diagram showing the structure of a convolutional neural network in a road surface marking recognition device according to an embodiment of the present invention.

[0015] Figure 4A is a diagram showing an example of an image of the surroundings of a vehicle captured.

[0016] Figure 4B is a diagram showing an example of shape information generated based on an image.

[0017] Figure 4C is a diagram showing an example of the distribution of labels set for each pixel.

[0018] Figure 4D is a diagram showing a reference example of the distribution of labels set for each pixel. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Next, embodiments of the present invention will be described in detail with reference to the drawings. In the description, the same symbols are assigned to the same parts and repeated descriptions are omitted.

[0020] (Structure of Road Surface Sign Recognition Device)

[0021] Figure 1 is a block diagram showing the structure of the road surface sign recognition device of the present embodiment. As Figure 1 shown, the road surface sign recognition device of the present embodiment includes: an acquisition unit 71, a database 73, a controller 100, and an output unit 400. The controller 100 is connected to the acquisition unit 71, the database 73, and the output unit 400 through a wired or wireless communication line.

[0022] The acquisition unit 71 acquires teacher data for machine learning (or deep learning) in the controller 100. In particular, the acquisition unit 71 acquires a "teacher label image" as teacher data. The "teacher label image" refers to an image obtained by giving a first label for recognizing a road surface sign that has entered the first image to each pixel constituting the first image of the first image of the front of the vehicle captured by an imaging unit (such as a camera) mounted on the vehicle.

[0023] Various road surface signs can be cited as the road surface signs recognized by the first label. The road surface signs include various signs such as "white solid line", "white dotted line", "stop line", "pedestrian crossing", "zebra crossing", "arrow", "bicycle guiding line", "blue solid line", "yellow solid line", and "other" that is not classified into any of the above. In addition, the "arrow" in the road surface signs can include "inside the intersection", "go straight and turn left", "turn left", "go straight", "turn right", "go straight and turn right", etc. The road surface signs are not limited to the examples listed here.

[0024] In addition, the teacher label image can also be generated by so-called "annotation" in which a person visually confirms the first image and gives a first label to each pixel constituting the first image. Furthermore, the teacher label image can also be an image generated by performing a viewpoint transformation based on a virtual posture that the imaging unit may capture on the teacher label image generated by "annotation". That is, the teacher label image can be an image obtained by "virtually expanding" the teacher label image generated by "annotation".

[0025] The teacher label image acquired by the acquisition unit 71 is input to the controller 100 and the database 73. In addition, the acquisition unit 71 can also be a device that acquires a second image of the surroundings of the vehicle captured by the imaging unit.

[0026] The database 73 stores the teacher label image. In addition, various information generated by the controller 100 can also be stored. For example, the database 73 can store a learning model described later.

[0027] The output unit 400 outputs various information generated by the controller 100. For example, the output unit 400 can output the output from the learning model described later to the outside.

[0028] The controller 100 (an example of a control unit or a processing unit) is a general-purpose computer including a CPU (Central Processing Unit), a memory, and an input / output unit. A computer program (information processing program) for functioning as a part of the road surface marking recognition device is installed in the controller 100. By executing the computer program, the controller 100 functions as a plurality of information processing circuits (110, 120, 130, 140, 150, 160) included in the road surface marking recognition device.

[0029] In addition, an example in which a plurality of information processing circuits (110, 120, 130, 140, 150, 160) included in the road surface marking recognition device are implemented by software is shown here. However, dedicated hardware for executing each of the following information processing may be prepared to constitute the information processing circuits (110, 120, 130, 140, 150, 160). In addition, the plurality of information processing circuits (110, 120, 130, 140, 150, 160) may also be composed of separate hardware.

[0030] As the plurality of information processing circuits (110, 120, 130, 140, 150, 160), the controller 100 includes: a color information acquisition unit 110, a shape information acquisition unit 120, a first learning unit 130, a second learning unit 140, a convolution setting unit 150, and an evaluation unit 160.

[0031] The color information acquisition unit 110 acquires color information for each pixel constituting the first image in the teacher label image as the teacher data.

[0032] For example, when the first image is a color image, the color information of each pixel includes information on the brightness of red ("R"), the brightness of green ("G"), and the brightness of blue ("B") in the pixel. In addition, when the first image is a black-and-white or gray-scale image, the color information of each pixel may be a gray scale representing the light intensity in the pixel. The color information of each pixel is determined based on the output from an imaging element such as a CCD (charge-coupled device) or a CMOS (complementary metal oxide semiconductor) mounted on the imaging unit.

[0033] The shape information acquisition unit 120 acquires shape information for each pixel constituting the first image in the teacher label image as the teacher data.

[0034] The shape information may include the edge information of the road surface markings reflected in the first image. The shape information acquisition unit 120 may apply a Sobel filter to the first image to generate the edge information. Here, the Sobel filter refers to a spatial filter used in the field of image processing to detect the contours of the images contained in the image. In addition, the shape information acquisition unit 120 may also use a smoothing filter, a differential filter, a Prewitt filter, or a filter formed by combining them to generate the edge information.

[0035] The shape information may include the corner information of the road surface markings reflected in the first image. The shape information acquisition unit 120 may apply a Harris filter to the first image to generate the corner information. Here, the Harris filter is a spatial filter used in the field of image processing to detect the corners of the images contained in the image.

[0036] The first learning unit 130 performs first machine learning using a first convolutional neural network N1 composed of multiple convolutional layers. The first learning unit 130 performs first machine learning to output a first label for each pixel from the first convolutional neural network N1 into which the color information has been input. As a result of the first machine learning of the first learning unit 130, the parameters for defining the first convolutional neural network N1 can be set.

[0037] The second learning unit 140 performs second machine learning using a second convolutional neural network N2 composed of multiple convolutional layers. In addition, the number of convolutional layers of the second convolutional neural network N2 is the same as the number of convolutional layers of the first convolutional neural network N1. In addition, the size of each convolutional layer of the second convolutional neural network N2 is the same as the size of each convolutional layer of the first convolutional neural network N1. The second learning unit 140 performs second machine learning to output a first label for each pixel from the second convolutional neural network N2 into which the shape information has been input. As a result of the second machine learning of the second learning unit 140, the parameters for defining the second convolutional neural network N2 can be set.

[0038] Figure 3 The structure of the first convolutional neural network N1 and the structure of the second convolutional neural network N2 are shown. Figure 3 It is a block diagram showing the structure of the convolutional neural network in the road surface marking recognition device of the present embodiment.

[0039] As Figure 3As shown, the first convolutional neural network N1 is composed of "2×N" convolutional layers EL11 to EL1N and DL11 to DL1N. Here, N is an integer greater than or equal to 1. The convolutional layers EL11 to EL1N and DL11 to DL1N are each composed of multiple units. In addition, the first convolutional neural network N1 may include: a pooling layer, a batch normalization layer, a rectified linear unit layer, an upsampling layer, a softmax output layer, etc.

[0040] The units between the convolutional layers EL11 to EL1N and DL11 to DL1N are combined with each other. More specifically, in the first convolutional neural network N1, multiple convolutional layers are arranged in the order of convolutional layers EL11, EL12, EL13,..., EL1N, DL1N,..., DL13, DL12, DL11, and the units of adjacent convolutional layers are combined with each other.

[0041] In addition, through the combination of the units with each other, the set of signals received from the convolutional layers adjacent to the convolutional layer of interest is regarded as an "image". Processing the set of signals in the order of convolutional layers EL11, EL12, EL13,..., EL1N, DL1N,..., DL13, DL12, DL11 can be regarded as the "image" processed in the above order of convolutional layers being sent sequentially. Therefore, the units constituting each convolutional layer can be regarded as the pixels of the "image" input to and output from each convolutional layer (hereinafter, sometimes the "unit" and the "pixel" are not particularly distinguished and used).

[0042] Each unit has an activation function (for example, a sigmoid function, a rectified linear unit function, a softmax function, etc.). The sum of the weights is calculated based on multiple inputs to the unit, and the value of the activation function with the sum value as a variable becomes the output of the unit. For example, in the first machine learning, the weights when calculating the sum in each unit of the first convolutional neural network N1 are adjusted as parameters defining the first convolutional neural network N1.

[0043] The convolutional layers EL11 to EL1N in the first convolutional neural network N1 constitute an encoder, generating a low-dimensional latent representation of the set of color information A1 about the input (the set of color information about the pixels of the first image as a whole). In addition, the convolutional layers DL11 to DL1N constitute a decoder, generating a set of label information A3 based on the low-dimensional latent representation.

[0044] Through the first machine learning, the first convolutional neural network N1 becomes a learning model that represents the relationship between the color information attached to each pixel of the first image and the first label for each first image.

[0045] AsFigure 3 As shown, the second convolutional neural network N2 is composed of "2×N" convolutional layers EL21 to EL2N and DL21 to DL2N. Here, N is an integer greater than or equal to 1. The convolutional layers EL21 to EL2N and DL21 to DL2N are each composed of multiple units. In addition, the second convolutional neural network N2 may further include a pooling layer, a batch normalization layer, a rectified linear unit layer, an upsampling layer, a softmax output layer, etc.

[0046] The units between the convolutional layers EL21 to EL2N and DL21 to DL2N are combined with each other. More specifically, in the second convolutional neural network N2, multiple convolutional layers are arranged in the order of convolutional layers EL21, EL22, EL23, …, EL2N, DL2N, …, DL23, DL22, DL21, and the units of adjacent convolutional layers are combined with each other.

[0047] In addition, through the combination of the units with each other, the set of signals received from the convolutional layers adjacent to the convolutional layer of interest is regarded as an "image". Processing the set of signals in the order of convolutional layers EL21, EL22, EL23, ..., EL2N, DL2N, ..., DL23, DL22, DL21 can be regarded as the "image" processed in the order of the above convolutional layers being sequentially transmitted. Therefore, the units constituting each convolutional layer can be regarded as pixels of the "image" input to and output from each convolutional layer (hereinafter, sometimes the "unit" and the "pixel" are not particularly distinguished and used).

[0048] Each unit has an activation function (e.g., sigmoid function, rectified linear unit function, softmax function, etc.). Based on the sum of the weights calculated from multiple inputs to the unit, the value of the activation function with the sum value as a variable becomes the output of the unit. For example, in the second machine learning, the sum of the weights calculated in each unit of the second convolutional neural network N2 is adjusted as a parameter defining the second convolutional neural network N2.

[0049] The convolutional layers EL21 to EL2N in the second convolutional neural network N2 constitute an encoder, generating a low-dimensional latent representation of the set of shape information A2 (the set of shape information about the pixels of the entire first image). In addition, the convolutional layers DL21 to DL2N constitute a decoder, generating a set of label information A3 based on the low-dimensional latent representation.

[0050] Through the second machine learning, the second convolutional neural network N2 becomes a learning model representing the relationship between the shape information attached to each pixel constituting the first image and the first label for each first image.

[0051] When generating a learning model by performing first machine learning and second machine learning, color information or shape information is input to the input layer of the neural network constituting the learning model, and a probability set represented for each label is output from the output layer. The parameters of the learning model are adjusted so that the probability in the first label is maximized in the probability set represented for each label. In addition, when color information or shape information is input to the input layer and a label is output from the output layer, the parameters of the first convolutional neural network N1 and the second convolutional neural network N2 are adjusted so that the difference between the output label and the first label becomes smaller.

[0052] When generating a learning model by performing first machine learning and second machine learning, for example, the gradient descent method, the probabilistic gradient descent method, etc. can be used to minimize the error with respect to the output of the neural network. In addition, in order to perform gradient calculation using the gradient descent method and the probabilistic gradient descent method, the error backpropagation algorithm can also be used.

[0053] In machine learning based on neural networks, generalization performance (the ability to discriminate unknown data) and overfitting (a phenomenon in which it is suitable for data used to create a learning model but the generalization performance cannot be improved) may become problems.

[0054] Then, when generating a learning model by performing first machine learning and second machine learning, in order to alleviate overfitting, methods such as regularization that restricts the degree of freedom of weights during learning can be used. In addition, methods such as dropout that probabilistically selects units in the neural network and invalidates other units can also be adopted. In addition, in order to improve generalization performance, methods such as data regularization, data normalization, and data augmentation that eliminate biases in the data can also be adopted.

[0055] During the execution of the first machine learning and the second machine learning, the convolution setting unit 150 selects the convolution operation of each pixel constituting the convolution layer of the first convolutional neural network based on the output value of each pixel of the convolution layer constituting the second convolutional neural network. Similarly, during the execution of the first machine learning and the second machine learning, the convolution setting unit 150 selects the convolution operation of each pixel constituting the convolution layer of the second convolutional neural network based on the output value of each pixel of the convolution layer constituting the first convolutional neural network.

[0056] First, the "convolution operation" will be described. Assume that the convolutional layer CL1 of interest in a plurality of convolutional layers constituting the first convolutional neural network N1 or the second convolutional neural network N2 is composed of a set of signals with an "image" of "W1×W2" received from an adjacent convolutional layer CL2. The values of the signals included in this "image" are represented by xij (where i = 0, 1, …, W1−1, j = 0, 1, …, W2−1).

[0057] On the other hand, in order to define the operation of convolution, consider a "filter" composed of a set of signals "H1×H2". The size of this "filter" ("filter size") is "H1×H2". The values of the signals contained in this "filter" are represented by hpq (where p = 0, 1, ..., H1 - 1, q = 0, 1, ..., H2 - 1). "H1×H2" is an integer greater than or equal to 2.

[0058] In this case, the operation of convolution "uij" is defined between the above-mentioned "image" and "filter" by the following Equation 1.

[0059] (Equation 1)

[0060]

[0061] That is, each unit of the convolution layer CL1 of interest is connected to the "H1×H2" units of the adjacent convolution layer CL2. Defining the values "hpq" of the "filter" for each unit becomes a parameter that defines the convolutional neural network (the first convolutional neural network N1 or the second convolutional neural network N2) to which the convolution layer CL1 belongs.

[0062] The convolution setting unit 150 selects the convolution operation of each unit of the convolution layer that constitutes the first convolutional neural network N1 based on the output values of each unit of the convolution layer that constitutes the second convolutional neural network N2. For example, the convolution setting unit 150 selects the convolution operation of each unit of the convolution layer EL11 that constitutes the first convolutional neural network N1 based on the output values of each unit of the convolution layer EL21 that constitutes the second convolutional neural network N2.

[0063] Similarly, the convolution operations in each unit that constitutes the convolution layers EL12, EL13, …, EL1N, DL1N, …, DL13, DL12, DL11 are respectively selected based on the output values of each unit that constitutes EL22, EL23, …, EL2N, DL2N, …, DL23, DL22, DL21.

[0064] For example, during the first machine learning and the second machine learning, the convolution setting unit 150 performs convolution with a filter size of 1 in the units of the convolution layer of the first convolutional neural network N1 corresponding to the units in the convolution layer of the second convolutional neural network N2 that have output values above a specified threshold.

[0065] To explain more specifically. Consider the case where the convolution operation "uij" in the output of the convolution layer unit of the second convolutional neural network N2 is above a specified threshold. In this case, the output "uij" of the corresponding unit in the convolution layer of the first convolutional neural network N1 can be "xij" or a constant multiple thereof, rather than being determined by Equation 1.

[0066] On the other hand, when the convolution operation "uij" in the output of the unit in the convolution layer of the second convolutional neural network N2 is less than a specified threshold, the output "uij" of the corresponding unit in the convolution layer of the first convolutional neural network N1 is determined by Equation 1.

[0067] When comparing at the same convolutional layer level, the "image" flowing through the first convolutional neural network N1 and the "image" flowing through the second convolutional neural network N2 have the same size. Therefore, it can be said that the convolution setting unit 150 overlaps the "image" flowing in the second convolutional neural network N2 with the "image" flowing in the first convolutional neural network N1, and feeds back the shape information in the analysis of color information.

[0068] The evaluation unit 160 inputs an image to the first convolutional neural network N1 and performs semantic segmentation on the input image. That is, the evaluation unit 160 inputs the second image obtained by photographing the surroundings of the vehicle by the photographing unit to the first convolutional neural network N1.

[0069] The evaluation unit 160 can also obtain a second label for identifying the road surface markings reflected in the second image by calculating the output from the first convolutional neural network N1 into which the second image is input.

[0070] The evaluation unit 160 can also obtain a second label for each pixel constituting the second image by calculating the output from the first convolutional neural network N1 into which the second image is input, so as to identify the road surface markings reflected in the second image.

[0071] In addition, the photographing unit that photographed the second image can be the same as or different from the photographing unit that photographed the first image for generating the learning model. In addition, the vehicle with the area where the second image is photographed in the surroundings can be the same as or different from the vehicle with the area where the first image is photographed in the surroundings.

[0072] In addition, the evaluation unit 160 can be structured to output to the outside, via the output unit 400, the second label obtained for each pixel constituting the second image.

[0073] In addition, the evaluation unit 160 can also perform additional learning to determine whether the first convolutional neural network N1 and the second convolutional neural network N2 need to be updated. More specifically, the evaluation unit 160 can use the teacher label image to calculate the confusion matrix (TP / TN / FP / FN) of the result estimated for each pixel by the learning model and the first label (correct label) assigned to each pixel in the teacher label image. Here, TP (true positive) represents the number of pixels accurately estimated as the correct label. TN (true negative) represents the number of pixels accurately estimated as not the correct label. FP (false positive) represents the number of pixels erroneously estimated as the correct label. FN (false negative) represents the number of pixels erroneously estimated as not the correct label.

[0074] Then, the evaluation unit 160 can also calculate the precision, accuracy, recall, and F-measure (F-measure) of the learning model based on the confusion matrix. Here, the precision of the learning model is calculated by "(TP + TN) / (TP + TN + FP + FN)". Accuracy is an index for measuring the degree to which the label estimated as the correct label is actually the correct label, and is calculated by "TP / (TP + FP)". Recall is an index for measuring the degree to which the label actually being the correct label is estimated as the correct label, and is calculated by "TP / (TP + FN)".

[0075] The F-measure is calculated by the harmonic mean of accuracy and recall, "2 × accuracy × recall / (accuracy + recall)".

[0076] The evaluation unit 160 can specify the learning object label based on the calculated F-measure. The processes in the color information acquisition unit 110, the shape information acquisition unit 120, the first learning unit 130, the second learning unit 140, and the convolution setting unit 150 can be repeatedly executed until a learning model that meets the target estimation accuracy is obtained.

[0077] For example, the evaluation unit 160 can calculate the recognition rate when recognizing each road surface marking using at least one of the learned first convolutional neural network or the learned second convolutional neural network. Then, the evaluation unit 160 can extract the road surface markings related to the recognition rate less than the specified value as the target road surface markings. Furthermore, the first image in which the number of pixels constituting the target road surface marking is equal to or more than the specified number can be extracted as the target image, and the first machine learning and the second machine learning can be performed based on the target image.

[0078] (Processing example of the road surface marking recognition device)

[0079] Figure 2 It is a flowchart showing a processing example of the road surface marking recognition device according to the present embodiment. Figure 2 The processing of the shown road surface marking recognition device can be repeatedly executed at a specified cycle.

[0080] In step S101, the acquisition unit 71 acquires an unselected teacher label image as teacher data.

[0081] In step S103, the color information acquisition unit 110 acquires color information for each pixel constituting the first image based on the first image included in the teacher label image. In addition, the shape information acquisition unit 120 acquires shape information for each pixel constituting the first image based on the first image included in the teacher label image.

[0082] In step S105, the output value of each pixel of one of the first convolutional neural network N1 and the second convolutional neural network N2 is calculated. For example, the first learning unit 130 calculates the output value of each pixel of the first convolutional neural network N1. The second learning unit 140 calculates the output value of each pixel of the second convolutional neural network N2.

[0083] In step S107, the convolution setting unit 150 extracts the pixels in which the output value of the pixels of one of the convolutional neural networks is greater than or equal to a specified threshold.

[0084] In step S109, the convolution setting unit 150 sets the convolution operation for each pixel of the other convolutional neural network among the first convolutional neural network N1 and the second convolutional neural network N2.

[0085] Specifically, the convolution setting unit 150 sets a "filter" with a filter size of 1 in the pixels of the first convolutional neural network N1 corresponding to the pixels in which the output value of the pixels of the second convolutional neural network N2 is greater than or equal to the specified threshold. On the other hand, the convolution setting unit 150 sets a "filter" with a filter size greater than 1 in the pixels of the first convolutional neural network N1 corresponding to the pixels in which the output value of the pixels of the second convolutional neural network N2 is less than the specified threshold.

[0086] Alternatively, the convolution setting unit 150 sets a "filter" with a filter size of 1 in the pixels of the second convolutional neural network N2 corresponding to the pixels in which the output value of the pixels of the first convolutional neural network N1 is greater than or equal to the specified threshold. On the other hand, the convolution setting unit 150 sets a "filter" with a filter size greater than 1 in the pixels of the second convolutional neural network N2 corresponding to the pixels in which the output value of the pixels of the first convolutional neural network N1 is less than the specified threshold.

[0087] In step S111, the output value of each pixel of the other convolutional neural network among the first convolutional neural network N1 and the second convolutional neural network N2 is calculated. For example, the second learning unit 140 calculates the output value of each pixel of the second convolutional neural network N2. The first learning unit 130 calculates the output value of each pixel of the first convolutional neural network N1.

[0088] In step S113, the controller 100 determines whether there is an unselected teacher label image.

[0089] In the case where it is determined that there is an unselected teacher label image (when it is "Yes" in step S113), the process returns to step S101. On the other hand, in the case where it is determined that there is no unselected teacher label image (when it is "No" in step S113), the process proceeds to step S115.

[0090] In step S115, the first learning unit 130 sets the parameters of the first convolutional neural network N1 based on the output value of the first convolutional neural network N1. In addition, the second learning unit 140 sets the parameters of the second convolutional neural network N2 based on the output value of the second convolutional neural network N2. Then, the processing of the road surface marking recognition device shown ends. Figure 2 The processing of the road surface marking recognition device shown.

[0091] (Effect of the embodiment)

[0092] As described in detail above, the road surface marking recognition method and the road surface marking recognition device of the present embodiment recognize the road surface markings reflected in the first image of the surroundings of the vehicle captured by the imaging unit mounted on the vehicle. For the relationship between the color information of the first image and the first label, the first machine learning is performed using the first convolutional neural network, and for the relationship between the shape information of the first image and the first label, the second machine learning is performed using the second convolutional neural network. During the first machine learning and the second machine learning, based on the output value of each pixel of the convolutional layer constituting one of the first convolutional neural network and the second convolutional neural network, the convolution operation of each pixel of the convolutional layer constituting the other is selected. Then, at least one of the first convolutional neural network and the second convolutional neural network is used to identify the road surface markings.

[0093] Thereby, in the recognition of traffic signs in an input image using a convolutional neural network, it is possible to suppress information loss in the long and thin regions present in the image, and it is possible to suppress misrecognition of traffic signs such as traffic signs with unclear contours and lanes of different colors in the distance.

[0094] For example, consider Figure 4A the case of performing semantic segmentation on an image of the surroundings of the vehicle. The road surface marking recognition method and the road surface marking recognition device of the present embodiment input the color information obtained for each pixel of the image shown into the first convolutional neural network. Figure 4A Moreover, the road surface marking recognition method and the road surface marking recognition device of the present embodiment generate shape information based on the image shown.

[0095] And, Figure 4A shown in the image.Figure 4B is a diagram showing an example of shape information generated based on Figure 4A the image shown. In Figure 4B , the outline part of the road surface markings in the image Figure 4A projected onto is emphasized in white. The road surface marking recognition method and the road surface marking recognition device of the present embodiment input the shape information obtained for each pixel of the image into a second convolutional neural network.

[0096] Figure 4C shows the result of semantic segmentation of an image of the surroundings of a vehicle by the road surface marking recognition method and the road surface marking recognition device of the present embodiment. Figure 4C is a diagram showing an example of the distribution of labels set for each pixel.

[0097] In Figure 4C , it shows the case where road surface markings TG1 representing the white lines beside the road, road surface markings TG2 representing the dividing lines, and road surface markings TG3 representing arrows are detected with high precision and labels are set. By using the signal flowing through the second convolutional neural network for the selection of the convolutional operations performed in the first convolutional neural network, the shape information is fed back into the analysis of the color information. As a result, as Figure 4C shown, road surface markings can be detected with high precision.

[0098] For comparison with Figure 4C , in Figure 4D , the result is shown in the case where the signal flowing through the second convolutional neural network is not used for the selection of the convolutional operations performed in the first convolutional neural network. Figure 4D is a diagram showing a reference example of the distribution of labels set for each pixel.

[0099] Different from Figure 4C , in Figure 4D , road surface markings TG1, road surface markings TG2 representing the dividing lines, and road surface markings TG3 representing arrows cannot be detected with high precision. This is because information loss occurs at the boundary part of the road surface markings due to the convolutional operations performed in the first convolutional neural network.

[0100] According to the road surface marking recognition method and the road surface marking recognition device of the present embodiment, by selecting the results of the convolutional operations in each unit of the convolutional layer constituting the first convolutional neural network, information loss at the boundary part of the road surface markings is suppressed, and misrecognition of traffic signs with unclear outlines and traffic signs such as lanes of different colors in the distance is suppressed.

[0101] In addition, in the road surface marking recognition method and the road surface marking recognition device of the present embodiment, during the execution of the first machine learning and the second machine learning, convolution with a filter size of 1 can also be performed on the pixels in the other convolutional layer corresponding to the pixels in one convolutional layer having an output value equal to or greater than a specified threshold. Thereby, information loss at the boundary portion of the road surface marking can be suppressed. In particular, since the filter size is 1, blurring of the contour position during convolution can be suppressed.

[0102] Furthermore, in the road surface marking recognition method and the road surface marking recognition device of the present embodiment, during the execution of the first machine learning and the second machine learning, convolution with a filter size greater than 1 can also be performed on the pixels in the other convolutional layer corresponding to the pixels in one convolutional layer having an output value less than the specified threshold. Thereby, it is possible to capture the appearance in which the road surface marking continues in the portion outside the boundary of the road surface marking. As a result, the road surface marking reflected in the image can be detected with high accuracy.

[0103] In addition, in the road surface marking recognition method and the road surface marking recognition device of the present embodiment, the shape information may also include the edge information of the road surface marking reflected in the first image. Thereby, information loss at the edge portion of the road surface marking can be suppressed.

[0104] Furthermore, in the road surface marking recognition method and the road surface marking recognition device of the present embodiment, a Sobel filter can also be applied to the first image to generate edge information. Thereby, it is possible to accurately capture the edge portion of the road surface marking reflected in the first image and reflect the edge portion in the machine learning.

[0105] In addition, in the road surface marking recognition method and the road surface marking recognition device of the present embodiment, the shape information may also include the corner information of the road surface marking reflected in the first image. Thereby, information loss at the corner portion of the road surface marking can be suppressed.

[0106] Furthermore, in the road surface marking recognition method and the road surface marking recognition device of the present embodiment, a Harris filter can also be applied to the first image to generate corner information. Thereby, it is possible to accurately capture the corner portion of the road surface marking reflected in the first image and reflect the corner portion in the machine learning.

[0107] In addition, the road surface marking recognition method and the road surface marking recognition device of the present embodiment can also calculate the recognition rate when recognizing the road surface marking using at least one of the learned first convolutional neural network and the learned second convolutional neural network for each road surface marking. It is also possible to extract the road surface marking related to a recognition rate less than a specified value as the target road surface marking. It is also possible to extract the first image in which the number of pixels constituting the target road surface marking is equal to or greater than a specified number as the target image. The first machine learning and the second machine learning can be executed based on the target image.

[0108] Therefore, it is possible to re-learn road signs with low recognition rates as objects. In particular, since it is possible to extract images containing many pixels related to road signs with low recognition rates for the first machine learning and the second machine learning, the efficiency of re-learning is improved.

[0109] Each function represented in the above-described embodiment can be implemented by one or more processing circuits. The processing circuit includes a programmed processor or an electric circuit, etc., and further includes a device such as an application-specific integrated circuit (ASIC) and circuit components configured to execute the functions.

[0110] As described above, the content of the present invention has been described according to the embodiment, but the present invention is not limited to these descriptions. It is obvious to those skilled in the art that various modifications and improvements can be made. The discussions and drawings forming a part of this disclosure should not be construed as limiting the present invention. Through this disclosure, those skilled in the art will clearly understand various alternative embodiments, examples, and application techniques.

[0111] The present invention naturally includes various embodiments not described herein. Therefore, the technical scope of the present invention is determined by the specific matters of the invention covered by the scope of the appropriate claims only according to the above description.

[0112] Symbol Explanation

[0113] 71: Acquisition unit

[0114] 73: Database

[0115] 100: Controller

[0116] 110: Color information acquisition unit

[0117] 120: Shape information acquisition unit

[0118] 130: First learning unit

[0119] 140: Second learning unit

[0120] 150: Convolution setting unit

[0121] 160: Evaluation unit

[0122] 400: Output unit

[0123] N1: First convolutional neural network

[0124] N2: Second convolutional neural network

Claims

1. A road surface marking recognition method, a controller for controlling the recognition of a road surface marking in a first image that captures the surroundings of the vehicle taken by a camera unit mounted on the vehicle, characterized in that: The controller performs the following processing: Based on the first image, color information and shape information are obtained for each pixel constituting the first image. Using a first convolutional neural network composed of multiple convolutional layers, perform first machine learning on each pixel to output a first label for recognizing the road surface marking from the first convolutional neural network into which the color information is input, and Using a second convolutional neural network composed of multiple convolutional layers, where the number of convolutional layers and the size of each convolutional layer are the same as those of the first convolutional neural network, perform second machine learning on each pixel to output the first label from the second convolutional neural network into which the shape information is input. During the first machine learning and the second machine learning, based on the output value of each pixel of the convolutional layer constituting one of the first convolutional neural network and the second convolutional neural network, select the convolution operation of each pixel of the convolutional layer constituting the other of the first convolutional neural network and the second convolutional neural network. Use at least one of the learned first convolutional neural network with parameters adjusted based on the first machine learning or the learned second convolutional neural network with parameters adjusted based on the second machine learning to recognize the road surface marking.

2. The road surface marking recognition method according to claim 1, characterized in that: During the first machine learning and the second machine learning, the controller performs convolution with a filter size of 1 on the pixels of the other convolutional layer corresponding to the pixels in the one convolutional layer having an output value above a specified threshold.

3. The road surface marking recognition method according to claim 1 or 2, characterized in that: During the first machine learning and the second machine learning, the controller performs convolution with a filter size greater than 1 on the pixels of the other convolutional layer corresponding to the pixels in the one convolutional layer having an output value less than the specified threshold.

4. The road surface marking recognition method according to any one of claims 1 to 3, characterized in that: The shape information includes edge information of the road surface marking reflected in the first image.

5. The road surface marking recognition method according to any one of claims 1 to 4, characterized in that: The shape information includes corner information of the road surface marking reflected in the first image.

6. The road surface marking recognition method according to any one of claims 1 to 5, characterized in that: The controller performs the following processing: Calculate the recognition rate when recognizing the road surface marking using at least one of the learned first convolutional neural network or the learned second convolutional neural network for each road surface marking. Extract the road surface marking related to the recognition rate less than a specified value as the target road surface marking. Extract the first image in which the number of pixels constituting the target road surface marking is equal to or more than a specified number as the target image. Execute the first machine learning and the second machine learning based on the object image.

7. A road surface marking recognition device, comprising an acquisition unit and a controller, characterized in that the acquisition unit acquires a first image of the surroundings of the vehicle captured by an imaging unit mounted on the vehicle, the controller performs the following processing: Based on the first image, color information and shape information are acquired for each pixel constituting the first image, Using a first convolutional neural network composed of a plurality of convolutional layers, perform first machine learning on each pixel so that a first label for recognizing a road surface marking reflected in the first image is output from the first convolutional neural network into which the color information is input, Using a second convolutional neural network composed of a plurality of convolutional layers, the number of convolutional layers and the size of each convolutional layer of which are the same as those of the first convolutional neural network, perform second machine learning on each pixel so that the first label is output from the second convolutional neural network into which the shape information is input, During the execution of the first machine learning and the second machine learning, based on the output value of each pixel of the convolutional layer constituting one of the first convolutional neural network and the second convolutional neural network, select the convolution operation of each pixel of the convolutional layer constituting the other of the first convolutional neural network and the second convolutional neural network, Recognize the road surface marking using at least one of the learned first convolutional neural network with parameters adjusted based on the first machine learning or the learned second convolutional neural network with parameters adjusted based on the second machine learning.

Citation Information

Patent Citations

  • Intelligent drive control method and device, vehicle, electronic device, medium and product

    JP2021504245A