A road arrow recognition method and device based on template self-supervision

Through the combination of self-supervised template comparison learning model and depth model, road arrow recognition without manual labeling is achieved, solving the problem of inefficiency reliance on manual labeling and transformation in the existing technology, and improving the recognition efficiency and accuracy.

CN113947763BActive Publication Date: 2025-05-06HOHAI UNIV CHANGZHOU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111220342.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-20
Publication Date
2025-05-06
Estimated Expiration
2041-10-20

AI Technical Summary

Technical Problem

The existing pavement arrow recognition methods require manual annotation or perspective transformation, which is inefficient and relies on manual supervision.

Method used

The recognition method based on template self-supervisation is adopted, through self-supervised template comparison learning model, the automatic transformation of the template generates arrow pictures from various perspectives, and a depth model that can judge the matching of the instance and the template is achieved to realize arrow recognition.

Benefits of technology

An effective arrow recognition model can be trained without manual annotation, which improves recognition efficiency and accuracy and reduces the dependence on manual supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113947763B_ABST
    Figure CN113947763B_ABST
Patent Text Reader

Abstract

The present invention discloses a road arrow recognition method and device based on template self-supervision, comprising: using a pre-constructed self-supervised template contrast learning model including an encoder and a projection layer to train on a pre-constructed extended data set to obtain a trained network, wherein the extended data set is an extended data set of arrow templates and their corresponding instance samples; using the trained network to identify the category of arrow images; the present invention can complete the training of the arrow recognition model by only using the template without any manual annotation, and obtains arrow images simulating various viewing angles through automatic transformation of the template, and then inputs the model in pairs to train a deep model that can judge whether the instance matches the template, which can be used for arrow recognition. Self-supervised contrast learning can effectively avoid the dependence on manual annotation, and train a reliable recognition model without the need for manual supervision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a road arrow recognition method and device based on template self-supervision, belonging to the technical field of image processing. Background Art

[0002] In the fields of intelligent driving and intelligent traffic facilities, it is necessary to automatically identify the guide arrows of lanes. The present invention discloses a more advanced road guide arrow recognition method that can be applied in the above fields. The current road arrow recognition method needs to first perform perspective transformation on the image acquired by the sensor to obtain a top view without deformation, and then use template matching to identify the arrow; or manually collect a large number of arrow images and mark them before training the classifier for recognition. The template self-supervised recognition method can complete the training of the arrow recognition model using only templates without any manual annotation. Arrow images simulating various perspectives are obtained through automatic transformation of the template, and then they are input into the model in pairs to train a deep model that can determine whether the instance matches the template, which can be used for arrow recognition. Summary of the invention

[0003] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a road arrow recognition method and device based on template self-supervision. The training of the arrow recognition model can be completed by using only the template without any manual annotation. Arrow images simulating various perspectives are obtained by automatic transformation of the template, and then they are input into the model in pairs to train a deep model that can determine whether the instance matches the template, which can be used for arrow recognition.

[0004] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0005] In a first aspect, the present invention provides a method for identifying road arrows based on template self-supervision, comprising:

[0006] Using a pre-built self-supervised template contrastive learning model including an encoder and a projection layer to train on a pre-built augmented dataset to obtain a trained network, wherein the augmented dataset is an augmented dataset of arrow templates and their corresponding instance samples;

[0007] The trained network is used to identify the category of the arrow image.

[0008] Furthermore, the specific data expansion method used when constructing the expanded data set is:

[0009] Each template image is transformed through perspective transformation to obtain 1200 arrow images at different angles. The transformation covers: azimuth angle 0 to 360 degrees, step interval 45 degrees, a total of 8 types; pitch angle 0 to -70 degrees, step interval -5 degrees, a total of 15 types; viewing angle 42 degrees to 132 degrees, step interval 10 degrees, a total of 10 types;

[0010] The expanded image is paired with its corresponding arrow template as a positive pair, and with the remaining arrow templates as a negative pair.

[0011] Furthermore, the specific network structure of the self-supervised template contrast learning model is:

[0012] Select any neural network with feature extraction capability as the encoder to extract features. The image is calculated by the encoder to obtain the feature vector:

[0013] ;

[0014] Use a two-layer fully connected network as a projection layer to project the feature map into the space used for contrast loss, and the projection vector is the feature vector Output at the projection layer:

[0015] ;

[0016] in, Represents the ReLu activation function;

[0017] Given a pair of images , use cosine similarity to define the similarity between them:

[0018]

[0019] Construct the loss function to be optimized:

[0020]

[0021] in It is a hyperparameter, interpreted as a temperature parameter, which can enhance the repulsion effect between negative pairs, improve the optimization ability of the loss function, and accelerate the convergence speed.

[0022] Furthermore, the pre-built self-supervised template contrastive learning model including an encoder and a projection layer is trained on a pre-built augmented data set, including:

[0023] The combined positive and negative pairs are input into the pre-built self-supervised template contrast learning model, and the image and template are calculated by the encoder to obtain feature vectors :

[0024] ; ;

[0025] Project the obtained feature vector into the mapping space for contrastive loss, are image feature vectors , Template Features Output at the projection layer:

[0026] ;

[0027] The network automatically adjusts its parameters based on the back-propagation of the loss function.

[0028] Furthermore, the method of using the trained network to identify the category of the arrow image includes:

[0029] Send the arrow image to be identified into the trained feature extraction network and obtain its feature vector;

[0030] Use the similarity calculation formula to calculate the similarity between the feature vector of the image and each template vector;

[0031] The arrow image is identified as the arrow category with the greatest similarity.

[0032] In a second aspect, the present invention provides a road arrow recognition device based on template self-supervision, comprising:

[0033] A training unit, used for training a pre-built self-supervised template contrast learning model including an encoder and a projection layer on a pre-built augmented data set to obtain a trained network, wherein the augmented data set is an augmented data set of arrow templates and their corresponding instance samples;

[0034] The recognition unit is used to use the trained network to recognize the category of the arrow image.

[0035] In a third aspect, the present invention provides a road arrow recognition device based on template self-supervision, comprising a processor and a storage medium;

[0036] The storage medium is used to store instructions;

[0037] The processor is used to operate according to the instructions to execute the steps of any of the methods described above.

[0038] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the above methods when executed by a processor.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] The present invention uses self-supervised contrastive learning in combination with images of template arrows in the national standard to achieve classification of arrow images through learning. The training of the arrow recognition model can be completed by using only the template without any manual annotation. Arrow images simulating various perspectives are obtained through automatic transformation of the template, and then they are input into the model in pairs to train a deep model that can determine whether the instance matches the template, which can be used for arrow recognition. Self-supervised contrastive learning can effectively avoid dependence on manual annotation, and train a reliable recognition model without the need for manual supervision. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flow chart of a road arrow recognition method based on template self-supervision provided by an embodiment of the present invention;

[0042] Figure 2 These are 11 arrow template images in the national standard of my country provided by the embodiment of the present invention;

[0043] Figure 3 It is a partial extended data display diagram provided by an embodiment of the present invention;

[0044] Figure 4 It is a diagram of the training process of a template self-supervised contrastive learning neural network provided by an embodiment of the present invention;

[0045] Figure 5 It is the composite direction arrow of a left turn actually photographed and provided by an embodiment of the present invention, and the template image successfully matched by the method. DETAILED DESCRIPTION

[0046] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0047] Embodiment 1: This embodiment introduces a road arrow recognition method based on template self-supervision, including:

[0048] Using a pre-built self-supervised template contrastive learning model including an encoder and a projection layer to train on a pre-built augmented dataset to obtain a trained network, wherein the augmented dataset is an augmented dataset of arrow templates and their corresponding instance samples;

[0049] The trained network is used to identify the category of the arrow image.

[0050] The following is a further detailed description of the technical solution of the present invention by taking 11 types of arrows in my country's national standard as examples.

[0051] The application process of the road arrow recognition method based on template self-supervision provided in this embodiment specifically involves the following steps:

[0052] Step S1: construct an extended dataset of arrow templates and their corresponding instance samples;

[0053] Step S2: construct a self-supervised template contrast learning model including an encoder and a projection layer;

[0054] Step S3: Use the model of step S1 to train on the expanded data set of step S2;

[0055] Step S4: Use the network trained in step S3 to identify the category of the arrow image.

[0056] According to the road arrow recognition method based on template self-supervision, Figure 1 As shown, the specific data expansion method in step S1 is:

[0057] S1.1 Each template image is transformed by perspective to obtain 1200 arrow images at different angles, such as Figure 2 The following are 11 template images in the national standard. Figure 3 These are some selected transformation image samples. The transformation covers: azimuth angle 0 to 360 degrees, with a step interval of 45 degrees, a total of 8 types; pitch angle 0 to -70 degrees, with a step interval of -5 degrees, a total of 15 types; viewing angle 42 degrees to 132 degrees, with a step interval of 10 degrees, a total of 10 types.

[0058] S1.2: The expanded image is paired with its corresponding arrow template as a positive pair, and paired with the remaining arrow templates as a negative pair.

[0059] According to the road arrow recognition method based on template self-supervision, the specific network structure of the self-supervised template comparison learning in step S2 is:

[0060] S2.1: Select any neural network with feature extraction capability as the encoder to extract features. All networks composed of convolutional layers and pooling layers, including VGGNet, GoogLeNet, ResNet, DenseNet, etc., can be used. The image is calculated by the encoder to obtain the feature vector:

[0061] ;

[0062] S2.2: Use two layers of nonlinear MLP as projection layers to project the feature vector into the space used for contrastive loss. The projection vector is the feature vector Output at the projection layer:

[0063] ;

[0064] in, Represents the ReLu activation function;

[0065] S2.3: Given a pair of images , use cosine similarity to define the similarity between them:

[0066]

[0067] S2.4: Use the loss function for back propagation and optimize network parameters:

[0068] ;

[0069] in, is the temperature parameter, and its value is 0.1.

[0070] According to the road arrow recognition method based on template self-supervision, Figure 4 As shown, the specific steps of training the network in step S3 are:

[0071] S3.1: The positive and negative pairs obtained in step S1 are input into the network structure constructed in step S2. The image and template are calculated by the encoder to obtain feature vectors. :

[0072] ; ;

[0073] S3.2: Project the obtained feature vector into the mapping space for contrastive loss, are image feature vectors , Template Features Output at the projection layer:

[0074] ;

[0075] S3.3: The network automatically adjusts parameters based on the back propagation of the loss function. This can be used in deep learning toolboxes such as PyTorch and Tensorflow.

[0076] According to the road arrow recognition method based on template self-supervision, the trained network in step S3 is applied to a specific classification task, and the specific steps are as follows:

[0077] S4.1: Send the arrow image to be identified into the feature extraction network trained in step S3, and obtain its feature vector;

[0078] S4.2: Using the similarity calculation formula in step S2.3, calculate the similarity between the feature vector of the image and each template vector;

[0079] S4.3: Identify the arrow image as the arrow category with the greatest similarity.

[0080] Figure 5 On the left is an actual photograph of a left-turn composite arrow image on the road surface. After comparing it with the arrows in the template library using the method of this patent, the left-turn composite arrow in the template library is successfully matched.

[0081] Embodiment 2: This embodiment provides a road arrow recognition device based on template self-supervision, comprising:

[0082] A training unit, used for training a pre-built self-supervised template contrast learning model including an encoder and a projection layer on a pre-built augmented data set to obtain a trained network, wherein the augmented data set is an augmented data set of arrow templates and their corresponding instance samples;

[0083] The recognition unit is used to use the trained network to recognize the category of the arrow image.

[0084] Embodiment 3, this embodiment provides a road arrow recognition device based on template self-supervision, including a processor and a storage medium;

[0085] The storage medium is used to store instructions;

[0086] The processor is configured to operate according to the instructions to perform the steps of any of the following methods:

[0087] Using a pre-built self-supervised template contrastive learning model including an encoder and a projection layer to train on a pre-built augmented dataset to obtain a trained network, wherein the augmented dataset is an augmented dataset of arrow templates and their corresponding instance samples;

[0088] The trained network is used to identify the category of the arrow image.

[0089] Embodiment 4, this embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of any of the following methods are implemented;

[0090] Using a pre-built self-supervised template contrastive learning model including an encoder and a projection layer to train on a pre-built augmented dataset to obtain a trained network, wherein the augmented dataset is an augmented dataset of arrow templates and their corresponding instance samples;

[0091] The trained network is used to identify the category of the arrow image.

[0092] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A road arrow recognition method based on template self-supervision, characterized in that: include: Using a pre-built self-supervised template contrastive learning model including an encoder and a projection layer to train on a pre-built augmented dataset to obtain a trained network, wherein the augmented dataset is an augmented dataset of arrow templates and their corresponding instance samples; Using the trained network to identify the category of the arrow image; The specific network structure of the self-supervised template contrast learning model is: Select any neural network with feature extraction capability as the encoder to extract features. The image is calculated by the encoder to obtain the feature vector: ; Use a two-layer fully connected network as a projection layer to project the feature map into the space used for contrast loss, and the projection vector is the feature vector Output at the projection layer: ; in, Represents the ReLu activation function; Given a pair of images , use cosine similarity to define the similarity between them: ; Construct the loss function to be optimized: ; in It is a hyperparameter, interpreted as a temperature parameter, which can enhance the repulsion effect between negative pairs, improve the optimization ability of the loss function, and accelerate the convergence speed.

2. The method for road arrow recognition based on template self-supervision according to claim 1, characterized in that: The specific data expansion method used when constructing the expanded data set is: Each template image is transformed through perspective transformation to obtain 1200 arrow images at different angles. The transformation covers: azimuth angle 0 to 360 degrees, step interval 45 degrees, a total of 8 types; pitch angle 0 to -70 degrees, step interval -5 degrees, a total of 15 types; viewing angle 42 degrees to 132 degrees, step interval 10 degrees, a total of 10 types; The expanded image is paired with its corresponding arrow template as a positive pair, and with the remaining arrow templates as a negative pair.

3. The method for road arrow recognition based on template self-supervision according to claim 2, characterized in that: The self-supervised template contrastive learning model using a pre-built encoder and projection layer is trained on a pre-built augmented dataset, including: The combined positive and negative pairs are input into the pre-built self-supervised template contrast learning model, and the image and template are calculated by the encoder to obtain feature vectors : ; ; Project the obtained feature vector into the mapping space for contrastive loss, are image feature vectors , Template Features Output at the projection layer: ; The network automatically adjusts its parameters based on the back-propagation of the loss function.

4. The method for road arrow recognition based on template self-supervision according to claim 3 is characterized in that: The step of using the trained network to identify the category of the arrow image includes: Send the arrow image to be identified into the trained feature extraction network and obtain its feature vector; Use the similarity calculation formula to calculate the similarity between the feature vector of the image and each template vector; The arrow image is identified as the arrow category with the greatest similarity.

5. A road arrow recognition device based on template self-supervision, characterized in that: include: A training unit is used to use a pre-built self-supervised template contrast learning model including an encoder and a projection layer to train on a pre-built extended data set to obtain a trained network, wherein the extended data set is an extended data set of arrow templates and their corresponding instance samples; wherein the specific network structure of the self-supervised template contrast learning model is: Select any neural network with feature extraction capability as the encoder to extract features. The image is calculated by the encoder to obtain the feature vector: ; Use a two-layer fully connected network as a projection layer to project the feature map into the space used for contrast loss, and the projection vector is the feature vector Output at the projection layer: ; in, Represents the ReLu activation function; Given a pair of images , use cosine similarity to define the similarity between them: ; Construct the loss function to be optimized: ; in It is a hyperparameter, interpreted as a temperature parameter, which can enhance the repulsion effect between negative pairs, improve the optimization ability of the loss function, and accelerate the convergence speed; The recognition unit is used to use the trained network to recognize the category of the arrow image.

6. A road arrow recognition device based on template self-supervision, characterized in that: including processor and storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on self-supervised contrast learning

    CN113011427A

  • Deep learning SAR image ship identification method based on self-supervision condition

    CN113378716A