Training Method for Traffic Signal Recognition Model and Traffic Signal Recognition Method

By cropping, scaling and merging the imaging content of traffic lights, combined with the multi-layer sub-network processing of neural network models, the problem of inaccurate traffic light recognition in the prior art is solved, and high-accurate traffic light recognition is achieved.

CN114332704BActive Publication Date: 2025-06-13BEIJING PACTERA JINXIN TECH LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111629580.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-06-13
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

The neural network model in the prior art cannot accurately identify traffic lights and is greatly affected by background noise.

Method used

By obtaining training data sets of multiple training samples, each sample contains at least two images marked with the actual position of the bounded box of the traffic light imaging content. Then multiple sample images are cropped, scaled and merged to form a merged image. The merged image is input to the multi-layer subnet of the initial neural network model for processing. Through adaptive anchor boxes, interlaced columns, and pixels are extracted and convolutional processing are gradually extracted and the position of the bounded box is adjusted. Finally, the traffic light recognition model is obtained through iterative training.

Benefits of technology

Accurate recognition of traffic lights is achieved, interference factors in the image are removed, and the accuracy of recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332704B_ABST
    Figure CN114332704B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a training method for a traffic signal recognition model, a recognition method for traffic signals, a training device for a traffic signal recognition model, a recognition device for traffic signals, an electronic device, a computer-readable storage medium, and a computer program product. It relates to the technical field of autonomous driving and assisted driving. The method includes: cropping, scaling, and arranging multiple sample images and then merging them to obtain a merged image; the actual positions of the bounding boxes marked with the imaging content of traffic signals in the sample images; inputting the merged image into an initial neural network, obtaining the predicted positions of the bounding boxes output by the initial neural network, calculating the loss function value based on the predicted positions and the actual positions, and iteratively updating the initial neural network model according to the loss function value to obtain a traffic signal recognition model. The traffic signal recognition model in the embodiment of the present application can accurately recognize the bounding boxes of traffic signals and improve the recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of autonomous driving and assisted driving. Specifically, the present application relates to a method for training a traffic signal recognition model, a method for recognizing a traffic signal, a device for training a traffic signal recognition model, a device for recognizing a traffic signal, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] As an important carrier of road traffic information, a traffic signal can provide current driving information for a driver. In difficult and dangerous driving environments, it can timely provide safety warnings to urge the driver to drive carefully. In the field of driverless or assisted driving, the traffic signal is recognized through an algorithm. Correctly recognizing the indication signals such as the indication color of the traffic light at the current intersection is the guarantee for ensuring the normal and safe operation of the vehicle at a traffic intersection with a large number of people and vehicles, and is also the guarantee for ensuring the continuous and normal driving of an autonomous vehicle on the operating road.

[0003] The existing solutions for traffic signal recognition mainly have the following two methods. The existing solutions for traffic signal recognition mainly have the following several methods. The first is the RGB color model method. The image containing the imaging content of the traffic signal is input into the RGB color model to recognize the color of the traffic signal in the image. This method can quickly recognize the color of the traffic signal, but due to the influence of background noise (such as the color of surrounding LED lights), the recognition of the traffic signal color is inaccurate to a large extent. The second is the HSI color model method. The HSI color space has characteristics such as illumination invariance, so it has good robustness. However, the HSI color model still cannot remove the influence of background noise, that is, the above two training models cannot accurately recognize traffic signals. Summary of the Invention

[0004] Embodiments of the present application provide a method for training a traffic signal recognition model, a method for recognizing a traffic signal, a device for training a traffic signal recognition model, a device for recognizing a traffic signal, an electronic device, a computer-readable storage medium, and a computer program product, which are used to solve the technical problem that the neural network model in the prior art cannot accurately recognize traffic signals.

[0005] According to the first aspect of the embodiments of the present application, a method for training a traffic signal recognition model is provided. The method includes:

[0006] Obtain a training data set including a plurality of training samples; each training sample includes at least two sample images, and the actual position of the bounding box of the imaging content of the traffic signal is marked on the sample image;

[0007] Crop multiple sample images, scale and arrange the multiple cropped sample images, and then merge them to obtain a merged image;

[0008] Input the merged image into the first sub-network of the initial neural network model to obtain a test image output by the first sub-network after performing adaptive anchor boxes and adaptive scaling on the merged image; the initial positions of all bounding boxes are marked in the test image;

[0009] Input the test image into the second sub-network of the initial neural network model to obtain multiple feature maps output after the second neural network extracts pixels row by row and column by column at multiple preset pixels in the test image and performs convolutional processing on the multiple sub-images obtained after pixel extraction;

[0010] Input the multiple feature maps into the third sub-network of the initial neural network model to obtain a fused feature map obtained by the third sub-network after fusing the features of the multiple feature maps;

[0011] Input the fused feature map and the test image into the fourth sub-network of the initial neural network model to obtain the target region mapped by the fused feature map matched by the fourth sub-network in the test image, and output the predicted position of the bounding box after adjusting the initial position of the bounding box according to the target region;

[0012] Calculate the loss function value of the preset neural network loss function according to the actual position and the predicted position, and iteratively train the initial neural network model according to the loss function value to obtain a traffic signal recognition model.

[0013] In a possible implementation, cropping multiple sample images includes:

[0014] For each sample image, obtain the horizontal resolution and vertical resolution of the sample image;

[0015] If any one of the horizontal resolution or the vertical resolution does not meet the multiple of the preset value, calculate the minimum value of the pixels to be cropped in the horizontal direction or the vertical direction;

[0016] Crop the sample image according to the minimum value so that the horizontal resolution and vertical resolution of the cropped sample image both meet the multiple of the preset value;

[0017] Wherein, the minimum value is less than the preset value.

[0018] In a possible implementation, scaling, arranging and then merging multiple cropped sample images to obtain a merged image includes:

[0019] Obtain a preset horizontal resolution and a preset vertical resolution, where both the preset horizontal resolution and the preset vertical resolution are multiples of a preset value;

[0020] Based on the preset vertical resolution, the preset vertical resolution, the horizontal resolution and the vertical resolution of the cropped sample image, determine a first scaling ratio in the horizontal direction and a second scaling ratio in the vertical direction for each cropped sample image; Scale the cropped sample images according to the first scaling ratio and the second scaling ratio;

[0021] Randomly arrange and merge the scaled sample images to obtain a merged image.

[0022] In a possible implementation, input the merged image into the first sub-network of the initial neural network model, and obtain a test image output by the first sub-network after performing adaptive anchor boxes and adaptive scaling on the merged image, including:

[0023] Perform adaptive anchor box operation on the merged image, mark the initial positions of all bounding boxes in the merged image, and obtain a marked image with the initial positions of all bounding boxes marked;

[0024] Perform adaptive scaling on the marked image to obtain a test image, where both the horizontal resolution and the vertical resolution of the test image are multiples of a preset value.

[0025] In a possible implementation, input the test image into the second sub-network of the initial neural network model, and obtain multiple feature maps output by the second neural network after extracting pixels at multiple preset pixels of the test image row by row and column by column and performing multiple convolutional processes on the multiple sub-images obtained after extracting the pixels, including:

[0026] Determine multiple preset pixels in the test image;

[0027] For any one of the preset pixels, extract pixels row by row and column by column starting from the preset pixel to obtain a sub-image of the test image;

[0028] For any one sub-image, perform multiple convolutional processes on the sub-image to extract the features of the sub-image and obtain the feature map corresponding to the sub-image.

[0029] In a possible implementation, input the fused feature map and the test image into the fourth sub-network of the initial neural network model, and obtain the target region mapped by the fused feature map matched by the fourth sub-network in the test image, and the predicted position of the bounding box output after adjusting the initial position of the bounding box according to the target region, including:

[0030] Extract the feature vector of the fused feature map; The feature vector includes pixel value features, texture features, shape features and spatial relationship features;

[0031] Determine the target area where the feature vector is mapped in the image to be measured; adjust the initial position of the bounding box according to the target area to obtain the predicted position of the bounding box.

[0032] In a possible implementation, the actual position is characterized by the actual coordinates of all pixels on the bounding box; the predicted position is characterized by the predicted coordinates of all pixels on the bounding box; calculate the loss function value of a preset neural network loss function according to the actual position and the predicted position, and iteratively train the initial neural network model according to the loss function value to obtain a traffic signal recognition model, including:

[0033] Input the actual coordinates of all pixels on the bounding box and the predicted coordinates of all pixels into a preset neural network loss function to obtain the loss function value;

[0034] If the loss function value does not meet the training end condition of the neural network model, continue to iteratively train the initial neural network model based on each training sample and the actual position of the bounding box marked in each sample image in each training sample until the training end condition of the neural network model is met to obtain a traffic signal recognition model.

[0035] According to the second aspect of the embodiments of the present application, a traffic signal recognition method is provided, and the method includes:

[0036] Obtain an image to be measured at a predetermined position; the image to be measured contains the imaging content of a traffic signal;

[0037] Input the image to be measured into the traffic signal recognition model to obtain the position of the bounding box for marking the imaging content output by the traffic signal recognition model in the image to be measured;

[0038] Identify the color of the area within the bounding box through a preset color model, use the color of the area within the bounding box as the color of the traffic signal, and determine the state of the traffic signal according to the color of the traffic signal;

[0039] The traffic signal recognition model is trained by the method of the first aspect.

[0040] In a possible implementation, the states of the traffic signal include allowing passage, warning, and prohibiting passage; determining the state of the traffic signal according to the color of the traffic signal includes:

[0041] If the color of the traffic signal is green, determine that the state of the traffic signal is running and allowing passage;

[0042] If the color of the traffic signal is yellow, determine that the state of the traffic signal is warning;

[0043] If the color of the traffic signal is yellow, determine that the status of the traffic signal is prohibited from passing.

[0044] According to the third aspect of the embodiments of the present application, a training device for a traffic signal recognition model is provided. The device includes: an acquisition module, configured to acquire a training data set including a plurality of training samples; each training sample includes at least two sample images, and the sample images are labeled with the actual positions of the bounding boxes of the imaging content of the traffic signal.

[0045] A cropping, scaling, arranging, and merging module, configured to crop multiple sample images, scale and arrange the multiple cropped sample images, and then merge them to obtain a merged image; the initial positions of all bounding boxes are labeled in the image to be measured.

[0046] An image-to-be-measured acquisition module, configured to input the merged image into the first-layer sub-network of the initial neural network model, and obtain an image to be measured output by the first-layer sub-network after performing adaptive anchor boxes and adaptive scaling on the merged image.

[0047] A feature map acquisition module, configured to input the image to be measured into the second-layer sub-network of the initial neural network model, and obtain multiple feature maps output by the second neural network after extracting pixels row by row and column by column starting from multiple preset pixels in the image to be measured and performing multiple convolution processes on the multiple sub-images obtained after pixel extraction.

[0048] A fused feature map acquisition module, configured to input the multiple feature maps into the third-layer sub-network of the initial neural network model, and obtain a fused feature map obtained by the third-layer sub-network after performing feature fusion on the multiple feature maps.

[0049] A predicted position output module, configured to input the fused feature map and the image to be measured into the fourth-layer sub-network of the initial neural network model, obtain the target area mapped by the fused feature map in the image to be measured matched by the fourth-layer sub-network, and output the predicted position of the bounding box after adjusting the initial position of the bounding box according to the target area.

[0050] An iteration module, configured to calculate the loss function value of a preset neural network loss function according to the actual position and the predicted position, and perform iterative training on the initial neural network model according to the loss function value to obtain a traffic signal recognition model.

[0051] According to the fourth aspect of the embodiments of the present application, a traffic signal recognition device is provided. The device includes:

[0052] An image-to-be-measured acquisition module, configured to acquire an image to be measured at a predetermined position; the image to be measured includes the imaging content of the traffic signal.

[0053] A bounded box position recognition module, configured to input a to-be-detected image into a traffic signal recognition model, and obtain the position of a bounded box for marking imaging content output by the traffic signal recognition model in the to-be-detected image;

[0054] A status recognition module, configured to recognize the color of the area within the bounded box through a preset color model, use the color of the area within the bounded box as the color of the traffic signal, and determine the status of the traffic signal according to the color of the traffic signal;

[0055] The traffic signal recognition model is trained according to the method of any one of claims 1 to 7.

[0056] According to a fifth aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes: the electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the methods provided in the first aspect and the second aspect are implemented.

[0057] According to a sixth aspect of the embodiments of the present application, a computer-readable storage medium is provided. A computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods provided in the first aspect and the second aspect are implemented.

[0058] According to a seventh aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. When a processor of a computer device reads the computer instructions from the computer-readable storage medium, the processor executes the computer instructions, so that the computer device executes the steps of the methods provided in the first aspect and the second aspect.

[0059] The beneficial effects brought by the technical solutions provided in the embodiments of the present application are as follows: The traffic signal recognition model trained in the embodiments of the present application can recognize the bounded box of the imaging content of the traffic signal in the image, remove the interference factors in the image, and can accurately recognize the traffic signal, improving the accuracy of traffic signal recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application.

[0061] Figure 1 It is a schematic flowchart of a method for training a traffic signal recognition model provided by an embodiment of the present application;

[0062] Figure 2Schematic diagram of the process of obtaining multiple sub-images by extracting pixels row by row and column by column at multiple preset pixels of the image to be measured provided by the embodiment of the present application;

[0063] Figure 3 Schematic flowchart of a traffic signal recognition method provided by the embodiment of the present application;

[0064] Figure 4 Schematic structural diagram of a training device for a traffic signal recognition model provided by the embodiment of the present application;

[0065] Figure 5 Schematic structural diagram of a traffic signal recognition device provided by the embodiment of the present application;

[0066] Figure 6 Schematic structural diagram of an electronic device provided by the embodiment of the present application. Detailed implementation manners

[0067] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0068] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "" and "the" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0069] To make the purpose, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0070] The target vehicle A is currently in the autonomous driving mode. There is a traffic light 500 meters ahead shown on the navigation map. At this time, the built-in camera in the target vehicle A will collect a video stream containing the imaging content of the traffic light ahead, decode the video stream to obtain the image to be detected, and detect the color of the traffic light through the following three methods.

[0071] The first is the RGB color model method. The image to be detected containing the imaging content of the traffic light is input into the RGB color model to identify the color of the traffic light in the image to be detected. However, the image to be detected contains a lot of background colors, such as the colors of the LED lights on the surrounding buildings. The background colors will largely lead to inaccurate recognition of the traffic light color.

[0072] The second is the HSI color model method. The image to be detected containing the imaging content of the traffic light is input into the HSI color model to identify the color of the traffic light in the image to be detected. The HSI color space has characteristics such as invariance to illumination and advantages such as good robustness, but it still cannot remove the influence of the background color.

[0073] A training method for a traffic light recognition model, a traffic light recognition method, a training device for a traffic light recognition model, a traffic light recognition method device, an electronic device, a computer-readable storage medium, and a computer program product provided by the present application aim to solve the above technical problems in the prior art.

[0074] Next, through the description of several exemplary embodiments, the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0075] An embodiment of the present application provides a method for training a traffic light recognition model, as Figure 1 shown. The method includes:

[0076] Step S101, obtaining a training data set including a plurality of training samples; each training sample includes at least two sample images, and the actual positions of the bounding boxes of the imaging content of the traffic light are marked on the sample images.

[0077] Each training sample in the embodiment of the present application includes multiple sample images, that is, in the model training stage, multiple sample images can be processed at one time.

[0078] The sample image in the embodiment of the present application is an image containing the imaging content of a traffic signal. In practical applications, traffic signals have various shapes, such as round lights, left arrow lights, or right arrow lights, etc. The imaging content of the traffic signal in the test image of the embodiment of the present application can be any of the above shapes, and the embodiment of the present application does not limit this.

[0079] During the driving of the vehicle, a video stream containing the imaging content of the traffic signal in front of the vehicle can be captured in real time by an image acquisition device (such as a camera) on the vehicle, and then the video stream can be decoded based on FFMPEG (Fast Forward Mpeg, a multimedia video processing tool) or OpenCV (a cross-platform computer vision and machine learning software library distributed under the BSD license) to extract the images of each frame containing the imaging content of the traffic signal.

[0080] In the sample image of the embodiment of the present application, the actual position of the bounding box marked with the imaging content of the traffic signal, that is, the actual position of the bounding box is the training label.

[0081] Step S102: Crop multiple sample images, scale and arrange the multiple cropped sample images, and then merge them to obtain a merged image.

[0082] After obtaining multiple sample images in the embodiment of the present application, each of the multiple sample images is cropped respectively to make the cropped sample images meet the requirements. The detailed cropping process is described in the following part.

[0083] In the magical embodiment of the present application, after cropping multiple sample images, the multiple cropped sample images are scaled, arranged, and then merged to obtain a merged image. Merging multiple images can enrich the detection targets in the image and improve the robustness of image processing.

[0084] Step S103: Input the merged image into the first sub-network of the initial neural network model, and the first sub-network outputs a test image after performing adaptive anchor boxes and adaptive scaling on the merged image; all the initial positions of the bounding boxes are marked in the test image.

[0085] After obtaining the merged image in the embodiment of the present application, the merged image is input into the first sub-network of the initial neural network model. In this first sub-network, adaptive anchor boxes are performed on the merged image. The adaptive anchor box is to determine the initial position of the bounding box of the traffic signal. The initial position is the possible position of the bounding box. After performing the adaptive anchor box operation on the merged image, a marked image marked with all the initial positions of the bounding boxes is obtained.

[0086] An image has a horizontal resolution and a vertical resolution. The horizontal resolution represents the number of pixels in the horizontal direction, and the vertical resolution represents the number of pixels in the vertical direction. After obtaining the marked image, it is also necessary to perform adaptive scaling on the image, scale the horizontal resolution and the vertical resolution of the marked image, and obtain the image to be measured after adaptive scaling.

[0087] Step S104: Input the image to be measured into the second sub-network of the initial neural network model, and obtain multiple feature maps output after the second neural network starts to extract pixels row by row and column by column at multiple preset pixels of the image to be measured and performs multiple convolutional processes on multiple sub-images obtained after pixel extraction.

[0088] In the embodiment of the present application, after obtaining the image to be measured output by the first sub-network, the image to be measured is input into the second sub-network. After receiving the image to be measured, the second sub-network starts to extract pixels row by row and column by column at multiple preset pixels of the image to be measured. The preset pixels can be 4 pixels in the overlapping part of the first two rows and the first two columns. Starting from each preset pixel, pixels are extracted row by row and column by column to obtain a sub-image. If there are 4 preset pixels, 4 sub-images can be extracted.

[0089] As Figure 2 shown, it exemplarily shows a schematic diagram of the process of obtaining multiple sub-images after extracting pixels row by row and column by column at multiple preset pixels of the image to be measured. Assume that the image to be measured is 4*4, and starting from the 4 pixels overlapping in the first two rows and the first two columns, pixels are extracted row by row and column by column to obtain four 2*2 sub-images.

[0090] After the embodiment of the present application performs the process of extracting pixels row by row and column by column on the image to be measured, multiple sub-images of the image to be measured are obtained. These sub-images can be combined completely to form the image to be measured without any information loss.

[0091] After the embodiment of the present application obtains multiple sub-images, convolutional processes are respectively performed on each sub-image to extract image features. Multiple convolutional processes can be performed to make the extracted image features relatively stable. The extracted image features include pixel value features, texture features, shape features, spatial relationship features, etc. of each pixel.

[0092] Step S105: Input the multiple feature maps into the third sub-network of the initial neural network model, and obtain a fused feature map obtained by the third sub-network performing feature fusion on the multiple feature maps.

[0093] Convolution processing is performed on multiple sub - images respectively to extract image features, and various features of the sub - images can be obtained completely. The features in the feature maps of each sub - image are also the features of the image to be measured. There are both identical and different parts in multiple feature maps. It is necessary to fuse multiple feature maps, fuse the identical features, and prevent excessive resource occupation. After obtaining multiple feature maps in the embodiments of the present application, the multiple feature maps are simultaneously input into the third - layer sub - network to obtain a fused feature map obtained by the third - layer sub - network after fusing the multiple feature maps.

[0094] Step S106: Input the fused feature map and the image to be measured into the fourth - layer sub - network of the initial neural network model, obtain the target area mapped by the fused feature map matched by the fourth - layer sub - network in the image to be measured, and output the predicted position of the bounding box after adjusting the initial position of the bounding box according to the target area.

[0095] After obtaining the fused feature map in the embodiments of the present application, the fused feature map and the image to be measured are input into the fourth - layer sub - network. The fourth - layer sub - network matches the target area mapped by the fused feature map in the image to be measured according to the features in the fused feature map. This target area is the area where the imaging content is located. The initial position of the bounding box is adjusted according to this target area, that is, the position of the bounding box is adjusted to the position where the target area is located. This position is the predicted position of the bounding box output by the fourth - layer sub - network.

[0096] Step S107: Calculate the value of the loss function of the preset neural network loss function according to the actual position and the predicted position, and perform iterative training on the initial neural network model according to the value of the loss function to obtain a traffic signal recognition model.

[0097] After determining the predicted position of the bounding box in the embodiments of the present application, calculate the value of the loss function of the preset neural network loss function according to this predicted position and the actual position. The neural network loss function is defined as

[0098]

[0099] where i represents the i - th pixel on the bounding box, m represents that there are m pixels in total on the bounding box, (xθ, yθ) is a certain pixel of the predicted position of the bounding box, and (x0, y0) is a certain pixel at the actual position of the bounding box. In the embodiments of the present application, the value of the loss function will be obtained after each operation. The neural network model is iteratively trained according to the value of the loss function, that is, the network parameters of the neural network model are updated to adjust the predicted position of the bounding box. When the value of the loss function is the smallest, that is, when the derivative of the loss function approaches 0 infinitely, the neural network loss function converges. In this case, the predicted position of the bounding box predicted is closest to the actual position of the bounding box. That is, the initial neural network model is iteratively trained to obtain a traffic signal recognition model.

[0100] The traffic signal recognition model trained in the embodiments of the present application can identify the bounding box of the imaging content of the traffic signal in the image, remove the interference factors in the image, accurately identify the traffic signal, and improve the accuracy of traffic signal recognition.

[0101] A possible implementation is provided in the embodiments of the present application. The cropping of multiple sample images includes:

[0102] For each sample image, obtain the horizontal resolution and vertical resolution of the sample image;

[0103] If any one of the horizontal resolution or vertical resolution does not meet the multiple of the preset value, calculate the minimum value of the pixels to be cropped in the horizontal direction or vertical direction;

[0104] Crop the sample image according to the minimum value, so that the horizontal resolution and vertical resolution of the cropped sample image both meet the multiple of the preset value;

[0105] Wherein, the minimum value is less than the preset value.

[0106] Each sample image in the embodiments of the present application has its corresponding size. The image size is reflected from two aspects, namely the horizontal resolution and the vertical resolution. The horizontal resolution represents the number of pixels in the horizontal direction, and the vertical resolution represents the number of pixels in the vertical direction.

[0107] It can be understood that, in order to facilitate the subsequent processing of the sample image, it is necessary to ensure that the size of the sample image meets the requirements. In the embodiments of the present application, when it is determined that both the horizontal resolution and the vertical resolution of the sample image meet the multiple of the preset value, it is determined that the sample meets the requirements. The preset value can be a multiple of 4. If any one of the horizontal resolution or vertical resolution does not meet the multiple of the preset value, it is determined that the sample image needs to be cropped so that the horizontal resolution and vertical resolution of the cropped sample image both meet the multiple of the preset value.

[0108] It should be noted that the cropping of the sample image in the embodiments of the present application is a fine cropping. When it is determined that any one of the horizontal resolution or vertical resolution does not meet the multiple of the preset value, excessive cropping may cause the imaging content of the traffic signal to be cropped off. To avoid excessive cropping, it is necessary to calculate the minimum value of the pixels that should be cropped when the horizontal resolution is cropped to meet the multiple of the preset value, and crop the sample image according to the minimum value, so that the horizontal resolution and vertical resolution of the cropped sample image both meet the multiple of the preset value.

[0109] Specifically, assume that the horizontal resolution of a sample image is 101, the vertical resolution is 102, and the preset value is 4. Since 101 has a remainder of 1 when divided by 4 and 102 has a remainder of 2 when divided by 4, it can be determined that the minimum value of the pixels cropped in the horizontal direction is 2, and the minimum value of the pixels cropped in the vertical direction is 1.

[0110] An embodiment of the present application provides a possible implementation manner. After scaling and arranging multiple cropped sample images, they are merged to obtain a merged image, including:

[0111] Obtain a preset horizontal resolution and a preset vertical resolution, and both the preset horizontal resolution and the preset vertical resolution are multiples of the preset value;

[0112] Based on the preset vertical resolution, the preset vertical resolution, the horizontal resolution and the vertical resolution of the cropped sample image, determine the first scaling ratio in the horizontal direction and the second scaling ratio in the vertical direction of each cropped sample image; scale the cropped sample image according to the first scaling ratio and the second scaling ratio;

[0113] Randomly arrange and merge the scaled sample images to obtain a merged image.

[0114] The sample image collected by the target vehicle during driving is usually an image directly in front. There are few detection targets (the detection target is the imaging content of a traffic signal light) in this sample image. To improve the detection accuracy, it is necessary to enrich the number of detection targets in the image input to the initial neural network model. By scaling, arranging and then merging multiple sample images, a merged image containing multiple detection targets can be obtained.

[0115] In the embodiment of the present application, the scaled sample image is set with a preset horizontal resolution and a preset vertical resolution, and both the preset horizontal resolution and the preset vertical resolution are multiples of the preset value. Through the preset vertical resolution, the preset vertical resolution, the horizontal resolution and the vertical resolution of the cropped sample image, calculate the first scaling ratio in the horizontal direction and the second scaling ratio in the vertical direction of each cropped sample image; scale the cropped sample image according to the first scaling ratio and the second scaling ratio to obtain a scaled sample image, and randomly arrange and merge the scaled sample images to obtain a merged image.

[0116] An embodiment of the present application provides a possible implementation manner. Input the merged image into the first-layer sub-network of the initial neural network model to obtain an image to be tested output by the first-layer sub-network after performing adaptive anchor boxes and adaptive scaling on the merged image, including:

[0117] Perform an adaptive anchor box operation on the merged image to mark the initial positions of all bounding boxes in the merged image, obtaining a marked image with the initial positions of all bounding boxes marked.

[0118] Perform an adaptive scaling on the marked image to obtain the image to be measured, where both the horizontal resolution and the vertical resolution of the image to be measured meet multiples of a preset value.

[0119] In the embodiment of the present application, after obtaining the merged image, the merged image is input into the first sub-network of the initial neural network model. In this first sub-network, an adaptive anchor box is performed on the merged image. The adaptive anchor box is to determine the initial positions of the bounding boxes of the traffic lights. The initial positions are the possible positions of the bounding boxes. After performing the adaptive anchor box operation on the merged image, a marked image with the initial positions of all bounding boxes marked is obtained. After obtaining the marked image, it is also necessary to perform an adaptive scaling on the image, scaling the horizontal resolution and the vertical resolution of the marked image. After the adaptive scaling, the image to be measured is obtained, and both the horizontal resolution and the vertical resolution of the image to be measured meet multiples of a preset value.

[0120] The embodiment of the present application provides a possible implementation. The image to be measured is input into the second sub-network of the initial neural network model to obtain multiple feature maps output after the second neural network extracts pixels row by row and column by column starting from multiple preset pixels in the image to be measured and performing multiple convolution processes on the multiple sub-images obtained after the pixel extraction, including:

[0121] Determine multiple preset pixels in the image to be measured;

[0122] For any one of the preset pixels, extract pixels row by row and column by column starting from the preset pixel to obtain a sub-image of the image to be measured;

[0123] For any one of the sub-images, perform multiple convolution processes on the sub-image to extract the features of the sub-image, obtaining the feature map corresponding to the sub-image.

[0124] In the embodiment of the present application, the pixel can be 4 pixels in the overlapping part of the first two rows and the first two columns. Starting from each preset pixel as the starting pixel, extract pixels row by row and column by column to obtain a sub-image. There are 4 preset pixels, so 4 sub-images can be extracted.

[0125] In the embodiment of the present application, after obtaining multiple sub-images, perform convolution processes on the multiple sub-images respectively to extract image features, and various features of the sub-images can be obtained, such as pixel value features, texture features, shape features, and spatial relationship features, etc. The features in the feature map of each sub-image are also the features of the image to be measured. There are the same parts and different parts in the multiple feature maps. To prevent occupying too much resources, it is necessary to fuse the multiple feature maps to obtain a fused feature map, and the fused feature map is the feature map of the image to be measured.

[0126] An embodiment of the present application provides a possible implementation. The fused feature map and the image to be tested are input into the fourth sub-network of the initial neural network model to obtain the target area mapped by the fused feature map matched by the fourth sub-network in the image to be tested, and the predicted position of the bounding box output after adjusting the initial position of the bounding box according to the target area, including:

[0127] Extract the feature vector of the fused feature map; the feature vector includes pixel value features, texture features, shape features, and spatial relationship features;

[0128] Determine the target area mapped by the feature vector in the image to be tested; adjust the initial position of the bounding box according to the target area to obtain the predicted position of the bounding box.

[0129] After obtaining the fused feature map in the embodiment of the present application, the feature vector in the fused feature map is extracted. Each parameter of the feature vector represents a kind of feature, including pixel value features, texture features, shape features, and spatial relationship features, etc.

[0130] After extracting the feature vector, determine the target area mapped by the feature vector in the image to be tested. This target area is the area where the imaging content of the traffic signal light is located. Adjust the initial position of the bounding box according to this target area to obtain the predicted position of the bounding box.

[0131] In addition, during the process of determining the predicted position of the bounding box, there may be a situation where there are multiple predicted positions. In this case, a best predicted position is determined through the Non-Maximum Suppression (NMS) algorithm, and this best predicted position is the predicted position of the bounding box to be output.

[0132] An embodiment of the present application provides a possible implementation. The actual position is characterized by the actual coordinates of all pixels on the bounding box; the predicted position is characterized by the predicted coordinates of all pixels on the bounding box; calculate the loss function value of the preset neural network loss function according to the actual position and the predicted position, and iteratively train the initial neural network model according to the loss function value to obtain a traffic signal light recognition model, including:

[0133] Input the actual coordinates of all pixels and the predicted coordinates of all pixels into the loss function of the preset neural network to obtain the loss function value;

[0134] If the loss function value does not meet the training end condition of the neural network model, continue to train the adjusted model based on each training sample and the actual position of the bounding box marked in each sample image in each training sample until the training end condition of the neural network model is met to obtain a traffic signal light recognition model.

[0135] It can be understood that the position of the bounding box can be characterized by the coordinates of all the pixels on the bounding box, that is, the actual position of the bounding box is characterized by the actual coordinates of all the pixels on the bounding box, and the predicted position of the bounding box is characterized by the predicted coordinates of all the pixels on the bounding box. By inputting the actual coordinates and predicted coordinates of all the pixels on the bounding box into a preset neural network loss function, the loss function value can be obtained; the preset loss function is:

[0136]

[0137] where J(θ) represents the loss function value, i represents the i-th pixel on the bounding box, m represents that there are m pixels in total on the bounding box, (x θ , y θ ) is a certain pixel of the predicted position of the bounding box, and (x 0 , y 0 ) is a certain pixel at the actual position of the bounding box.

[0138] If the loss function value does not meet the neural network model training end condition, it means that the neural network training model has not been trained to convergence, that is, when the loss function value is not the minimum value, it is necessary to continue iterative training on the adjusted model according to each training sample and the actual position of the bounding box marked in each sample image in each training sample until the neural network model training end condition is met to obtain a traffic signal recognition model.

[0139] An embodiment of the present application provides a traffic signal recognition method, as Figure 3 shown, including:

[0140] Step S301, obtain a to-be-detected image at a predetermined position; the to-be-detected image includes the imaging content of the traffic signal;

[0141] Step S302, input the to-be-detected image into the traffic signal recognition model to obtain the position of the bounding box for marking the imaging content output by the traffic signal recognition model in the to-be-detected image;

[0142] Step S303, identify the color of the area within the bounding box through a preset color model, use the color of the area within the bounding box as the color of the traffic signal, and determine the state of the traffic signal according to the color of the traffic signal;

[0143] The traffic signal recognition model is trained by the method in the above embodiment.

[0144] In the embodiment of the present application, the predetermined position can be, for example, the position preset by the traffic signal. At this position, a to-be-detected image containing the imaging content of the traffic signal is collected, and the to-be-detected image is directly input into the already trained traffic signal recognition model, and the position of the bounding box for marking the imaging content output by the traffic signal recognition model in the to-be-detected image can be obtained.

[0145] After obtaining the position of the bounding box, the color of the area within the bounding box is recognized through a preset color model, and the preset model can be an RGB color model, an HIS color model, an HSV color model, etc.

[0146] The RGB (Red-Green-Blue color mode) color model obtains various colors through the changes of the three color channels of red, green, and blue and their superposition.

[0147] The HIS (Hue-Saturation-Intensity) color model describes the color characteristics with three parameters H, S, where H defines the frequency of the color and becomes the hue; S represents the depth of the color and becomes the saturation; I represents the intensity or brightness.

[0148] The HSV (Hue-Saturation-Value) color model describes the color characteristics with three parameters H, S, V, where H is defined as the hue; S represents the saturation of the color; V represents the lightness.

[0149] Of course, the preset color model in the embodiment of the present application can also be other color models in addition to the above models, and the embodiment of the present application does not limit this.

[0150] After recognizing the color of the area within the bounding box, the color of the area within the bounding box is used as the color of the traffic signal, and the state of the traffic signal is determined according to the color of the traffic signal.

[0151] The embodiment of the present application provides a possible implementation manner. The states of the traffic signal include allowing passage, warning, and prohibiting passage; determining the state of the traffic signal according to the color of the traffic signal includes:

[0152] If the color of the traffic signal is green, it is determined that the state of the traffic signal is running and allowing passage;

[0153] If the color of the traffic signal is yellow, it is determined that the state of the traffic signal is warning;

[0154] If the color of the traffic signal is red, it is determined that the state of the traffic signal is prohibiting passage.

[0155] In the application embodiment, after the traffic light recognition model identifies the position of the bounding box of the imaging content of the traffic light in the image to be tested, the color within the bounding box is directly identified, and the color within the bounding box is the color of the traffic light. The state of the traffic light is determined according to the color of the traffic light. Specifically, if the color within the bounding box is green, the state of the traffic light is determined to be allowing passage; if the color within the bounding box is red, the state of the traffic light is determined to be prohibiting passage; if the color within the bounding box is yellow, the state of the traffic light is determined to be warning. In the field of autonomous driving or assisted driving, whether passage is allowed can be determined based on the identified state of the traffic light.

[0156] A traffic light recognition method provided by an embodiment of the present application will be introduced below with a specific embodiment. A target vehicle A turns on an unmanned driving mode on a certain road. There is a traffic light at the intersection ahead. A video containing imaging content of the traffic light in front is collected in real time by a camera carried by the vehicle, and video frames of the video are extracted in real time to obtain multiple images to be tested. The multiple images to be tested are input into a traffic light recognition model to obtain the position of a bounding box output by the traffic light recognition model for detecting the imaging content of the traffic light in the image to be tested. After determining the position of the bounding box, the color of the area within the bounding box is identified by an RGB color model. The color within the bounding box area is green. It is determined that the current color of the traffic light is green, so it can be determined that the current state of the traffic light is to allow passage. The target vehicle A can directly pass through the intersection according to the determined state.

[0157] The present application embodiment provides a training device 40 for a traffic light recognition model, such as Figure 4 As shown, the training device 40 of the traffic light recognition model may include:

[0158] The training data set acquisition module 410 is used to acquire a training data set including a plurality of training samples; each training sample includes at least two sample images, and the sample images are annotated with the actual position of the bounding box of the imaging content of the traffic light;

[0159] The cropping and merging module 420 is used to crop multiple sample images, and to merge the cropped sample images after scaling and arranging them to obtain a merged image;

[0160] The first layer sub-network processing module 430 is used to input the merged image into the first layer sub-network of the initial neural network model, and output a test image after the first layer sub-network performs adaptive anchor frame and adaptive scaling on the merged image; the test image is marked with the initial positions of all bounding boxes;

[0161] The second - layer sub - network processing module 440 is configured to input the image to be tested into the second - layer sub - network of the initial neural network model, and obtain multiple feature maps output after starting to extract pixels row - by - row and column - by - column at multiple preset pixels of the second - layer neural network for the image to be tested and performing convolution processing on multiple sub - images obtained after pixel extraction;

[0162] The third - layer sub - network processing module 450 is configured to input the multiple feature maps into the third - layer sub - network of the initial neural network model, and obtain a fused feature map obtained by performing feature fusion on the multiple feature maps by the third - layer sub - network;

[0163] The fourth - layer sub - network processing module 460 is configured to input the fused feature map and the image to be tested into the fourth - layer sub - network of the initial neural network model, and obtain the target region mapped by the fused feature map matched by the fourth - layer sub - network in the image to be tested, and output the predicted position of the bounding box after adjusting the initial position of the bounding box according to the target region;

[0164] The iterative training module 470 is configured to calculate the value of the loss function of a preset neural network loss function according to the actual position and the predicted position, and perform iterative training on the initial neural network model according to the value of the loss function to obtain a traffic signal recognition model.

[0165] The traffic signal recognition model trained in the embodiments of the present application can identify the bounding box of the imaging content of the traffic signal in the image, remove the interference factors in the image, accurately identify the traffic signal, and improve the accuracy of traffic signal recognition.

[0166] The embodiments of the present application provide a possible implementation manner. The cropping and merging module includes:

[0167] The horizontal - resolution and vertical - resolution obtaining sub - module is configured to, for each sample image, obtain the horizontal resolution and the vertical resolution of the sample image;

[0168] The cropping sub - module is configured to crop the sample image according to the minimum value, so that the horizontal resolution and the vertical resolution of the cropped sample image are both multiples of a preset value; wherein, the minimum value is less than the preset value.

[0169] The embodiments of the present application provide a possible implementation manner. The cropping and merging module further includes:

[0170] The preset - resolution obtaining sub - module is configured to obtain a preset horizontal resolution and a preset vertical resolution, and both the preset horizontal resolution and the preset vertical resolution are multiples of the preset value;

[0171] A scaling sub-module, configured to determine a first scaling ratio in the horizontal direction and a second scaling ratio in the vertical direction for each cropped sample image based on a preset vertical resolution, a preset vertical resolution, the horizontal resolution and the vertical resolution of the cropped sample image; and scale the cropped sample image according to the first scaling ratio and the second scaling ratio;

[0172] A merging sub-module, configured to randomly arrange and merge the scaled sample images to obtain a merged image.

[0173] An embodiment of the present application provides a possible implementation manner. The first-layer sub-module processing module includes:

[0174] An adaptive anchor box sub-module, configured to perform an adaptive anchor box operation on the merged image to mark the initial positions of all bounding boxes in the merged image, and obtain a marked image with the initial positions of all bounding boxes marked;

[0175] An adaptive scaling sub-module, configured to adaptively scale the marked image to obtain a to-be-tested image, where the horizontal resolution and the vertical resolution of the to-be-tested image both meet multiples of a preset value.

[0176] An embodiment of the present application provides a possible implementation manner. The second-layer sub-network processing module includes:

[0177] A preset pixel determination sub-module, configured to determine a plurality of preset pixels in the to-be-tested image;

[0178] An interlaced and interleaved extraction sub-module, configured to, for any one of the preset pixels, extract pixels in an interlaced and interleaved manner starting from the preset pixel to obtain a sub-image of the to-be-tested image;

[0179] A feature map obtaining sub-module, configured to, for any one of the sub-images, perform multiple convolutional processes on the sub-image to extract the features of the sub-image and obtain a feature map corresponding to the sub-image.

[0180] An embodiment of the present application provides a possible implementation manner. The fourth-layer sub-network processing module includes:

[0181] A feature vector extraction sub-module, configured to extract a feature vector of the fused feature map; the feature vector includes pixel value features, texture features, shape features and spatial relationship features;

[0182] An adjustment sub-module, configured to determine a target area mapped by the feature vector in the to-be-tested image; and adjust the initial position of the bounding box according to the target area to obtain the predicted position of the bounding box.

[0183] An embodiment of the present application provides a possible implementation manner. The actual position is represented by the actual coordinates of all pixels on the bounding box; the predicted position is represented by the predicted coordinates of all pixels on the bounding box; the iterative training module includes:

[0184] A loss function value determination sub-module, configured to input the actual coordinates of all pixels on the bounding box and the predicted coordinates of all pixels into a preset neural network loss function to obtain a loss function value;

[0185] An iterative training sub-module, configured to, if the loss function value does not meet the neural network model training end condition, continue to iteratively train the initial neural network model based on each training sample and the actual positions of the bounding boxes marked in each sample image in each training sample until the neural network model training end condition is met, to obtain a traffic signal recognition model.

[0186] An embodiment of the present application provides a traffic signal recognition device 50, as Figure 5 shown. The traffic signal recognition device 50 may include:

[0187] A to-be-tested image acquisition module 510, configured to acquire a to-be-tested image at a predetermined position; the to-be-tested image includes imaging content of a traffic signal;

[0188] A position determination module 520, configured to input the to-be-tested image into the traffic signal recognition model to obtain the position of the bounding box for marking the imaging content output by the traffic signal recognition model in the to-be-tested image;

[0189] A color recognition module 530, configured to recognize the color of the region within the bounding box through a preset color model, use the color of the region within the bounding box as the color of the traffic signal, and determine the state of the traffic signal according to the color of the traffic signal;

[0190] The traffic signal recognition model is trained according to the above embodiment.

[0191] An embodiment of the present application provides a possible implementation manner. The states of the traffic signal include allowing passage, warning, and prohibiting passage;

[0192] The color recognition module includes a state determination sub-module, and the state determination sub-module is configured to, if the color of the traffic signal is green, determine that the state of the traffic signal is running and allowing passage;

[0193] if the color of the traffic signal is yellow, determine that the state of the traffic signal is warning;

[0194] if the color of the traffic signal is red, determine that the state of the traffic signal is prohibiting passage.

[0195] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, and details are not described herein again.

[0196] An electronic device is provided in an embodiment of the present application, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the traffic signal recognition model training method and the traffic signal recognition method. Compared with the related art, it can be achieved that: the traffic signal recognition model trained in the embodiment of the application can recognize the bounding box of the imaging content of the traffic signal in the image, remove the interference factors in the image, and can accurately recognize the traffic signal, improving the accuracy of traffic signal recognition.

[0197] In an optional embodiment, an electronic device is provided, as Figure 6 shown. Figure 6 The electronic device 6000 shown includes: a processor 6001 and a memory 6003. Among them, the processor 6001 and the memory 6003 are connected, such as connected through a bus 6002. Optionally, the electronic device 6000 may further include a transceiver 6004, and the transceiver 6004 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 6004 is not limited to one, and the structure of the electronic device 6000 does not constitute a limitation to the embodiment of the present application.

[0198] The processor 6001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 6001 may also be a combination for implementing computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0199] The bus 6002 may include a path for transmitting information among the above components. The bus 6002 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 6002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 it is only represented by a thick line in Figure 6 , but it does not mean that there is only one bus or one type of bus.

[0200] The memory 6003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0201] The memory 6003 is used to store the computer program for implementing the embodiments of the present application, and is controlled by the processor 6001 to execute. The processor 6001 is used to execute the computer program stored in the memory 6003 to implement the steps shown in the foregoing method embodiments.

[0202] Among them, the electronic device includes but is not limited to: the electronic device package can include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The illustrated electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0203] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiment can be implemented. Compared with the prior art, it can be achieved that: the traffic signal recognition model trained in the embodiment of the application can recognize the bounding box of the imaging content of the traffic signal in the image, remove the interference factors in the image, and can accurately recognize the traffic signal, improving the accuracy of traffic signal recognition.

[0204] An embodiment of the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiment can be implemented. Compared with the prior art, it can be achieved that: the traffic signal recognition model trained in the embodiment of the application can recognize the bounding box of the imaging content of the traffic signal in the image, remove the interference factors in the image, and can accurately recognize the traffic signal, improving the accuracy of traffic signal recognition.

[0205] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown or described in words.

[0206] It should be understood that although the flowcharts in the embodiments of the present application indicate each operation step by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0207] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the technical concept of the solution of the present application, other similar implementation means based on the technical idea of the present application also belong to the protection scope of the embodiments of the present application.

Claims

1. A training method for a traffic signal recognition model, characterized in that, it includes: Obtain a training data set including multiple training samples; Each training sample includes at least two sample images, and the actual positions of the bounding boxes of the imaging content of the traffic signals are marked on the sample images; Crop the multiple sample images, scale and randomly arrange the multiple cropped sample images and then merge them to obtain a merged image; Input the merged image into the first sub-network of the initial neural network model, and obtain a to-be-detected image output by the first sub-network after performing adaptive anchor boxes and adaptive scaling on the merged image; the initial positions of all bounding boxes are marked on the to-be-detected image; Input the to-be-detected image into the second sub-network of the initial neural network model, and obtain multiple feature maps output after the second neural network extracts pixels row by row and column by column at multiple preset pixels of the to-be-detected image and performs convolution processing on the multiple sub-images obtained after the pixel extraction; Input the multiple feature maps into the third sub-network of the initial neural network model, and obtain a fused feature map obtained by the third sub-network performing feature fusion on the multiple feature maps; Input the fused feature map and the to-be-detected image into the fourth sub-network of the initial neural network model, and obtain the target area mapped by the fused feature map in the to-be-detected image matched by the fourth sub-network, and output the predicted position of the bounding box after adjusting the initial position of the bounding box according to the target area; Calculate the loss function value of a preset neural network loss function according to the actual position and the predicted position, and perform iterative training on the initial neural network model according to the loss function value to obtain a traffic signal recognition model.

2. The method according to claim 1, characterized in that, The cropping of the multiple sample images includes: For each sample image, obtain the horizontal resolution and vertical resolution of the sample image; If any one of the horizontal resolution or the vertical resolution does not conform to a multiple of a preset value, calculate the minimum value of the pixels to be cropped in the horizontal direction or the vertical direction; Crop the sample image according to the minimum value, so that the horizontal resolution and vertical resolution of the cropped sample image both conform to a multiple of the preset value; wherein, the minimum value is less than the preset value.

3. The method according to any one of claims 1 or 2, characterized in that, The scaling, arranging and then merging of the multiple cropped sample images to obtain a merged image includes: Obtain a preset horizontal resolution and a preset vertical resolution, and both the preset horizontal resolution and the preset vertical resolution conform to a multiple of the preset value; Based on the preset vertical resolution, the preset vertical resolution, the horizontal resolution and vertical resolution of the cropped sample image, determine the first scaling ratio in the horizontal direction and the second scaling ratio in the vertical direction for each cropped sample image; scale the cropped sample image according to the first scaling ratio and the second scaling ratio; Randomly arrange and merge the scaled sample images to obtain a merged image.

4. The method according to claim 1, wherein, the step of inputting the merged image into the first sub-network of the initial neural network model to obtain a to-be-detected image output by the first sub-network after performing adaptive anchor boxes and adaptive scaling on the merged image includes: Performing an adaptive anchor box operation on the merged image to mark the initial positions of all bounding boxes in the merged image, obtaining a marked image marked with the initial positions of all bounding boxes; Performing adaptive scaling on the marked image to obtain a to-be-detected image, where the horizontal resolution and vertical resolution of the to-be-detected image are both multiples of a preset value.

5. The method according to claim 1, wherein, the step of inputting the to-be-detected image into the second sub-network of the initial neural network model to obtain multiple feature maps output by the second neural network after extracting pixels row by row and column by column starting from multiple preset pixels in the to-be-detected image and performing multiple convolution processes on the multiple sub-images obtained after pixel extraction includes: Determining multiple preset pixels in the to-be-detected image; For any one of the preset pixels, extracting pixels row by row and column by column starting from the preset pixel to obtain a sub-image of the to-be-detected image; For any one sub-image, performing multiple convolution processes on the sub-image to extract the features of the sub-image, obtaining the feature map corresponding to the sub-image.

6. The method according to claim 1, wherein, the step of inputting the fused feature map and the to-be-detected image into the fourth sub-network of the initial neural network model to obtain the target region mapped by the fused feature map in the to-be-detected image matched by the fourth sub-network and outputting the predicted position of the bounding box after adjusting the initial position of the bounding box according to the target region includes: Extracting the feature vector of the fused feature map; the feature vector includes pixel value features, texture features, shape features and spatial relationship features; Determining the target region mapped by the feature vector in the to-be-detected image; adjusting the initial position of the bounding box according to the target region to obtain the predicted position of the bounding box.

7. The method according to claim 1, wherein, the actual position is represented by the actual coordinates of all pixels on the bounding box; the predicted position is represented by the predicted coordinates of all pixels on the bounding box; the step of calculating the loss function value of a preset neural network loss function according to the actual position and the predicted position and iteratively training the initial neural network model according to the loss function value to obtain a traffic signal recognition model includes: Input the actual coordinates of all pixels on the bounding box and the predicted coordinates of all pixels into the preset neural network loss function to obtain a loss function value; If the loss function value does not meet the neural network model training end condition, continue to iteratively train the initial neural network model based on each of the training samples and the actual positions of the bounding boxes marked in each sample image in each training sample until the neural network model training end condition is met, and obtain a traffic signal recognition model.

8. A method for recognizing traffic signals, Characterized in that, It includes: Obtain a to-be-detected image at a predetermined position; The to-be-detected image contains the imaging content of the traffic signal; Input the to-be-detected image into the traffic signal recognition model to obtain the position of the bounding box for marking the imaging content output by the traffic signal recognition model in the to-be-detected image; Identify the color of the area within the bounding box through a preset color model, use the color of the area within the bounding box as the color of the traffic signal, and determine the state of the traffic signal according to the color of the traffic signal; The traffic signal recognition model is trained according to the method described in any one of claims 1 to 7.

9. According to the method described in claim 8, Characterized in that, The states of the traffic signal include allowing passage, warning, and prohibiting passage; determining the state of the traffic signal according to the color of the traffic signal includes: If the color of the traffic signal is green, determine that the state of the traffic signal is running and allowing passage; If the color of the traffic signal is yellow, determine that the state of the traffic signal is warning; If the color of the traffic signal is red, determine that the state of the traffic signal is prohibiting passage.

10. A training device for a traffic signal recognition model, Characterized in that, It includes: An acquisition module, configured to acquire a training data set including a plurality of training samples; each training sample includes at least two sample images, and the actual positions of the bounding boxes of the imaging content of the traffic signal are marked in the sample images; A cropping, scaling, arranging, and merging module, configured to crop the multiple sample images, scale and randomly arrange the multiple cropped sample images and then merge them to obtain a merged image; The initial positions of all bounding boxes are marked in the to-be-detected image; A to-be-detected image acquisition module, configured to input the merged image into the first-layer sub-network of the initial neural network model to obtain a to-be-detected image output by the first-layer sub-network after performing adaptive anchor boxes and adaptive scaling on the merged image; A feature map acquisition module, configured to input the to-be-detected image into the second-layer sub-network of the initial neural network model to obtain multiple feature maps output after the second-layer neural network extracts pixels row by row and column by column starting from multiple preset pixels in the to-be-detected image and performing multiple convolutional processes on the multiple sub-images obtained after pixel extraction; It should be noted that in the original text, the content in has an error. It should be "If the color of the traffic signal is red" instead of "If the color of the traffic signal is yellow" to correctly represent the condition for determining the "prohibiting passage" state. The above translation has been corrected accordingly. A fused feature map obtaining module, configured to input the multiple feature maps into a third sub-network of the initial neural network model, and obtain a fused feature map obtained by the third sub-network after performing feature fusion on the multiple feature maps; A predicted position output module, configured to input the fused feature map and the image to be measured into a fourth sub-network of the initial neural network model, obtain a target region mapped by the fused feature map matched by the fourth sub-network in the image to be measured, and output a predicted position of the bounding box after adjusting an initial position of the bounding box according to the target region; An iterative module, configured to calculate a loss function value of a preset neural network loss function according to the actual position and the predicted position, and perform iterative training on the initial neural network model according to the loss function value to obtain a traffic signal recognition model.

11. An apparatus for recognizing traffic signals, characterized in that it includes: An image to be measured obtaining module, configured to obtain an image to be measured at a predetermined position; the image to be measured includes imaging content of a traffic signal; A bounding box position recognition module, configured to input the image to be measured into a traffic signal recognition model, and obtain a position of a bounding box for marking imaging content output by the traffic signal recognition model in the image to be measured; A state recognition module, configured to recognize a color of a region within the bounding box through a preset color model, use the color of the region within the bounding box as the color of the traffic signal, and determine a state of the traffic signal according to the color of the traffic signal; The traffic signal recognition model is trained according to the method described in any one of claims 1 to 7.

12. An electronic device, including a memory, a processor, and a computer program stored on the memory, characterized in that the processor executes the computer program to implement the steps of the method described in any one of claims 1-9.

13. A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method described in any one of claims 1-9 are implemented.

14. A computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the steps of the method described in any one of claims 1-9 are implemented.

Citation Information

Patent Citations

  • Remote sensing scene classification method based on attention network scale feature fusion

    CN113408594A