Method, device and storage medium for identifying landmarks based on regions of interest
Through the road sign recognition method of the area of interest, the feature extraction network and unified operation of size is used to solve the problem of low recognition accuracy of traffic signs in harsh environments, and higher recognition accuracy and recall rate are achieved.
Patent Information
- Application Number
- CN202011507550.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-18
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2040-12-18
AI Technical Summary
The existing traffic sign recognition technology has a low recognition accuracy in application scenarios such as rainy days, nights, and signs that are blocked or signs blurred.
The road sign recognition method based on the region of interest is adopted, and the road sign feature map in the image is extracted through the feature extraction network, multiple CBM operations and convolution operations are performed, combined with unified size operations, to improve feature expression capabilities, support dynamic changes in image resolution, and avoid the influence of image distortion.
It improves the recognition accuracy and recall rate of road traffic signs in harsh environments, enhances the expressive ability of feature maps, and improves the extraction quality of road sign features.
Smart Images

Figure CN112560708B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence detection technology, which is applied in smart transportation, and specifically to a method, device and storage medium for identifying road signs based on regions of interest. Background Art
[0002] Traffic signs are important road safety facilities, and their automatic detection and recognition has become one of the important technical links in the intelligent transportation system, and has attracted more and more attention from researchers.
[0003] The existing traffic sign recognition application scenarios (such as weather) are relatively high. In some specific application scenarios, such as rainy days, nighttime, when the sign is blocked, or when the sign is blurred, the recognition accuracy is relatively low. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, and storage medium for identifying landmarks in a region of interest (ROI) that can improve the accuracy of specific application scenarios.
[0005] In a first aspect, an embodiment of the present application provides a method for identifying road signs based on a region of interest, which is applied to an electronic device. The method includes:
[0006] The acquired original image is passed through a feature extraction network to extract a first feature map of the landmarks in the image;
[0007] Outputting the first feature map into an output matrix of fixed size through a resizing operation;
[0008] Performing a convolution operation on the output matrix after performing multiple CBM operations to obtain a second feature map;
[0009] The second feature map is used to detect and identify landmarks to obtain the category and coordinates of the reference frame, and the coordinates are mapped to the original input image to determine the category of the landmark corresponding to the original input image.
[0010] The technical solution provided by this application provides a road traffic sign recognition device based on a region of interest, which is applied to an electronic device, and the device includes:
[0011] an extraction unit, configured to extract a first feature map of a landmark in the acquired original image through a feature extraction network;
[0012] a resizing unit, configured to output the first feature map into an output matrix of a fixed size through a resizing operation;
[0013] an operation unit, configured to perform a convolution operation on the output matrix after performing multiple CBM operations to obtain a second feature map;
[0014] The recognition unit is used to detect and recognize landmarks in the second feature map to obtain the category and coordinates of the reference frame, and map the coordinates to the original input image to determine the category of the landmark corresponding to the original input image.
[0015] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the program comprises instructions for executing the steps in the first aspect of the embodiment of the present application.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the above-mentioned computer-readable storage medium stores a computer program for electronic data exchange, wherein the above-mentioned computer program enables a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application.
[0017] In a fifth aspect, embodiments of the present application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform some or all of the steps described in the first aspect of the embodiments of the present application. The computer program product may be a software installation package.
[0018] The implementation of the embodiments of this application has the following beneficial effects:
[0019] It can be seen that the technical solution provided by the present application adopts an extraction network to better extract the original image to obtain a feature map with better quality, and then based on the first feature map with better quality, the first feature map can better represent the pixel changes in harsh environments, and the first feature map is operated to obtain a second feature map. The first feature map with better quality greatly enhances the feature expression ability of the feature map in difficult scenarios such as rainy days, nighttime, obstructed signs, blurred signs, etc., thereby improving the overall accuracy and recall rate of road traffic sign recognition; in addition, the operation of size unification is added, which can support dynamic changes in image resolution and avoid the influence of image distortion on recognition results, effectively improving the extraction quality of road signs (such as road traffic signs), thereby improving the overall accuracy and recall rate of road signs (such as road traffic signs). BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0022] Figure 2 This is a flow chart of a method for identifying road traffic signs based on regions of interest provided by an embodiment of the present application;
[0023] Figure 3a This is a flow chart of feature extraction provided by an embodiment of the present application;
[0024] Figure 3b This is a schematic diagram of the CBM process provided in the embodiment of the present application;
[0025] Figure 3c This is a schematic diagram of the CResX process provided in an embodiment of the present application;
[0026] Figure 3d 1 is a flow chart of CRes1 provided in an embodiment of the present application;
[0027] Figure 4 This is a schematic diagram of a road traffic sign flow based on an area of interest provided by an embodiment of the present application;
[0028] Figure 5 This is a structural diagram of a road traffic sign device based on an area of interest provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0030] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0031] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0032] The technical solutions provided in the embodiments of the present application are executed based on the hardware structure of the electronic device. The electronic device may be a portable electronic device that also includes other functions such as a personal digital assistant and / or a music player function, such as a mobile phone, a tablet computer, a wearable electronic device with a wireless communication function (such as a smart watch), etc. Of course, fixed electronic devices may also be included, such as traffic cameras, surveillance cameras, and the like. Exemplary embodiments of portable electronic devices and fixed electronic devices include but are not limited to portable electronic devices or fixed electronic devices equipped with an IOS system, an Android system, a Microsoft system, or other operating systems (such as an embedded operating system). The above-mentioned portable electronic device may also be other portable electronic devices, such as a laptop computer (Laptop), etc. It should also be understood that in some other embodiments, the above-mentioned electronic device may not be a portable electronic device, but a desktop computer or a fixed camera.
[0033] An electronic device is a device deployed in indoor or outdoor environments to transmit and receive signals. For example, a signal transceiver device can be an evolved Node B (eNB), a radio network controller (RNC), a Node B (NB), a base station controller (BSC), a base transceiver station (BTS), a home base station (e.g., home evolved Node B, or HNB), an access controller (AC), a Wi-Fi access point (AP), etc.
[0034] The software and hardware operating environment of the technical solution disclosed in this application is introduced as follows.
[0035] For example, Figure 11 shows a schematic diagram of the structure of an electronic device 100. The electronic device may be a smart camera. Of course, the electronic device may also be a wearable device, specifically a smart watch, a smart bracelet, or other wearable portable device. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone jack 170D, a sensor module 180, a compass 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195.
[0036] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent components or integrated into one or more processors. In some embodiments, the electronic device 100 may also include one or more processors 110. The controller may generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. In other embodiments, the processor 110 may also include a memory for storing instructions and data. Exemplarily, the memory in the processor 110 may be a cache memory. This memory may store instructions or data that have just been used or are being recycled by the processor 110. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the electronic device 100 in processing data or executing instructions.
[0037] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface. Among them, the USB interface 130 is an interface that complies with the USB standard specification, and specifically can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 101, and can also be used to transmit data between the electronic device 101 and peripheral devices. The USB interface 130 can also be used to connect headphones to play audio through the headphones.
[0038] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0039] See Figure 2 , Figure 2 This application provides a road sign recognition algorithm based on a region of interest. The region of interest (ROI) is a part of the image. The part can be a task-related area, such as the area where the road sign is located in the image. The region of interest can be extracted by an extraction network. The method can be performed as follows: Figure 1 The electronic device shown executes the method as shown in Figure 2 Shown, including:
[0040] Step S201: extracting a first feature map of the road signs in the acquired original image through a feature extraction network;
[0041] In an optional embodiment, the feature extraction network may be DarkNet98. The DarkNet98 is used to extract features of the original image. Its specific architecture is as follows: Figure 3a shown.
[0042] The above step S201 may specifically include:
[0043] After performing CBM on the original image, the CBM result is obtained. Then, the CBM result is subjected to 6 cross-layer residual connection (CResX) operations to obtain the first feature map of the landmarks in the image. The X in the 6 CResX operations can be different, and X is the number of Res units (residual units) in the CRes operation. The specific structure of CResX is as follows: Figure 3c shown.
[0044] The above CBM is formed by the initials of three calculations: convolution, batch normalization (BN), and self-adjusting non-monotonic neural activation function (Mish). The specific architecture of CBM can be as follows: Figure 3b As shown, it can specifically include: performing convolution operations, batch normalization, and self-adjusting non-monotonic neural activation function (Mish) activation function on the CBM input data in sequence to obtain the CBM result.
[0045] See Figure 3c , the CResX operation may specifically include: dividing the CResX input data into two paths to perform operations to obtain two results, and concat (joining) the two results to obtain the CResX operation result. Figure 3c As shown, the first-path operation may be to perform a residual connection (Resx) operation on the input data of CResX to obtain a residual connection result, and then perform a CBM operation on the residual connection result to obtain a first-path result. The residual connection may include: after performing a CBM operation once, performing x residual unit operations on the CBM operation result to obtain a residual connection result. The second-path operation is to perform a CBM operation on the input data of CResX to obtain a second-path result of the two-path result.
[0046] In an optional solution, the operation diagram of the DarkNet98 network model extracting the first feature map of the landmark in the image is as follows: Figure 3a As shown, the specific values of X in the 6 CResX can be: 1, 2, 8, 16, 8, 4.
[0047] Specifically, the first feature map may be a feature map matrix (i.e., the first feature map) obtained by performing multiple feature extraction operations (i.e., 6 times CResX) on the original input image using the DarkNet98 network model. The feature map matrix includes landmarks.
[0048] The above CBM Figure 3bAs shown, it can specifically include: performing a convolution operation to obtain a convolution result, performing a batch normalization operation on the convolution result to obtain a batch normalization result, and performing a self-adjusting non-monotonic neural activation function (Mish) operation on the batch normalization result to obtain a CBM result.
[0049] The Mish operation is a self-adjusting non-monotonic neural activation function operation. The above CResX operation is as follows Figure 3c As shown, this may include:
[0050] If X=1 in CResX, the specific method can be as follows Figure 3d Specifically, it may include:
[0051] The CBM result is divided into two paths. The first path may include: performing three CBM operations on the CBM result, performing a weighted sum (the weight can be 1) on the sum to obtain the first path result; the second path may include: performing one CBM operation on the CBM result to obtain the second path result, and concatenating the first path result and the second path result to obtain the result of CRes1. The above concatenation may specifically include: concatenating only the number of channels of the two paths, and not performing concatenation on the height and width.
[0052] Specific splicing may include, for example, the first result is: 16*32*10, where 10 represents the number of channels, and the second result is: 16*32*22, then the splicing result of CRes1 may be: 16*32*32.
[0053] If X is a value greater than 1, the number of Res units is increased by a corresponding amount. For example, when X=2, it may specifically include:
[0054] Similar to the operation when X=1, when X=2, the second-path operation will not change, and the first-path operation will change. When X=2, for the first-path operation, the CBM result is weighted summed after executing 3 CBM operations (the weight can be 1), and the sum is weighted summed again after executing 2 CBM operations to obtain the sum and the CBM operation is executed to obtain the first-path result.
[0055] Step S202: Output the first feature map into an output matrix of a fixed size through a resizing operation;
[0056] The above-mentioned size operation may specifically include: a region of interest alignment (RoIAlign) operation. The RoIAlign is a regional feature aggregation operation performed on the region of interest. In this application, the regional feature may specifically be a landmark.
[0057] The specific implementation method may include: performing a Block operation on the feature map, which may include: dividing it into blocks of a fixed size (specifically 416*416), each block size is (1600 / 416)*(1400 / 416)=3.85*3.37, and then using bilinear interpolation to perform a Sample (pixel sampling, assuming the number of sampling points is set to 6, which is equivalent to changing 3.85*3.37=12.97 pixels to 6 pixels) operation on each block, and finally performing a MaxPooling operation on each block to obtain a fixed-size output matrix.
[0058] The fixed size may be other sizes. The fixed size may be determined by pre-configuration. In practical applications, it may also be configured by higher-layer signaling, including but not limited to radio resource control (RRC) signaling or media access control unit (MAC CE).
[0059] Step S203: Perform multiple CBM operations on the output matrix and then perform convolution operations to obtain a second feature map y1;
[0060] The above CBM Figure 3b As shown, it can specifically include: performing a convolution operation to obtain a convolution result, performing a BN (batch normalization) operation on the convolution result to obtain a batch normalization result, and performing a Mish operation on the batch normalization result to obtain a CBM result.
[0061] Compared with the above steps, the input data of the CBM operation here is different, and the output result is also different, that is, the input data here is the output matrix.
[0062] Step S204: Detect and identify landmarks on the second feature map (y1) to obtain the category and coordinates of the anchor box, and map the coordinates to the original input image to determine the category of the landmark corresponding to the original input image.
[0063] The category of the above-mentioned reference frame can be determined by existing methods. For example, a softmax function operation can be performed on the data in the reference frame to determine the soft result, and the category corresponding to the reference frame is determined based on the range of the softmax result. For example, if the softmax result belongs to the first range, and the first range prohibits non-motor vehicles from passing, then the category is determined to be non-motor vehicle passage.
[0064] The technical solution provided in the present application adopts an extraction network (such as DarkNet98) to better extract the original image to obtain a first feature map with better quality, and then performs an operation based on the first feature map with better quality to obtain a second feature map (y1). The good quality first feature map greatly enhances the feature expression ability of the feature map in difficult scenarios such as rainy days, nighttime, obstructed signs, blurred signs, etc., thereby improving the overall accuracy and recall rate of road sign (such as road traffic signs) recognition; in addition, the size unification operation is added to support the dynamic change of image resolution and avoid the influence of image distortion on the recognition results, effectively improving the extraction quality of road sign (such as road traffic signs) features, thereby improving the overall accuracy and recall rate of road sign (such as road traffic signs) recognition.
[0065] Softmax function, softmax is used in the multi-classification process. It maps the output of multiple neurons to the interval (0, 1), which can be understood as probability, so as to perform multi-classification.
[0066] Suppose we have an array, V, Vi represents the i-th element in V, then the softmax value of this element specifically includes:
[0067]
[0068] The above-mentioned mapping of the coordinates onto the original input image to determine the category of the road sign (traffic sign) corresponding to the original input image can specifically include: after determining the zoom factor of the reference frame, enlarging the reference frame by a corresponding factor (i.e., the inverse of the zoom factor, for example, if it is zoomed 1 / 16, it is enlarged 16 times), mapping the enlarged reference frame (due to the reverse enlargement, the size of the reference frame is now consistent with the size of the original input image) to the coordinate position corresponding to the original input image, determining the coordinate position as the position of the road sign in the original input image, and determining the category of the reference frame as the category of the road sign.
[0069] The identification of the above categories can be achieved through a general classifier, which includes but is not limited to: a support vector machine, a deep neural network model, etc.
[0070] In an optional solution, the above method may further include, before step S204:
[0071] The output matrix is subjected to multiple CBM operations and then up-sampled to obtain a sampling result, and the sampling result is weightedly summed with the first feature map of the road sign (such as a road traffic sign). The obtained sum is subjected to multiple CBM operations and convolution operations to obtain the third feature map y2.
[0072] The third feature map y2 is used to detect and identify road signs (such as road traffic signs) to obtain the category and coordinates of the reference frame, and the coordinates are mapped to the original input image to determine the category and coordinates of the traffic sign corresponding to the original input image.
[0073] The specific implementation method may include: after determining the scaling factor of the reference frame, enlarging the reference frame of the third feature map y2 by a corresponding factor (i.e., the inverse of the scaling factor, where the scaling factor is generally larger than that of the second feature map, for example, 1 / 32, then the magnification factor is 32 times), and then mapping the enlarged reference frame to the coordinate position corresponding to the original input image as the position of the signpost in the original input image, and determining that the category of the reference frame is the category of the signpost.
[0074] In an optional solution, the above method may further include:
[0075] Perform weighted summation on the up-sampled sampling result and the first feature map, perform multiple CBM operations on the obtained sum, and perform secondary upsampling to obtain a secondary sampling result, perform weighted summation on the secondary sampling result and the first feature map, perform multiple CBM operations and convolution operations on the obtained sum, and obtain a fourth feature map y3;
[0076] The fourth feature map y3 is used to detect and identify road signs to obtain the category and coordinates of the reference frame, and the coordinates are mapped to the original input image to determine the category of the traffic sign corresponding to the original input image.
[0077] The technical solution of the present application adopts a weighted summation method, which can more effectively integrate shallow and deep features, improve the overall feature quality of road traffic signs, and thus further improve the overall accuracy and recall rate of road traffic sign recognition.
[0078] The identification of the above categories can be achieved through a general classifier, which includes but is not limited to: a support vector machine, a deep neural network model, etc.
[0079] The specific implementation method may be similar to the third feature map y2 and the second feature map y1.
[0080] The identification of the above categories can be achieved through a general classifier, which includes but is not limited to: a support vector machine, a deep neural network model, etc.
[0081] The specific process of implementing the above y1, y2, and y3 is as follows Figure 4As shown in the figure, CBM (i.e., convolution + batch normalization + Mish activation function), Conv (convolution), Upsample (upsampling) and other operations are performed multiple times to obtain feature maps y1, y2, and y3 of three different scales. Among them, feature maps y2 and y3 mainly have an additional Sum operation compared to y1 (i.e., the feature map obtained in the previous layer is upsampled and then weighted summed with the corresponding intermediate layer in the DarkNet98 network (weights are set to 0.6 and 0.4 respectively)). The above weights can also be other values, for example, weights of 0.5 and 0.5 respectively.
[0082] The above-mentioned method of obtaining the reference frame can specifically include: using the 9 anchor boxes (reference boxes) obtained by clustering the feature maps y1, y2, and y3 using the k-means clustering algorithm in advance to perform road traffic sign detection and recognition on the three feature maps of different scales y1, y2, and y3, and predicting the coordinates (i.e., sign positions) and categories (i.e., 50 types of road traffic signs: no non-motor vehicles allowed, no pedestrians allowed, no entry, no honking, etc.) of 3 different anchor boxes on each feature map.
[0083] The k-means algorithm constructs k partitioning clusters from a given dataset of n data objects, each of which is a cluster. This method divides the data into n clusters, each containing at least one data object. Each data object must belong to exactly one cluster. Furthermore, data objects within the same cluster must be highly similar, while data objects in different clusters must be less similar. Cluster similarity is calculated using the mean of the objects within each cluster.
[0084] The k-means algorithm's processing flow includes the following: First, k data objects are randomly selected, each representing a cluster center, i.e., k initial centers are selected. For each remaining object, based on its similarity (distance) to each cluster center, it is assigned to the cluster corresponding to the cluster center with which it is most similar. Then, the average value of all objects in each cluster is recalculated, and this is used as the new cluster center. This process is repeated until the criterion function converges, meaning that the cluster centers do not change significantly. The mean square error (MSE) is typically used as the criterion function, which minimizes the sum of the squares of the distances from each point to the nearest cluster center.
[0085] The new cluster center calculation method is to calculate the average value of all objects in the cluster. In other words, the average value of each dimension of all objects is taken to get the cluster center. For example, if a cluster contains the following three data objects {(6,4,8), (8,2,2), (4,6,2)}, then the cluster center is ((6+8+4) / 3, (4+2+6) / 3, (8+2+2) / 3) = (6,4,4).
[0086] The k-means algorithm uses distance to describe the similarity between two data objects. Distance functions include Ming-style distance, Euclidean distance, Mahalanobis distance, and Langmuir distance, with Euclidean distance being the most commonly used.
[0087] The k-means algorithm terminates when the criterion function reaches the optimal value or the maximum number of iterations. When using Euclidean distance, the criterion function is generally to minimize the sum of the squares of the distances from the data object to its cluster center.
[0088] In an optional solution, the technical solution of the present application can be configured to set up a separate AI chip to perform a convolution operation, which can be a multi-layer convolution operation (or a multi-stage convolution operation). The AI chip includes: an allocation calculation processing circuit and x calculation processing circuits. The AI chip obtains the matrix size CI*CH of the input data. If the convolution kernel size in the n-layer convolution operation is a 3*3 convolution kernel, the allocation calculation processing circuit divides CI*CH into CI / x data blocks in the CI direction (assuming CI is an integer of x), and allocates the CI / x data blocks to x calculation processing circuits in sequence. The x calculation processing circuits respectively receive 1 allocated data block and perform the i-th layer convolution operation with the i-th layer convolution kernel to obtain the i-th convolution result (that is, the x result matrices (CI / x-2)*(CH-2) of the x calculation processing circuits are combined in sequence to obtain the i-th convolution result), and the results of the two edge columns of the i-th convolution result (the two columns calculated by different calculation processing circuits for the results of adjacent columns are determined as edge columns) are sent to The distribution processing circuit, the x calculation processing circuits perform a convolution operation on the i-th layer convolution result and the (i+1)-th layer convolution kernel to obtain the (i+1)-th convolution result, and send the (i+1)-th convolution result to the distribution calculation circuit. The distribution calculation processing circuit performs a convolution operation on the (CI / x-1) combined data blocks and the i-th layer convolution kernel to obtain the i-th combination result, and splices the i-th combination result with the results of the two edge columns of the i-th convolution result (insert the i-th combination result into the edge two columns according to the mathematical rules of the convolution operation). The (i+1)th combined data block is obtained by performing a convolution operation on the (i+1)th combined data block and the (i+1)th convolution kernel to obtain the (i+1)th combined result, and the (i+1)th combined result is inserted between the edge columns of the (i+1)th convolution result (the results of adjacent columns are calculated by different calculation processing circuits) to obtain the (i+1)th layer convolution result. The AI chip performs the remaining convolution layer (the convolution kernel after the i+1 layer) operation based on the (i+1)th layer convolution result to obtain the nth layer convolution operation result. The above-mentioned combined data block can be a 4*CI matrix composed of 4 columns of data between 2 adjacent data blocks, for example, a 4*CH matrix composed of the last 2 columns of the first data block (the data block assigned by the first calculation processing circuit) and the first 2 columns of the second data block (the data block assigned by the second calculation processing circuit).
[0089] The operations of the remaining convolutional layers mentioned above can also refer to the calculations of the i-th layer and the (i+1)-th layer, where i is an integer ≥1 and less than or equal to n, the above n is the total number of convolutional layers of the AI model, i is the layer number of the convolutional layer, CI is the column value of the matrix, and CH is the row value of the matrix.
[0090] Setting up a separate AI chip to perform convolution operations can increase the speed of convolution operations and reduce IO overhead, so it has the advantages of saving costs and reducing power consumption.
[0091] It is understandable that, in order to implement the above functions, the electronic device includes hardware and / or software modules that perform the corresponding functions. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to be beyond the scope of this application.
[0092] In this embodiment, the electronic device can be divided into functional modules according to the above-mentioned method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into a single processing module. The above-mentioned integrated modules can be implemented in the form of hardware. It should be noted that the module division in this embodiment is illustrative and is only a logical functional division. In actual implementation, other division methods may be used.
[0093] In the case of dividing each functional module into corresponding functional modules, Figure 5 A road traffic sign recognition device based on a region of interest is shown, which is applied to an electronic device. The device includes:
[0094] An extraction unit 501 is configured to extract a first feature map of a road sign from an acquired original image through a feature extraction network;
[0095] A resizing unit 502 is configured to output the first feature map into an output matrix of a fixed size through a resizing operation;
[0096] An operation unit 503 is configured to perform a convolution operation on the output matrix after performing multiple CBM operations to obtain a second feature map y1;
[0097] The recognition unit 504 is used to detect and recognize landmarks in the second feature map y1 to obtain the category and coordinates of the reference frame, and map the coordinates to the original input image to determine the category of the landmark corresponding to the original input image.
[0098] The extraction unit 501 may be used to support the electronic device in executing the above step 201 and / or other processes for the technology described herein, and the resizing unit 502 may be used to support the electronic device in executing the above step 202 and / or other processes for the technology described herein.
[0099] The computing unit 503 may be used to support the electronic device in executing the above-mentioned step 203 and / or other processes of the technology described herein.
[0100] The identification unit 504 may be used to support the electronic device in executing the above-mentioned step 204 and / or other processes of the technology described herein.
[0101] It should be noted that all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0102] The electronic device provided in this embodiment is used to execute the above-mentioned location determination method, and thus can achieve the same effect as the above-mentioned implementation method.
[0103] When integrated, the electronic device may include a processing module, a storage module, and a communication module. The processing module may be used to control and manage the electronic device's operations. For example, it may be used to support the electronic device in executing the steps performed by the extraction unit 501, the resizing unit 502, the computing unit 503, and the recognition unit 504. The storage module may be used to support the electronic device in executing and storing program code and data. The communication module may be used to support communication between the electronic device and other devices.
[0104] The processing module may be a processor or a controller. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, and so on. The storage module may be a memory. The communication module may specifically be a device that interacts with other electronic devices, such as a radio frequency circuit, a Bluetooth chip, or a Wi-Fi chip.
[0105] In one embodiment, when the processing module is a processor and the storage module is a memory, the electronic device involved in this embodiment may be a Figure 1 Device with the structure shown.
[0106] An embodiment of the present application further provides an electronic device, including a processor and a memory, wherein the memory is used to store one or more programs and is configured to be executed by the processor, wherein the programs include instructions for performing the following steps, which may specifically include:
[0107] The acquired original image is passed through a feature extraction network to extract a feature map of road traffic signs in the image;
[0108] Unify the feature map into a fixed-size output matrix through size operations;
[0109] Perform multiple CBM operations on the output matrix and then perform convolution operations to obtain a feature map y1;
[0110] The feature map y1 is used to detect and identify road traffic signs to obtain the category and coordinates of the reference frame, and the coordinates are mapped to the original input image to determine the category of the traffic sign corresponding to the original input image.
[0111] The instructions executed by the above program may also include: Figure 2 The detailed solution of step S201, step S202, step S203 and step S204 is shown. Of course, the above program instructions can also include the following Figure 2 The optional solutions in the illustrated embodiment will not be described in detail here.
[0112] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute part or all of the steps of any method described in the above method embodiments.
[0113] The present application also provides a computer program product comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package.
[0114] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0115] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0116] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0117] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0118] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0119] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the above-mentioned methods of each embodiment of the present application. The aforementioned memory includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0120] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0121] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A method for identifying road signs based on regions of interest, characterized in that: Applied to electronic equipment, the method includes: Perform CBM on the acquired original image to obtain a CBM result, and perform CRes1, CRes2, CRes8, CRes16, and CRes4 operations on the CBM result in sequence to obtain a first feature map of the landmarks in the image; the CRes1 result is obtained as follows: perform CBM operations on the CBM result three times and then perform a weighted sum on the sum to obtain a first-path result, perform a CBM operation on the CBM result once to obtain a second-path result, and splice the first-path result and the second-path result to obtain a CRes1 result. The splicing is used to splice the number of channels of the first-path result and the second-path result, and the splicing operation is not performed on the height and width; Outputting the first feature map into a fixed-size output matrix through a resizing operation, wherein the resizing operation includes a region of interest alignment algorithm operation, wherein the region of interest alignment algorithm operation includes: dividing the feature map into fixed-size blocks, performing a pixel sampling operation on each block using a bilinear interpolation method, and performing a maximum pooling operation on each block to obtain a fixed-size output matrix; Performing a convolution operation on the output matrix after performing multiple CBM operations to obtain a second feature map; Performing multiple CBM operations on the output matrix and then performing upsampling to obtain a sampling result, performing weighted summation on the sampling result and the first feature map, and performing multiple CBM operations and convolution operations on the obtained sum again to obtain a third feature map; Performing a weighted summation on the sampling result and the first feature map, performing multiple CBM operations and secondary upsampling on the obtained sum to obtain a secondary sampling result, performing a weighted summation on the secondary sampling result and the first feature map, performing multiple CBM operations and convolution operations on the obtained sum to obtain a fourth feature map; The second feature map, the third feature map, and the fourth feature map are respectively used to detect and identify road traffic signs to obtain categories and coordinates of reference frames, and the coordinates are mapped to the original input image to determine the category of the road sign corresponding to the original input image; wherein the method of obtaining the reference frame includes: clustering the second feature map, the third feature map, and the fourth feature map using the k-means clustering algorithm in advance to obtain multiple reference frames, and performing road traffic sign detection and recognition on the obtained second feature map, the third feature map, and the fourth feature map, three feature maps of different scales, respectively, and predicting the coordinates and categories of multiple different reference frames on each feature map, wherein the categories are used to indicate road traffic signs, and the road traffic signs include no non-motor vehicles allowed, no pedestrians allowed, no entry, and no honking; Among them, the convolution operation is performed by a separately configured AI chip, and the convolution operation is a multi-layer convolution operation. The AI chip includes: an allocation calculation processing circuit and x calculation processing circuits. The AI chip is used to obtain the matrix size CI*CH of the input data. If the convolution kernel size in the n-layer convolution operation is a 3*3 convolution kernel, the allocation calculation processing circuit divides CI*CH into CI / x data blocks in the CI direction, and allocates the CI / x data blocks to the x calculation processing circuits in sequence. The x calculation processing circuits respectively perform the i-th convolution operation on the received 1 data block and the i-th convolution kernel to obtain the i-th convolution result, and send the results of the two edge columns of the i-th convolution result to the allocation processing circuit. The x calculation processing circuits perform a convolution operation on the i-th convolution result and the (i+1)-th convolution kernel to obtain the (i+1)-th convolution result, and the (i+1)-th convolution result The data is sent to the distribution calculation circuit. The distribution calculation processing circuit performs the i-th convolution operation on the (CI / x-1) combined data blocks and the i-th convolution kernel to obtain the i-th combination result, splices the i-th combination result with the results of the two edge columns of the i-th convolution result to obtain the (i+1)-th combined data block, performs a convolution operation on the (i+1)-th combined data block and the (i+1)-th convolution kernel to obtain the (i+1)-th combination result, inserts the (i+1)-th combination result between the edge columns of the (i+1)-th convolution result to obtain the (i+1)-th layer convolution result, and the AI chip performs the remaining convolution layer operation based on the (i+1)-th layer convolution result to obtain the n-th layer convolution operation result; the combined data block is a 4*CI matrix composed of 4 columns of data between 2 adjacent data blocks; n is the total number of convolution layers corresponding to the AI chip, i is the layer number of the convolution layer, CI is the column value of the matrix, and CH is the row value of the matrix.
2. The method according to claim 1, characterized in that The CBM operation includes: convolution operation, batch normalization BN and self-adjusting non-monotonic neural activation function Mish activation function.
3. The method according to claim 1, characterized in that Mapping the coordinates to the original input image to determine the category of the landmark corresponding to the original input image specifically includes: After determining the zoom factor of the reference frame, the reference frame is enlarged by a reciprocal of the corresponding zoom factor to obtain an enlarged reference frame; Map the enlarged reference frame to the coordinate position corresponding to the original input image; The coordinate position is determined to be the coordinate position of the road sign in the original input image, and the category of the reference frame is determined to be the category of the road sign.
4. A road sign recognition device based on region of interest, characterized in that: Applied to electronic equipment, the device comprises: An extraction unit is configured to perform CBM on the acquired original image to obtain a CBM result, and sequentially perform CRes1, CRes2, CRes8, CRes16, and CRes4 operations on the CBM result to obtain a first feature map of the landmark in the image; a resizing unit, configured to output the first feature map into an output matrix of a fixed size through a resizing operation; the resizing operation includes a region of interest alignment algorithm operation, the region of interest alignment algorithm operation including: dividing the first feature map into blocks of a fixed size, performing a pixel sampling operation on each block using a bilinear interpolation method, and performing a maximum pooling operation on each block to obtain an output matrix of a fixed size; an operation unit, configured to perform a convolution operation on the output matrix after performing multiple CBM operations to obtain a second feature map; The recognition unit is used to perform multiple CBM operations on the output matrix and then perform upsampling to obtain a sampling result, perform weighted summation on the sampling result and the first feature map, perform multiple CBM operations and convolution operations on the obtained sum again to obtain a third feature map; perform weighted summation on the sampling result and the first feature map, perform multiple CBM operations and secondary upsampling on the obtained sum again to obtain a secondary sampling result, perform weighted summation on the secondary sampling result and the first feature map, perform multiple CBM operations and convolution operations on the obtained sum to obtain a fourth feature map; perform detection and recognition of road traffic signs on the second feature map, the third feature map and the fourth feature map respectively to obtain a reference The category and coordinates of the frame are mapped to the original input image to determine the category of the road sign corresponding to the original input image; wherein the method of obtaining the reference frame includes: using the second feature map, the third feature map, and the fourth feature map to obtain multiple reference frames by pre-clustering the second feature map, the third feature map, and the fourth feature map using the k-means clustering algorithm, respectively, performing road traffic sign detection and recognition on the obtained second feature map, the third feature map, and the fourth feature map at three different scales, respectively, and predicting the coordinates and categories of multiple different reference frames on each feature map, wherein the categories are used to indicate road traffic signs, and the road traffic signs include no non-motorized vehicles allowed, no pedestrians allowed, no entry, and no honking; Among them, the convolution operation is performed by a separately configured AI chip, and the convolution operation is a multi-layer convolution operation. The AI chip includes: an allocation calculation processing circuit and x calculation processing circuits. The AI chip is used to obtain the matrix size CI*CH of the input data. If the convolution kernel size in the n-layer convolution operation is a 3*3 convolution kernel, the allocation calculation processing circuit divides CI*CH into CI / x data blocks in the CI direction, and allocates the CI / x data blocks to the x calculation processing circuits in sequence. The x calculation processing circuits respectively perform the i-th convolution operation on the received 1 data block and the i-th convolution kernel to obtain the i-th convolution result, and send the results of the two edge columns of the i-th convolution result to the allocation processing circuit. The x calculation processing circuits perform a convolution operation on the i-th convolution result and the (i+1)-th convolution kernel to obtain the (i+1)-th convolution result, and the (i+1)-th convolution result The data is sent to the distribution calculation circuit. The distribution calculation processing circuit performs the i-th convolution operation on the (CI / x-1) combined data blocks and the i-th convolution kernel to obtain the i-th combination result, splices the i-th combination result with the results of the two edge columns of the i-th convolution result to obtain the (i+1)-th combined data block, performs a convolution operation on the (i+1)-th combined data block and the (i+1)-th convolution kernel to obtain the (i+1)-th combination result, inserts the (i+1)-th combination result between the edge columns of the (i+1)-th convolution result to obtain the (i+1)-th layer convolution result, and the AI chip performs the remaining convolution layer operation based on the (i+1)-th layer convolution result to obtain the n-th layer convolution operation result; the combined data block is a 4*CI matrix composed of 4 columns of data between 2 adjacent data blocks; n is the total number of convolution layers corresponding to the AI chip, i is the layer number of the convolution layer, CI is the column value of the matrix, and CH is the row value of the matrix.
5. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store one or more programs and is configured to be executed by the processor, wherein the programs include instructions for executing the steps of the method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Traffic signal lamp identification method and device based on artificial intelligence, equipment and medium
CN111738212A