Violation Detection Method, Device and Storage Medium Based on Fast Recurrent Network
Through the violation detection method based on the fast circular network, the problem of low recognition accuracy in harsh environments in the existing technology is solved, and a higher violation detection accuracy and recall rate is achieved.
Patent Information
- Application Number
- CN202011562635.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-12-25
AI Technical Summary
The existing violation detection technology has a low recognition accuracy rate in specific scenarios such as rainy days, nighttime, signs with obstructions or signs with blurred signs.
The violation detection method based on the fast cyclic network is adopted, and the violation feature maps in the image are extracted through the feature extraction network, convolutional operations and regional proposal network operations are performed, and the violation coordinates of the violation are obtained to realize the violation detection of the image.
It improves the feature expression ability of violation detection in harsh environments, and enhances the recognition accuracy and recall rate.
Smart Images

Figure CN112613412B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of artificial intelligence detection, and is applied to intelligent transportation. Specifically, it relates to a method, device, and storage medium for violation detection based on a fast recurrent network. Background Art
[0002] Violation detection is a necessary means for traffic order. The bus lane is an independent right-of-way lane specifically set for buses and is a very important infrastructure in urban traffic. Therefore, the detection of bus lanes has also become an indispensable part of the intelligent transportation system.
[0003] Existing violation detections have high requirements for application scenarios (such as weather). In some specific application scenarios, such as rainy days, nights, signs being blocked, signs being blurred, etc., the recognition accuracy is relatively low. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, and storage medium for violation detection based on a fast recurrent network, which can improve the accuracy in specific application scenarios.
[0005] In a first aspect, the embodiments of the present application provide a method for violation detection based on a fast recurrent network, which is applied to an electronic device. The method includes:
[0006] Extracting a first feature map representing violations in the obtained original image through a first network model;
[0007] Performing a convolution operation on the first feature map to obtain a second feature map, where the second feature map is a feature map matrix including all channels;
[0008] Performing a region proposal network operation on the second feature map to obtain an output result, performing a mapping process on the output result to obtain the coordinates of the violation features on the output result in the coordinates of the original image, and implementing the violation detection of the original image based on the coordinates of the original image.
[0009] The technical solution provided by the present application provides a device for violation detection based on a fast recurrent network, which is applied to an electronic device. The device includes:
[0010] An extraction unit, configured to extract a first feature map of violations in the image from the obtained original image through a feature extraction network;
[0011] An operation unit, configured to perform convolution-related operation on the first feature map to obtain a second feature map; performing a region proposal network-related operation on the second feature map to obtain an output result; where the second feature map is a feature map matrix including all channels;
[0012] An identification unit is configured to perform a mapping process on the output result to obtain the coordinates of the violation features on the output result in the coordinates of the original image, and implement the violation detection of the original image based on the coordinates of the original image.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, a communication interface, and one or more programs. Among them, the above one or more programs are stored in the above memory and are configured to be executed by the above processor. The above programs include instructions for performing the steps in the first aspect of the embodiment of the present application.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. Among them, the above computer-readable storage medium stores a computer program for electronic data exchange. Among them, the above computer program enables a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application.
[0015] In a fifth aspect, an embodiment of the present application provides a computer program product. Among them, the above computer program product includes a non-transitory computer-readable storage medium storing a computer program. The above computer program is operable to enable a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application. The computer program product can be a software installation package.
[0016] Implementing the embodiments of the present application has the following beneficial effects:
[0017] It can be seen that the technical solution provided by the present application uses an extraction network to better extract the original image to obtain a first feature map with better quality. Then, based on the first feature map with better quality, the first feature map can better represent the pixel changes in a harsh environment. The first feature map is subjected to convolution-related operations to obtain a second feature map. The first feature map with better quality greatly enhances the feature expression ability of the feature map in difficult scenarios such as rainy days, nights, blocked signs, and blurred signs, thereby improving the feature expression ability of the second feature map. Therefore, it can improve the overall accuracy and recall rate of violations. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0020] Figure 2 It is a schematic flowchart of a traffic violation detection method based on a fast recurrent network provided by an embodiment of the present application;
[0021] Figure 3 It is a schematic flowchart of a lane traffic violation detection method based on a fast recurrent network provided by an embodiment of the present application;
[0022] Figure 4 It is a schematic structural diagram of the CBL operation provided by an embodiment of the present application;
[0023] Figure 5 It is a schematic structural diagram of a traffic violation detection device based on a fast recurrent network provided by an embodiment of the present application. Detailed implementation manners
[0024] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0025] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0026] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0027] The technical solution provided by the embodiments of the present application is executed based on the hardware structure of an electronic device. The electronic device may be a portable electronic device that further includes other functions such as a personal digital assistant and / or a music player function, such as a mobile phone, a tablet computer, a wearable electronic device with wireless communication function (such as a smart watch), etc. Of course, it may also include a fixed electronic device, such as a traffic camera, a surveillance camera, and so on. Exemplary embodiments of the portable electronic device and the fixed electronic device include, but are not limited to, a portable electronic device or a fixed electronic device equipped with an IOS system, an Android system, a Microsoft system, or other operating systems (such as an embedded operating system). The above portable electronic device may also be other portable electronic devices, such as a laptop computer. It should also be understood that in some other embodiments, the above electronic device may not be a portable electronic device, but a desktop computer or a fixed camera.
[0028] An electronic device is a device deployed in an indoor environment or an outdoor environment for receiving and transmitting signals. For example, the signal transceiver device may be an evolved Node B (eNB), a radio network controller (RNC), a Node B (NB), a Base Station Controller (BSC), a Base Transceiver Station (BTS), a home base station (for example, Home evolved Node B, or Home Node B, HNB), an access controller (AC), a WIFI access point (Access Point, AP), etc.
[0029] The software and hardware operating environment of the technical solution disclosed in the present application is introduced as follows.
[0030] Exemplarily, Figure 1The structural schematic diagram of the electronic device 100 is shown. The electronic device may specifically be an intelligent camera. Of course, the above-mentioned electronic device may also be a wearable device, specifically, wearable portable devices such as smart watches and smart bracelets. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a compass 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0031] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent components or integrated in one or more processors. In some embodiments, the electronic device 100 may also include one or more processors 110. Among them, the controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions. In some other embodiments, a memory may also be provided in the processor 110 for storing instructions and data. Exemplarily, the memory in the processor 110 may be a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. In this way, repeated access is avoided, the waiting time of the processor 110 is reduced, and thus the efficiency of the electronic device 100 in processing data or executing instructions is improved.
[0032] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. Among them, the USB interface 130 is an interface that conforms to the USB standard specification, and specifically may be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 101, and can also be used to transfer data between the electronic device 101 and peripheral devices. The USB interface 130 can also be used to connect headphones to play audio through the headphones.
[0033] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is only for illustrative purposes and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0034] See Figure 2 , Figure 2 The present application provides a violation detection method based on a fast recurrent network. This method can be executed by an electronic device as shown in Figure 1 shown, and this method is as shown in Figure 2 shown, including:
[0035] Step S201: Extract a first feature map of violations in the obtained original image through a feature extraction network;
[0036] The above violations may specifically be dedicated lane violations or seat belt violations.
[0037] In an alternative embodiment, the above feature extraction network may be: a SpineNet network or a VoVNet network.
[0038] The above step S201 may specifically include:
[0039] Perform a resizing operation to obtain an image of a preset size, and perform a feature extraction operation on the image of the preset size to obtain a first feature map of the violation in the image.
[0040] The above resizing operation can be implemented through the Resize function in C++. The above preset size can be 1000*600. Of course, in actual applications, it can also be other sizes.
[0041] The violation is a dedicated lane violation (it can be a violation of driving in a dedicated lane type, such as a bus lane or an emergency lane, or it can be a violation of occupying a dedicated lane type, such as stopping at a specific lane position (emergency lane)). The specific steps for extracting the violation features in the image can include: performing a convolution operation on the image of the preset size through a 1*1 convolutional layer to obtain a convolution result, and performing an upsampling (Upsample) function operation of the SpineNet network on the convolution result and then performing another 1*1 convolutional layer to obtain a first feature map of the dedicated lane.
[0042] If the violation is a seat belt, the specific steps for extracting the violation features in the image can include: performing a convolution operation on the image of the preset size through three 3*3 convolutional layers to obtain a convolution result, and performing the OSA (One-Shot Aggregation) of the VoVNet network four times on the convolution result to obtain a first feature map of the seat belt features.
[0043] Step S202: Perform a convolution-related operation on the first feature map to obtain a second feature map;
[0044] The specific implementation method of the above step S202 can include:
[0045] Perform a convolution operation on the first feature map and then perform an activation function operation to obtain a second feature map containing all channels. The above second feature map is a feature map matrix containing all channels.
[0046] For example, if the first feature map can be a feature map of 1000*600, after performing step S202, a feature map matrix containing all channels can be obtained. For example, a second feature map of 60*40*512 can be obtained.
[0047] If the violation is a dedicated lane violation, the above convolution operation can be a 3*3 convolution operation, and the above activation function can specifically be: the relu activation function.
[0048] If the violation is a seat belt violation, the above convolution-related operation can be a CBL operation. Refer to Figure 4 The CBL operation specifically can include: convolution + batch normalization + Leaky Relu activation function.
[0049] Step S203: Perform region proposal network related operations on the second feature map to obtain an output result, perform mapping processing on the output result to obtain the coordinates of the violation features on the output result in the coordinates of the original image, and implement violation detection of the original image based on the coordinates of the original image.
[0050] The above-mentioned mapping processing of the output result to obtain the coordinates of the violation features on the output result in the coordinates of the original image may specifically include:
[0051] After determining the scaling factor of the reference box, enlarge the reference box by the reciprocal of the corresponding scaling factor to obtain an enlarged reference box;
[0052] Map the enlarged reference box to the original image, extract the coordinates of the violation features of the reference box corresponding to the original image; determine that this coordinate position is the coordinate position of the violation features in the original image.
[0053] The implementation method of the above-mentioned step S203 may specifically include:
[0054] Perform fast region proposal network operations on the second feature map to obtain an output result. Specifically:
[0055] Perform CBL operations on the second feature map to obtain a CBL result, perform CLS function operations on the CBL result to obtain a CLS result, perform reg function operations on the CBL result to obtain a reg result, perform CBL operations, CS operations, and convolution operations on the CLS result and the reg result in sequence to obtain two convolution result matrices, merge the two convolution result matrices to obtain a merged matrix, and perform deduplication on the merged matrix to obtain an output result.
[0056] The above-mentioned CBL operations include: convolution operations, batch normalization BN, and leaky ReLU activation function operations; the above-mentioned CS operations include: 1*1 convolution operations and sigmoid activation function operations.
[0057] For example: Assume that the size of the second feature map (i.e., the Input matrix) is 32*32*512. After the CBL operation, the size of the matrix remains 32*32*512. Then, after operations such as CBL, CS, and Conv in the cls and reg branches, two matrices (i.e., the classification matrix and the coordinate matrix) are obtained, with sizes of 32*32*(3*2), 32*32*6 (where 3 represents the number of anchor boxes and 2 represents the binary classification of seat belt and non-seat belt), and 32*32*(3*4) = 32*32*12 (where 3 represents the number of anchor boxes and 4 represents the coordinates of each anchor box (i.e., the center point coordinates x, y of the anchor box and the width and height w, h of the anchor box)). Finally, the obtained coordinate matrix 32*32*12 (i.e., all predicted bounding boxes) is subjected to NMS (i.e., sorting the predicted bounding boxes in descending order according to the classification probability, retaining the predicted bounding box with the highest probability, and filtering out the duplicate boxes using IOU) to remove duplicates and obtain the final predicted result Output (one output result).
[0058] The implementation of the illegal behavior detection of the original image based on the output result specifically may include:
[0059] If the violation is a seat belt, map the seat belt coordinates predicted on multiple output results to the coordinates on the original image Image, and determine whether the position of the coordinates belongs to the specified area of the seat belt, so as to realize the detection of the driver's seat belt in the image.
[0060] If the violation is a dedicated lane, perform two-layer FCR (i.e., fully connected + relu activation function) and two different branches of FC (i.e., fully connected) and FCS (fully connected + softmax activation function) on the output result (i.e., the candidate bounding box Cbox) to perform regression (reg) and classification (cls) of the detection bounding box respectively to obtain the dedicated lane coordinates predicted on the feature map; map the dedicated lane coordinates predicted on the feature map to the coordinates on the original image Image, and determine whether the area of the coordinates is in the dedicated lane area, so as to realize the detection of the dedicated lane in the image.
[0061] In an alternative solution, before step S203, the above method may further include:
[0062] Perform multiple convolution-related operations and upsampling operations on the first feature map to obtain multiple feature maps with different sizes; these multiple feature maps with different sizes (such as Figure 3 FMap2 and FMap3 in) are feature maps containing all channels with different sizes from the second feature map;
[0063] Perform region proposal network related operations on multiple feature maps of different sizes respectively to obtain multiple output results, and implement violation detection of the original image based on the multiple output results. The above-mentioned operations of performing Pyramid-RPN on multiple feature maps respectively obtain candidate boxes Cbox.
[0064] Refer to Figure 3 , perform a CBL operation on the feature map matrix to obtain a first CR result (i.e., a feature map FMap1 of one scale), perform an Upsample operation on the first CR result to obtain a second CR result (i.e., a feature map FMAP2 of one scale), and perform an Upsample operation on the second CR result to obtain a third CR result (i.e., a feature map FMAP3 of one scale). The above-mentioned FMap1, 2, and 3 are feature maps of multiple different scales. Perform PCR operations on the feature maps of multiple scales respectively to obtain the results of multiple PCR operations, and perform merge (merge sort algorithm) on the results of multiple PCR operations to obtain candidate boxes Cbox.
[0065] The above-mentioned PCR operation can be: convolution operations of multiple convolutional layers, and the multiple convolutional layers can be convolutional layers with a pyramid structure (the number of layers can be configured, such as 3-layer pyramid, 4-layer pyramid, etc., and the size of the convolutional kernel can also be configured, as long as it is ensured to be from large to small, such as 7*7->5*5->3*3, 9*9->7*7->5*5->3*3). For example: 7*7 convolution + relu activation function -> 5*5 convolution + relu activation function -> 3*3 convolution + relu activation function.
[0066] Here, FMap1 among the three feature maps FMap1, FMap2, and FMap3 is used to elaborate on the specific execution process of the Pyramid - RPN generation network. Assume that the size of the feature map matrix FMap1 is 60 * 40 * 512. Then, after the PCR operation, the size of the matrix remains 60 * 40 * 512. After the CS and Conv operations in the cls and reg branches respectively, the sizes of the obtained matrices are 60 * 40 * (9 * 2) = 60 * 40 * 18 (9 represents the number of anchor boxes, and 2 represents binary classification of foreground and background), 60 * 40 * (9 * 4) = 60 * 40 * 36 (9 represents the number of anchor boxes, and 4 represents the coordinates of each anchor box (i.e., the center point coordinates x, y of the anchor box and the width and height w, h of the anchor box)). Next, the obtained coordinate matrix 60 * 40 * 36 (i.e., all candidate boxes) is subjected to the NMS operation to remove overlapping candidate boxes (since the NMS operation requires descending sorting of the probabilities of candidate boxes and screening using IOU, the classification matrix 60 * 40 * 18 is needed) to obtain all the screened candidate boxes. Finally, all the obtained candidate boxes are Cut on the input feature map FMap1 to obtain the feature map matrix Output of all the final candidate boxes (i.e., Figure 1 CBox1 in
[0067] Similarly, the other two groups of candidate box feature maps CBox2 and CBox3 can be obtained. Then, these three groups of candidate boxes (assuming that 50, 100, and 150 candidate boxes are respectively retained after NMS screening for these three groups, and the retained number can be configured) are merged to obtain a total of 300 candidate boxes CBox.
[0068] Specific example of the Cut operation: Assume that the size of the FMap1 matrix is 60 * 40 * 512, and there are 50 candidate boxes screened by NMS. Assume that the coordinates of one of the candidate boxes are (9, 15, 20, 30)). Then, the Cut operation means cutting out a candidate box with a width and height of (20, 30) at the position (9, 15) on the feature map FMap1 (i.e., the candidate box feature map matrix with a size of 20 * 30 * 512). Similarly, finally, 50 different - sized candidate box feature map matrices can be obtained.
[0069] In an alternative solution, the technical solution of the present application can be configured to set a separate AI chip to perform convolution operations, and the convolution operations can be multi-layer convolution operations (or multi-stage convolution operations). The AI chip includes: an allocation calculation processing circuit and x calculation processing circuits. The AI chip obtains the matrix size CI*CH of the input data. If the convolution kernel size in the n-layer convolution operation is a 3*3 convolution kernel, the allocation calculation processing circuit divides CI*CH into CI / x data blocks in the CI direction (assuming CI is an integer multiple of x), and sequentially allocates the CI / x data blocks to the x calculation processing circuits. The x calculation processing circuits respectively perform the i-th layer convolution operation on the 1 data block received and allocated with the i-th layer convolution kernel to obtain the i-th convolution result (that is, combining the x result matrices (CI / x - 2)*(CH - 2) of the x calculation processing circuits in sequence to obtain the i-th convolution result), and send the results of the 2 columns at the edge of the i-th convolution result (the 2 columns where the results of adjacent columns are calculated by different calculation processing circuits are determined as the edge columns) to the allocation processing circuit. The x calculation processing circuits perform the convolution operation on the i-th convolution result and the (i + 1)-th layer convolution kernel to obtain the (i + 1)-th convolution result, and send the (i + 1)-th convolution result to the allocation calculation circuit. The allocation calculation processing circuit performs the i-th layer convolution operation on the (CI / x - 1) combined data blocks and the i-th layer convolution kernel to obtain the i-th combined result, splices the i-th combined result with the results of the 2 columns at the edge of the i-th convolution result (inserting the i-th combined result into the middle of the 2 edge columns according to the mathematical rules of the convolution operation) to obtain the (i + 1)-th combined data block, performs the convolution operation on the (i + 1)-th combined data block and the (i + 1)-th layer convolution kernel to obtain the (i + 1)-th combined result, and inserts the (i + 1)-th combined result between the edge columns of the (i + 1)-th convolution result (the results of adjacent columns are calculated by different calculation processing circuits) to obtain the (i + 1)-th layer convolution result. The AI chip performs the operations of the remaining convolution layers (convolution kernels after the (i + 1)-th layer) based on the (i + 1)-th layer convolution result to obtain the n-th layer convolution operation result. The above combined data block can be a 4*CI matrix composed of 4 columns of data between two adjacent data blocks. For example, a 4*CH matrix composed of the last 2 columns of the first data block (the data block allocated to the first calculation processing circuit) and the first 2 columns of the second data block (the data block allocated to the second calculation processing circuit).
[0070] For the operations of the above remaining convolution layers, reference can also be made to the calculations of the i-th layer and the (i + 1)-th layer, where i is an integer greater than or equal to 1 and less than or equal to n. The above n is the total number of convolution layers of the AI model, and i is the layer number of the convolution layer. The CI is the column value of the matrix, and the CH is the row value of the matrix.
[0071] Setting a separate AI chip to perform convolution operations can improve the speed of convolution operations and reduce the IO overhead. Therefore, it has the advantages of cost savings and power consumption reduction.
[0072] It can be understood that, in order to implement the above functions, the electronic device includes the corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the manner of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the present application.
[0073] In this embodiment, the electronic device can be divided into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0074] In the case of dividing each functional module corresponding to each function, Figure 5 The illegal detection device based on the fast recurrent network is shown, which is applied to an electronic device. The device includes:
[0075] An extraction unit 501, configured to extract a first feature map of the illegal situation in the acquired original image through a feature extraction network;
[0076] An operation unit 502, configured to perform convolution-related operation operations on the first feature map to obtain a second feature map; and perform region proposal network-related operations on the second feature map to obtain an output result;
[0077] An identification unit 503, configured to implement the illegal detection of the original image according to the output result.
[0078] Among them, the extraction unit 501 can be used to support the electronic device to execute the above step 201 and / or other processes of the technology described in this article. The operation unit 502 can be used to support the electronic device to execute the above step 202 and / or other processes of the technology described in this article.
[0079] The identification unit 503 can be used to support the electronic device to execute the above step 203, etc., and / or other processes of the technology described in this article.
[0080] It should be noted that all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be repeated here.
[0081] The electronic device provided in this embodiment is used to execute the above-mentioned position determination method, so the same effect as the above implementation method can be achieved.
[0082] In the case of adopting an integrated unit, the electronic device may include a processing module, a storage module, and a communication module. Among them, the processing module can be used to control and manage the operations of the electronic device. For example, it can be used to support the electronic device to execute the steps performed by the above-mentioned extraction unit 501, operation unit 502, and recognition unit 503. The storage module can be used to support the electronic device to execute storing program codes and data, etc. The communication module can be used to support the communication between the electronic device and other devices.
[0083] Among them, the processing module can be a processor or a controller. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of this application. The processor can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, and so on. The storage module can be a memory. The communication module can specifically be a device for interacting with other electronic devices, such as a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, etc.
[0084] In one embodiment, when the processing module is a processor and the storage module is a memory, the electronic device involved in this embodiment can be a device with Figure 1 the structure shown.
[0085] This application embodiment also provides an electronic device, including a processor and a memory. The memory is used to store one or more programs and is configured to be executed by the processor. The programs include instructions for performing the following steps, and the following steps may specifically include:
[0086] Extract the first feature map of violations in the acquired original image through a feature extraction network;
[0087] Perform a convolution-related operation on the first feature map to obtain a second feature map;
[0088] Perform region proposal network-related operations on the second feature map to obtain an output result, and implement the violation detection of the original image based on the output result.
[0089] The instructions executed by the above program may also include the refinement solutions of steps S201, S202, and S203 as shown in Figure 2 and of course the optional solutions in the embodiments as shown in Figure 2 . Details are not described here again.
[0090] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute some or all of the steps of any one of the methods described in the foregoing method embodiments.
[0091] An embodiment of the present application also provides a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute some or all of the steps of any one of the methods described in the foregoing method embodiments. The computer program product may be a software installation package.
[0092] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, some steps may be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0093] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0094] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0095] The units described as separate components above may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit exists physically alone, or two or more units are integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0097] When the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of this application. The aforementioned memory includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs, etc., all kinds of media that can store program codes.
[0098] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory. The memory can include: flash drives, read-only memories (abbreviation: ROM, English: Read-Only Memory), random access memories (abbreviation: RAM, English: Random Access Memory), magnetic disks, or optical discs, etc.
[0099] The above has introduced the embodiments of this application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A violation detection method based on a fast recurrent network, characterized in that, applied to an electronic device, the method includes: extracting a first feature map representing violations in the original image from the obtained original image through a first network model; performing a convolution-related operation on the first feature map to obtain a second feature map, where the second feature map is a feature map matrix containing all channels; performing a fast region proposal network operation on the second feature map to obtain an output result, specifically including: performing a CBL operation on the second feature map to obtain a CBL result; performing a CLS function operation and a reg function operation on the CBL result respectively. Obtaining a CLS result by performing the CLS function operation and obtaining a reg result by performing the reg function operation; performing a CBL operation, a CS operation, and a convolution operation on the CLS result and the reg result in sequence respectively to obtain two convolution result matrices; merging the two convolution result matrices to obtain a merged matrix, performing deduplication on the merged matrix to obtain an output result, performing a mapping process on the output result to obtain the coordinates of the violation features on the output result in the coordinates of the original image, and implementing violation detection of the original image based on the coordinates of the original image.
2. The method according to claim 1, characterized in that, extracting a first feature map of violations in the image from the obtained original image through a feature extraction network specifically includes: adjusting the original image according to a preset size adjustment rule to obtain a preset size image; performing a feature extraction operation on the preset size image to obtain a first feature map of violations in the image.
3. The method according to claim 1, characterized in that, performing a convolution-related operation on the first feature map to obtain a second feature map specifically includes: performing an activation function operation on the first feature map after convolution operation to obtain a feature map matrix containing all channels.
4. The method according to claim 1, characterized in that, the CBL operation includes: convolution operation, batch normalization BN, and leaky ReLU activation function operation; the CS operation includes: 1*1 convolution operation and sigmoid activation function operation.
5. The method according to claim 1, characterized in that, performing a mapping process on the output result to obtain the coordinates of the violation features on the output result in the coordinates of the original image specifically includes: after determining the scaling factor of the output result, magnifying the output result by the reciprocal of the corresponding scaling factor to obtain a magnified output result; mapping the magnified output result frame to the original image, and extracting the coordinates of the violation features of the output result corresponding to the original image.
6. The method according to claim 1, characterized in that, before performing a region proposal network operation on the second feature map to obtain an output result, the method further includes: performing multiple convolution operations and upsampling operations on the second feature map to obtain multiple feature maps of different sizes; the multiple feature maps of different sizes are feature maps containing all channels with different sizes from the second feature map; Perform operations related to the region proposal network on feature maps of multiple different sizes respectively to obtain multiple output results, perform mapping processing on the output results to obtain the coordinates of the violation features on the output results in the coordinates of the original image, and implement the violation detection of the original image based on the coordinates of the original image.
7. A violation detection device based on a fast recurrent network, characterized in that it is applied to an electronic device, and the device includes: An extraction unit, configured to extract a first feature map of violations in the image from the acquired original image through a feature extraction network; An operation unit, configured to perform convolution-related operation operations on the first feature map to obtain a second feature map; wherein the second feature map is a feature map matrix including all channels; performing a fast region proposal network operation on the second feature map to obtain an output result, specifically including: Performing CBL operation on the second feature map to obtain a CBL result; performing CLS function operation and row reg function operation on the CBL result respectively, obtaining a CLS result by performing the CLS function operation, and obtaining a reg result by performing the reg function operation; performing CBL operation, CS operation and convolution operation on the CLS result and the reg result in sequence respectively to obtain two convolution result matrices; merging the two convolution result matrices to obtain a merged matrix, and performing duplicate removal on the merged matrix to obtain an output result; An identification unit, configured to perform mapping processing on the output result to obtain the coordinates of the violation features on the output result in the coordinates of the original image, and implement the violation detection of the original image based on the coordinates of the original image.
8. An electronic device, characterized in that it includes a processor and a memory, the memory is used to store one or more programs, and is configured to be executed by the processor, and the programs include instructions for performing the steps in the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Head and shoulder region detection method and device
CN108805016A
Method and device for identifying traffic light signal, readable medium and electronic equipment
CN108830199A