Object Detection Method and Related Devices
By extracting and fusion the RGB images and IR images of the monitoring area to generate a fusion feature map, the problem of low target detection accuracy in the prior art is solved, and more efficient and accurate target detection is achieved, reducing supervision costs.
Patent Information
- Application Number
- CN202011642072.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-12-31
AI Technical Summary
In the prior art, when detecting early warning areas in the image, there is a problem of low detection accuracy, which leads to actual travel vendors but not detection, or no travel vendors detecting travel vendors, which increases the on-site confirmation and processing costs of supervisors.
By using feature extraction and fusion technology of RGB images and IR images in the monitoring area, the RGB feature map and IR feature map are obtained, and fused to generate a fusion feature map to improve the accuracy of object detection.
Through the fusion feature map for object detection, the detection accuracy can be improved, false detection can be reduced, supervision efficiency can be improved, and manual supervision costs can be reduced.
Smart Images

Figure CN114694000B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to an object detection method and related devices. Background Art
[0002] In today's society, efforts are being made to build civilized cities. However, unregulated street vendors still occasionally appear in people's sight. Currently, the main solution for managing street vendors is to collect images through cameras, detect warning areas in the images, and then have supervisors go to the scene to confirm, drive them away, impose fines, etc. However, this solution has a high cost of manual supervision. In addition, in the prior art, when detecting warning areas in images, there is a problem of low detection accuracy, resulting in the situation that there are actually street vendors but they are not detected; or there are actually no street vendors but it is detected that there are street vendors, causing supervisors to find no street vendors when they go to the scene to confirm. Summary of the Invention
[0003] Embodiments of this application disclose an object detection method and related devices, which are beneficial to improving the accuracy of object detection.
[0004] A first aspect of an embodiment of this application discloses an object detection method, the method includes: extracting features from an RGB image to obtain an RGB feature map, and extracting features from an IR image to obtain an IR feature map, where the RGB image and the IR image are obtained by collecting images of the same image acquisition area; fusing the RGB feature map and the IR feature map to obtain a fused feature map; and detecting whether there is an object in the image acquisition area according to the fused feature map.
[0005] In an embodiment of this application, a monitoring area is used as the image acquisition area, RGB images and IR images of the image acquisition area are collected, an RGB feature map is obtained by extracting features from the RGB image, and an IR feature map is obtained by extracting features from the IR image; then the RGB feature map and the IR feature map are fused to obtain a fused feature map; and then it is detected whether there is an object in the monitoring area according to the fused feature map. Since the RGB image has characteristics such as high resolution and rich color texture information, and the IR image has the characteristic of being insensitive to light, the fused feature map obtained by fusing the RGB feature map and the IR feature map has the characteristics of both the RGB image and the IR image at the same time; the embodiment of this application uses the fused feature map for object detection, which can improve the accuracy of object detection and obtain more accurate detection results.
[0006] In a possible implementation, the step of fusing the RGB feature map and the IR feature map to obtain a fused feature map includes: mapping the first pixel value of each point of the IR feature map to a value between 0 and 1 to obtain the second pixel value of each point of the IR feature map; multiplying the second pixel value of each point of the IR feature map by the third pixel value of the corresponding point of the RGB feature map to obtain a plurality of fourth pixel values; and combining the plurality of fourth pixel values to obtain the fused feature map.
[0007] In this implementation, the first pixel value of each point of the IR feature map is mapped to a value between 0 and 1 and used as the weight of the corresponding point in the RGB feature map. Since the IR image is not affected by factors such as illumination, the features of the IR feature map are concentrated on the target, and unimportant factors such as the background in the image can be ignored. After each point of the RGB feature map is multiplied by the weight, that is, after each point of the RGB feature map is multiplied by the second pixel value of the IR feature map mapped to a value between 0 and 1, unimportant pixel points are filtered out, highlighting the target. Thus, the obtained fused feature map is a feature map that filters out unimportant pixel points and highlights the target. Using this fused feature map for object detection can improve the accuracy of object detection. In addition, after mapping the first pixel value of each point of the IR feature map to a value between 0 and 1 and then using it as the weight of the corresponding point in the RGB feature map, the computational amount can be reduced and the feature fusion speed can be accelerated.
[0008] In a possible implementation, the step of extracting features from the RGB image to obtain an RGB feature map and extracting features from the IR image to obtain an IR feature map includes: annotating a prediction box in the RGB image and annotating the prediction box in the IR image, where the prediction box is used to frame a candidate object; extracting features from the region of the RGB image framed by the prediction box to obtain the RGB feature map, and extracting features from the region of the IR image framed by the prediction box to obtain the IR feature map.
[0009] In this implementation manner, the RGB image and the IR image are images that include the entire monitoring area, and the target appears at a certain position in the monitoring area. Prediction boxes are respectively marked in the RGB image and the IR image to frame the candidate objects with the prediction boxes, and the framed candidate objects are the candidate objects that may be the target; then features are respectively extracted within the prediction boxes of the RGB image and the IR image to obtain an RGB feature map and an IR feature map; since the RGB feature map extracted from the RGB image is the feature of the regional image including the candidate object in the RGB image, and the IR feature map extracted from the IR image is the feature of the regional image including the candidate object in the IR image, the features in the RGB feature map and the IR feature map are concentrated on the candidate object, which is beneficial to improving the accuracy of target detection; and compared with directly extracting the RGB feature map from the entire RGB image and directly extracting the IR feature map from the entire IR image, the total number of features of the RGB feature map and the IR feature map extracted by using this implementation manner is less, which can reduce the amount of computation and is thus beneficial to improving the speed of target detection.
[0010] In a possible implementation manner, the detecting whether there is a target in the image acquisition area according to the fused feature map includes: performing class prediction on each point of the fused feature map to obtain the class and classification confidence of each point of the fused feature map, where the class includes a target class and a non-target class, the target class is the class to which the target belongs, and the non-target class is the class to which the non-target belongs; counting the number of points predicted as the target class in the fused feature map, and counting the number of points predicted as the non-target class in the fused feature map; calculating a first average classification confidence according to the number of points predicted as the target class and the classification confidence of each point predicted as the target class, and calculating a second average classification confidence according to the number of points predicted as the non-target class and the classification confidence of each point predicted as the non-target class; if the number of points predicted as the target class is greater than the number of points predicted as the non-target class, and the first average classification confidence is greater than the second average classification confidence, then determine that the candidate object is a target; otherwise, determine that the candidate object is a non-target.
[0011] In this implementation manner, category prediction is performed on each point of the fused feature map to obtain the category and classification confidence of each point of the fused feature map. Among them, the categories include the target category and the non-target category. The target category is the category to which the target belongs, and the non-target category is the category to which the non-target belongs. Then, the first average classification confidence of the points classified as the target category is calculated, and the second average classification confidence of the points classified as the non-target category is calculated. If the first average classification confidence is greater than the second average classification confidence, it is determined that the candidate object is the target. If the number of points predicted to be the target category is greater than the number of points predicted to be the non-target category, and the first average classification confidence is not greater than the second average classification confidence, it is determined that the candidate object is a non-target. Otherwise, it is determined that the candidate object is a non-target. Since it is only determined that the candidate object is the target when the proportion of the target in the fused feature map is greater than the proportion of the non-target and the average classification confidence of the target is greater than the average classification confidence of the non-target, it is beneficial to improve the accuracy of target detection.
[0012] In a possible implementation manner, if there is a target in the image acquisition area, the method further includes: extracting a face image from the RGB image to obtain a target face image; comparing the target face image with multiple pre-stored template face images; if the comparison is successful, obtaining the identity information corresponding to the target template face image, and performing a preset operation according to the identity information, where the target template face image is the template face image that is successfully compared with the target face image.
[0013] In this implementation manner, after determining that there is a target in the image acquisition area, that is, after determining that there is a target in the monitoring area, a target face image is extracted from the RGB image, and the extracted target face image is compared with multiple pre-stored template face images. Then, the identity information corresponding to the target template face image that is successfully compared with the target face image is obtained, and a preset operation is performed according to this identity information, so as to automatically monitor the target, without the need for the supervisor to go to the monitoring area for on-site management, reducing the monitoring cost.
[0014] In a possible implementation manner, if the comparison fails, the method further includes: sending a prompt message to the handheld terminal, where the prompt message includes the RGB image, and the prompt message is used to indicate processing of the target. The handheld terminal is the terminal held by the supervisor.
[0015] In this implementation manner, if the target face image is not successfully compared with any of the multiple pre-stored template face images, automatic monitoring cannot be performed, and a prompt message including the RGB image is sent to the terminal held by the supervisor, instructing the supervisor to process the target, so as to ensure that the target can be monitored.
[0016] In a second aspect of the embodiments of the present application, a target detection device is disclosed. The device includes: an extraction unit configured to perform feature extraction on an RGB image to obtain an RGB feature map and perform feature extraction on an IR image to obtain an IR feature map, where the RGB image and the IR image are obtained by performing image acquisition on the same image acquisition area; a fusion unit configured to fuse the RGB feature map and the IR feature map to obtain a fused feature map; and a detection unit configured to detect whether a target exists in the image acquisition area according to the fused feature map.
[0017] In a possible implementation manner, the fusion unit is specifically configured to: map the first pixel value of each point of the IR feature map to between 0 and 1 to obtain the second pixel value of each point of the IR feature map; multiply the second pixel value of each point of the IR feature map by the third pixel value of the corresponding position point of the RGB feature map to obtain a plurality of fourth pixel values; and combine the plurality of fourth pixel values to obtain the fused feature map.
[0018] In a possible implementation manner, in terms of performing feature extraction on the RGB image to obtain the RGB feature map and performing feature extraction on the IR image to obtain the IR feature map, the extraction unit is specifically configured to: mark a prediction box in the RGB image and mark the prediction box in the IR image, where the prediction box is used to frame a candidate object; perform feature extraction on the area of the RGB image framed by the prediction box to obtain the RGB feature map, and perform feature extraction on the area of the IR image framed by the prediction box to obtain the IR feature map.
[0019] In a possible implementation manner, in terms of detecting whether a target exists in the image acquisition area according to the fused feature map, the detection unit is specifically configured to: perform class prediction on each point of the fused feature map to obtain the class and classification confidence of each point of the fused feature map, where the class includes a target class and a non-target class, the target class is the class to which the target belongs, and the non-target class is the class to which the non-target belongs; count the number of points predicted as the target class in the fused feature map and count the number of points predicted as the non-target class in the fused feature map; calculate a first average classification confidence according to the number of points predicted as the target class and the classification confidence of each point predicted as the target class, and calculate a second average classification confidence according to the number of points predicted as the non-target class and the classification confidence of each point predicted as the non-target class; if the number of points predicted as the target class is greater than the number of points predicted as the non-target class and the first average classification confidence is greater than the second average classification confidence, determine that the candidate object is a target; otherwise, determine that the candidate object is a non-target.
[0020] In a possible implementation, if there is a target in the image acquisition area, the detection unit is further configured to: extract a face image from the RGB image to obtain a target face image; compare the target face image with multiple pre-stored template face images; if the comparison is successful, obtain the identity information corresponding to the target template face image, and perform a preset operation according to the identity information, where the target template face image is the template face image that is successfully compared with the target face image.
[0021] In a possible implementation, if the comparison fails, the detection unit is further configured to: send a prompt message to the handheld terminal, where the prompt message includes the RGB image, and the prompt message is used to indicate processing of the target, and the handheld terminal is a terminal held by a supervisor.
[0022] A third aspect of the embodiments of the present application discloses an electronic device, including a processor, a memory, a communication interface, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the processor, and the programs include instructions for performing the steps in the method according to any one of the first aspects of the embodiments of the present application.
[0023] A fourth aspect of the embodiments of the present application discloses a chip, including: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the method according to any one of the first aspects of the embodiments of the present application.
[0024] A fifth aspect of the embodiments of the present application discloses a computer-readable storage medium, which stores a computer program for electronic data exchange, where the computer program enables a computer to execute the method according to any one of the first aspects of the embodiments of the present application.
[0025] A sixth aspect of the embodiments of the present application discloses a computer program product, where the computer program product enables a computer to execute the method according to any one of the first aspects of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0027] Figure 1 It is a schematic flowchart of a target detection method provided by the embodiments of the present application.
[0028] Figure 2 It is a schematic structural diagram of an object detection model provided by an embodiment of the present application.
[0029] Figure 3 It is a schematic structural diagram of a feature map fusion module provided by an embodiment of the present application.
[0030] Figure 4 It is a schematic structural diagram of an object detection device provided by an embodiment of the present application.
[0031] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0032] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0033] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an object detection method provided by an embodiment of the present application. The object detection method can be applied to an electronic device, and the object detection method includes but is not limited to the following steps.
[0034] Step 101: Extract features from the RGB image to obtain an RGB feature map, and extract features from the IR image to obtain an IR feature map, where the RGB image and the IR image are obtained by image acquisition of the same image acquisition area.
[0035] The embodiments of the present application can be applied to the "new governance" road monitoring scenario, such as monitoring the scenario of street vendors occupying road resources. For example, determining the location of street vendors, and at the same time intercepting the faces of street vendors and reporting them to relevant management agencies.
[0036] Among them, the RGB image and the IR image are aligned, that is, the pictures in the RGB image and the IR image are the same; the image acquisition area can be a monitoring area; the RGB image and the IR image can be obtained by synchronous acquisition of the same hardware, or by synchronous acquisition of different hardware. Specifically, the electronic device includes an IR-RGB binocular camera, and the electronic device acquires images of the monitoring area through the IR-RGB binocular camera to obtain the RGB image and the IR image; or, the electronic device includes an IR camera and an RGB camera, and the electronic device acquires images of the monitoring area through the IR camera to obtain the IR image, and the electronic device synchronously acquires images of the monitoring area through the RGB camera to obtain the RGB image; or, the RGB image and the IR image are acquired by other devices and sent to the electronic device for processing, and the other devices synchronously acquire the RGB image and the IR image.
[0037] In the embodiments of the present application, the RGB image and the IR image of the monitoring area are synchronously collected. The advantage of the IR image is that it can be used in scenarios such as at night, but its resolution is slightly lower than that of the RGB image and it cannot be used for face recognition. The RGB image contains richer color texture information and can be used to capture faces for ID matching, etc. If the RGB image and the IR image are combined for object detection, more accurate detection results can be obtained.
[0038] In a possible implementation manner, the feature extraction of the RGB image to obtain the RGB feature map and the feature extraction of the IR image to obtain the IR feature map include: annotating a prediction box in the RGB image and annotating the prediction box in the IR image, where the prediction box is used to frame a candidate object; performing feature extraction on the area of the RGB image framed by the prediction box to obtain the RGB feature map, and performing feature extraction on the area of the IR image framed by the prediction box to obtain the IR feature map.
[0039] Among them, the monitoring area is a relatively large area, and the target to be monitored may only occupy a very small area of the monitoring area. Therefore, the target also occupies a very small area of the captured image. Since there are too many background areas in the image, if feature extraction is performed on the entire captured image, the proportion of the effective features extracted in the total number of features is very small, so the effective features are not prominent and it is not conducive to improving the accuracy of object detection. Therefore, a prediction box can be annotated in the captured image, and the prediction box is used to select a candidate object. The candidate object may be the target to be monitored, and then features are extracted in the prediction box. Since there are fewer background areas in the prediction box, more effective features can be extracted from the prediction box, which is conducive to improving the accuracy of object detection.
[0040] Specifically, prediction boxes are annotated in both the RGB image and the IR image. The RGB feature map is extracted from the prediction box in the RGB image, and the IR feature map is extracted from the prediction box in the IR image. Among them, the image content or the image frame framed by the prediction box in the RGB image and the prediction box in the IR image is the same. In addition, the target can be a mobile vendor, and the mobile vendor includes a food truck, a booth, a person, etc.
[0041] In this implementation, the RGB image and the IR image are images that cover the entire monitoring area, and the target appears at a certain position in the monitoring area. Prediction boxes are respectively marked in the RGB image and the IR image to frame the candidate objects with the prediction boxes, and the framed candidate objects are the candidate objects that may be the target. Then, features are respectively extracted within the prediction boxes of the RGB image and the IR image to obtain an RGB feature map and an IR feature map. Since the RGB feature map extracted from the RGB image is the feature of the regional image including the candidate object in the RGB image, and the IR feature map extracted from the IR image is the feature of the regional image including the candidate object in the IR image, the features in the RGB feature map and the IR feature map are concentrated on the candidate object, which is beneficial to improving the accuracy of target detection. Moreover, compared with directly extracting the RGB feature map from the entire RGB image and directly extracting the IR feature map from the entire IR image, the total number of features of the RGB feature map and the IR feature map extracted by using this implementation is less, which can reduce the amount of computation and is thus beneficial to improving the speed of target detection.
[0042] Step 102: Fuse the RGB feature map and the IR feature map to obtain a fused feature map.
[0043] In a possible implementation, the fusing of the RGB feature map and the IR feature map to obtain a fused feature map includes: mapping the first pixel value of each point of the IR feature map to between 0 and 1 to obtain the second pixel value of each point of the IR feature map; multiplying the second pixel value of each point of the IR feature map by the third pixel value of the corresponding position point of the RGB feature map to obtain a plurality of fourth pixel values; and combining the plurality of fourth pixel values to obtain the fused feature map.
[0044] Among them, fusing the RGB feature map and the IR feature map means using the normalized IR feature map as a weight to multiply with the RGB feature map for fusion, thereby implementing an attention mechanism through multi-modal (RGB and IR are two modalities). Specifically, the first pixel value of each point of the IR feature map is restricted to between 0 and 1 and used as the weight of the corresponding point of the RGB feature map. Since the IR image itself is not affected by factors such as illumination, the features of the extracted feature map are concentrated on the pixels of food trucks, stalls, and people, ignoring unimportant factors such as the background. After multiplying the RGB feature map by this weight, unimportant pixel points are filtered out and the main body is highlighted, that is, the attention area of target detection.
[0045] Among them, the first pixel value of each point in the IR feature map is restricted to between 0 and 1, aiming to facilitate calculation and improve the operation speed; moreover, the pixel values of the original IR feature map may be very large or very small. After multiplying with the RGB feature map, the pixel values of the obtained fused feature image may be very large or very small, which may lead to the consequence of subsequent calculation overflow.
[0046] In this implementation, the first pixel value of each point in the IR feature map is mapped to between 0 and 1 and used as the weight of the corresponding point in the RGB feature map. Since the IR image is not affected by factors such as illumination, the features of the IR feature map are concentrated on the target, and unimportant factors such as the background in the image can be ignored; after each point in the RGB feature map is multiplied by the weight, that is, after each point in the RGB feature map is multiplied by the second pixel value of the IR feature map mapped to between 0 and 1, the unimportant pixel points are filtered out, highlighting the target. Thus, the obtained fused feature map is a feature map that filters out unimportant pixel points and highlights the target. Using this fused feature map for object detection can improve the accuracy of object detection. In addition, after mapping the first pixel value of each point in the IR feature map to between 0 and 1 and then using it as the weight of the corresponding point in the RGB feature map, the amount of calculation can be reduced and the feature fusion speed can be accelerated.
[0047] Step 103: Detect whether there is a target in the image acquisition area according to the fused feature map.
[0048] In a possible implementation, the detecting whether there is a target in the image acquisition area according to the fused feature map includes: performing class prediction on each point of the fused feature map to obtain the class and classification confidence of each point of the fused feature map, where the class includes a target class and a non-target class, the target class is the class to which the target belongs, and the non-target class is the class to which non-targets belong; counting the number of points predicted as the target class in the fused feature map, and counting the number of points predicted as the non-target class in the fused feature map; calculating a first average classification confidence according to the number of points predicted as the target class and the classification confidence of each point predicted as the target class, and calculating a second average classification confidence according to the number of points predicted as the non-target class and the classification confidence of each point predicted as the non-target class; if the number of points predicted as the target class is greater than the number of points predicted as the non-target class, and the first average classification confidence is greater than the second average classification confidence, then determine that the candidate object is a target; otherwise, determine that the candidate object is a non-target.
[0049] Among them, the target can be a mobile vendor, and the target class can be a food truck, a booth, a person, etc.; the non-target is an object other than the target, such as an object other than a mobile vendor, and the target class is a class other than the target class.
[0050] In this implementation manner, category prediction is performed on each point of the fused feature map to obtain the category and classification confidence of each point of the fused feature map. Among them, the categories include the target category and the non-target category. The target category is the category to which the target belongs, and the non-target category is the category to which the non-target belongs. Then, the first average classification confidence of the points classified as the target category is calculated, and the second average classification confidence of the points classified as the non-target category is calculated. If the first average classification confidence is greater than the second average classification confidence, it is determined that the candidate object is the target. If the number of points predicted as the target category is greater than the number of points predicted as the non-target category, and the first average classification confidence is not greater than the second average classification confidence, it is determined that the candidate object is the non-target. Otherwise, it is determined that the candidate object is the non-target. Since the candidate object is recognized as the target only when the proportion of the target in the fused feature map is greater than the proportion of the non-target and the average classification confidence of the target is greater than the average classification confidence of the non-target, it is beneficial to improve the accuracy of target detection.
[0051] Among them, the above steps 101 to 103 can be implemented by a target detection model, that is, target detection of the monitoring area is performed by a pre-trained target detection model. For example, a food truck, a booth, and a person of a street vendor are detected by a pre-trained street vendor detection model to determine the street vendor area.
[0052] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a target detection model provided by an embodiment of the present application. As Figure 2As shown in the figure, the object detection model includes two backbone networks (NN backbone), a feature map fusion module, and a YOLOv3 network. Specifically, based on the YOLOv3 network, the object detection model uses two backbone networks to extract features from IR images and RGB images respectively to obtain IR feature maps and RGB feature maps. Among them, the backbone network can be ResNet50 (residual network). ResNet50 has a shortcut and a bottleneck structure. The shortcut can make the model deeper, and the bottleneck structure can improve the running speed of the model. The combination of the two can better extract features. The feature map fusion module is used to fuse the IR feature map and the RGB feature map to obtain a fused feature map. Then, through YOLOv3, the fused feature map is mapped to the final object detection result. Specifically, the YOLO layer in YOLOv3 is used to map the fused feature map to the final object detection result. Among them, the final object detection result mainly includes the coordinate points of the detection box and the classification confidence of the detection box, that is, the final object detection result includes the detection box marked in the RGB image and / or the IR image. The detection box frames a target, and the final object detection result also includes the classification confidence that the object framed by the detection box is the target.
[0053] Among them, the input RGB image and IR image of the object detection model are aligned and marked with prediction boxes. For example, collect the aligned IR images and RGB images of street vendors in the monitoring scenario, and the IR images and RGB images are marked with prediction boxes whose framed objects may be street vendors.
[0054] Please refer to Figure 3 , Figure 3 is a schematic structural diagram of a feature map fusion module provided by an embodiment of the present application. As Figure 3 shown, the process of the feature map fusion module fusing the feature maps is as follows: perform a sigmoid operation on each point of the IR feature map, so as to limit the pixel value of each point of the IR feature map between 0 and 1, and then multiply the pixel value of each point of the IR feature map limited between 0 and 1 by the pixel value of the corresponding point of the RGB feature map to obtain a fused feature map.
[0055] Among them, the advantage of fusing the IR feature map and the RGB feature map is that: obtain the attention area of the RGB feature map through the IR feature map that is not affected by light, thereby improving the detection accuracy; since IR is not sensitive to light, object detection can be stably used in scenarios such as at night.
[0056] Among them, the object detection model calculates the loss based on the following formula and uses the batch stochastic gradient descent algorithm for model training:
[0057]
[0058] There are a total of S×S grids, that is, the image input to the network of the target detection model is divided into S×S grids; each grid generates B prediction boxes (anchor boxes), and each prediction box will eventually get the corresponding detection box (bounding box) through the network, and finally S×S×B detection boxes will be obtained. Indicates whether the j-th prediction box of the i-th grid is responsible for a certain object. If so, otherwise The so-called responsible means that when the intersection-over-union ratio (IoU) of the jth prediction box of the i-th grid and the ground truth box of the target is the largest among the B prediction boxes of the i-th grid and the ground truth box of the target, the jth prediction box of the i-th grid is responsible for the target; since the shape and size of the jth prediction box of the i-th grid best match the target, at this time Indicates that the j-th prediction box of the i-th grid is not responsible for the target. coord and λ noobj is the weight parameter, and the default setting is 1. i Indicates the x value of the upper left corner coordinate point of the detection box. Indicates the x value of the upper left corner coordinate point of the prediction box. i Indicates the y value of the upper left corner coordinate point of the detection box. Indicates the y value of the upper left corner coordinate point of the prediction box. i Indicates the width of the detection box. Indicates the width of the prediction box. h i Indicates the height of the detection box. Indicates the height of the prediction box. Represents the parameter confidence. Indicates the confidence level of the predicted parameters. Represents the probability (0 or 1) that the true i-th box belongs to the j-th class. It indicates the probability that the predicted i-th box belongs to the j-th class. c represents: class, that is, there are as many c as there are boxes of different classes.
[0059] In an embodiment of the present application, the monitoring area is used as the image acquisition area, and RGB image acquisition and IR image acquisition are performed on the image acquisition area to obtain the RGB image and the IR image of the monitoring area. Feature extraction is performed on the RGB image to obtain an RGB feature map, and feature extraction is performed on the IR image to obtain an IR feature map. Then, the RGB feature map and the IR feature map are fused to obtain a fused feature map. Furthermore, it is detected whether there is a target in the monitoring area according to the fused feature map. Since the RGB image has characteristics such as high resolution and rich color texture information, and the IR image is insensitive to light, the fused feature map obtained by fusing the RGB feature map and the IR feature map has the characteristics of both the RGB image and the IR image at the same time. In the embodiment of the present application, the fused feature map is used for target detection, which can improve the accuracy of target detection and obtain a more accurate detection result.
[0060] In a possible implementation manner, if there is a target in the image acquisition area, the method further includes: extracting a face image from the RGB image to obtain a target face image; comparing the target face image with multiple pre-stored template face images; if the comparison is successful, obtaining the identity information corresponding to the target template face image, and performing a preset operation according to the identity information, where the target template face image is the template face image that is successfully compared with the target face image.
[0061] Among them, the target face image can also be obtained by performing face extraction on the RGB feature map or the fused feature map. The extraction and comparison of the target face image can be realized by a face recognition model, and the multiple pre-stored template face images are ID photos of personnel stored in the database. For example, the operator of a street vendor is identified through a face recognition model, and the ID information of the operator is automatically identified and uploaded to a relevant supervision website, thereby realizing automatic supervision. When a street vendor is detected, the target face image is obtained and input into the face recognition model to obtain face features, and then the face features are compared with the face features of the ID photos of the personnel in the database. If the comparison is successful, it means that the ID information corresponding to the street vendor actor has been found.
[0062] In this implementation manner, after it is determined that there is a target in the image acquisition area, that is, after it is determined that there is a target in the monitoring area, a target face image is extracted from the RGB image, and the extracted target face image is compared with multiple pre-stored template face images. Then, the identity information corresponding to the target template face image that is successfully compared with the target face image is obtained, and a preset operation is performed according to the identity information, so as to automatically supervise the target, without the need for supervisors to go to the monitoring area for on-site management, reducing the monitoring cost.
[0063] In a possible implementation, if the comparison fails, the method further includes: sending a prompt message to the handheld terminal, where the prompt message includes the RGB image, and the prompt message is used to indicate processing of the target, and the handheld terminal is a terminal held by a supervisor.
[0064] In this implementation, if the target face image fails to match any of the multiple pre-stored template face images, automatic supervision cannot be performed. Then, a prompt message including the RGB image is sent to the terminal held by the supervisor, indicating that the supervisor should process the target, thereby ensuring that the target can be supervised.
[0065] In summary, the embodiments of the present application use the fusion of IR images and RGB images for target detection, which can resist the influence of light. It can achieve very stable detection and recognition accuracy in scenarios with poor light conditions such as at night, but with a very high demand for detection and recognition accuracy.
[0066] The method of the embodiments of the present application is described in detail above. Next, the device of the embodiments of the present application is provided.
[0067] Please refer to Figure 4 , Figure 4 FIG. is a schematic structural diagram of a target detection device 400 provided by an embodiment of the present application. The target detection device can be applied to an electronic device. The target detection device 400 may include an extraction unit 401, a fusion unit 402, and a detection unit 403. The detailed descriptions of each unit are as follows:
[0068] The extraction unit 401 is configured to extract features from the RGB image to obtain an RGB feature map, and extract features from the IR image to obtain an IR feature map, where the RGB image and the IR image are obtained by collecting images of the same image acquisition area;
[0069] The fusion unit 402 is configured to fuse the RGB feature map and the IR feature map to obtain a fused feature map;
[0070] The detection unit 403 is configured to detect whether a target exists in the image acquisition area according to the fused feature map.
[0071] In a possible implementation, the fusion unit 402 is specifically configured to: map the first pixel value of each point of the IR feature map to between 0 and 1 to obtain the second pixel value of each point of the IR feature map; multiply the second pixel value of each point of the IR feature map by the third pixel value of the corresponding position point of the RGB feature map to obtain a plurality of fourth pixel values; and combine the plurality of fourth pixel values to obtain the fused feature map.
[0072] In a possible implementation manner, in terms of extracting features from the RGB image to obtain an RGB feature map and extracting features from the IR image to obtain an IR feature map, the extraction unit 401 is specifically configured to: label a prediction box in the RGB image and label the prediction box in the IR image, where the prediction box is used to frame a candidate object; extract features from the region of the RGB image framed by the prediction box to obtain the RGB feature map, and extract features from the region of the IR image framed by the prediction box to obtain the IR feature map.
[0073] In a possible implementation manner, in terms of detecting whether there is a target in the image acquisition area according to the fusion feature map, the detection unit 403 is specifically configured to: perform class prediction on each point of the fusion feature map to obtain the class and classification confidence of each point of the fusion feature map, where the class includes a target class and a non-target class, the target class is the class to which the target belongs, and the non-target class is the class to which the non-target belongs; count the number of points predicted as the target class in the fusion feature map and count the number of points predicted as the non-target class in the fusion feature map; calculate a first average classification confidence according to the number of points predicted as the target class and the classification confidence of each point predicted as the target class, and calculate a second average classification confidence according to the number of points predicted as the non-target class and the classification confidence of each point predicted as the non-target class; if the number of points predicted as the target class is greater than the number of points predicted as the non-target class and the first average classification confidence is greater than the second average classification confidence, determine that the candidate object is a target; otherwise, determine that the candidate object is a non-target.
[0074] In a possible implementation manner, if there is a target in the image acquisition area, the detection unit 403 is further configured to: extract a face image from the RGB image to obtain a target face image; compare the target face image with multiple pre-stored template face images; if the comparison is successful, obtain the identity information corresponding to the target template face image and perform a preset operation according to the identity information, where the target template face image is the template face image that is successfully compared with the target face image.
[0075] In a possible implementation manner, if the comparison fails, the detection unit 403 is further configured to: send a prompt message to the handheld terminal, where the prompt message includes the RGB image, and the prompt message is used to indicate processing of the target, and the handheld terminal is a terminal held by a supervisor.
[0076] It should be noted that the implementation of each unit can also be correspondingly referred to Figure 1The corresponding description of the method embodiments shown. Of course, the target detection device 400 provided in the embodiments of the present application includes, but is not limited to, the above unit modules. For example, the target detection device 400 may further include a storage unit 404, and the storage unit 404 may be used to store the program code and data of the target detection device 400.
[0077] In Figure 4 In the described target detection device 400, the monitoring area is used as the image acquisition area, RGB image acquisition and IR image acquisition are performed on the image acquisition area to obtain the RGB image and IR image of the monitoring area, feature extraction is performed on the RGB image to obtain an RGB feature map, and feature extraction is performed on the IR image to obtain an IR feature map; then the RGB feature map and the IR feature map are fused to obtain a fused feature map; and then it is detected whether there is a target in the monitoring area according to the fused feature map. Since the RGB image has characteristics such as high resolution and rich color texture information, and the IR image has the characteristic of being insensitive to light, the fused feature map obtained by fusing the RGB feature map and the IR feature map has the characteristics of both the RGB image and the IR image at the same time; in the embodiments of the present application, the fused feature map is used for target detection, which can improve the accuracy of target detection and obtain a more accurate detection result.
[0078] Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of an electronic device 510 provided in the embodiments of the present application. The electronic device 510 includes a processor 511, a memory 512, and a communication interface 513, and the above processor 511, memory 512, and communication interface 513 are interconnected through a bus 514.
[0079] The memory 512 includes, but is not limited to, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM). The memory 512 is used for related computer programs and data. The communication interface 513 is used to receive and send data.
[0080] The processor 511 may be one or more central processing units (CPUs). When the processor 511 is a single CPU, the CPU may be a single-core CPU or a multi-core CPU.
[0081] The processor 511 in the electronic device 510 is configured to read the computer program code stored in the memory 512 and execute Figure 1 the method shown.
[0082] It should be noted that the implementation of each operation can also be correspondingly referred to Figure 1 the corresponding description of the method embodiment shown.
[0083] In Figure 5 the described electronic device 510, the monitoring area is used as an image acquisition area, RGB image acquisition and IR image acquisition are performed on the image acquisition area to obtain the RGB image and IR image of the monitoring area, feature extraction is performed on the RGB image to obtain an RGB feature map, and feature extraction is performed on the IR image to obtain an IR feature map; then the RGB feature map and the IR feature map are fused to obtain a fused feature map; and then whether there is a target in the monitoring area is detected according to the fused feature map. Since the RGB image has characteristics such as high resolution and rich color texture information, and the IR image is insensitive to light, the fused feature map obtained by fusing the RGB feature map and the IR feature map has the characteristics of both the RGB image and the IR image at the same time; in the embodiment of the present application, the fused feature map is used for target detection, which can improve the accuracy of target detection and obtain more accurate detection results.
[0084] The embodiment of the present application further provides a chip, the chip includes at least one processor, a memory and an interface circuit, the memory, the transceiver and the at least one processor are interconnected through lines, and a computer program is stored in the at least one memory; when the computer program is executed by the processor, Figure 1 the method flow shown is realized.
[0085] The embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when it runs on a computer, Figure 1 the method flow shown is realized.
[0086] The embodiment of the present application further provides a computer program product, and when the computer program product runs on a computer, Figure 1 the method flow shown is realized.
[0087] Among them, the processor mentioned in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0088] Still among them, the memory mentioned in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0089] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) is integrated in the processor.
[0090] It should be noted that the memories described herein are intended to include but not be limited to these and any other suitable types of memories.
[0091] Among them, the first, second, third, fourth, and various numerical numbers involved in this article are only for the convenience of description and are not used to limit the scope of this application.
[0092] Among them, the term "and / or" in this article is only an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0093] Among them, in various embodiments of this application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0094] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0095] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.
[0096] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.
[0097] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0098] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit.
[0099] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods shown in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0100] The steps in the method of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs.
[0101] The modules in the device of the embodiments of the present application can be combined, divided, and deleted according to actual needs.
[0102] The above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present application.
Claims
1. A target detection method, characterized in that, the method includes: performing feature extraction on the RGB image to obtain an RGB feature map, and performing feature extraction on the IR image to obtain an IR feature map, where the RGB image and the IR image are obtained by collecting images of the same image acquisition area; including: annotating a prediction box in the RGB image, and annotating the prediction box in the IR image, where the prediction box is used to frame a candidate object; performing feature extraction on the area of the RGB image framed by the prediction box to obtain the RGB feature map, and performing feature extraction on the area of the IR image framed by the prediction box to obtain the IR feature map; fusing the RGB feature map and the IR feature map to obtain a fused feature map; detecting whether there is a target in the image acquisition area according to the fused feature map; including: performing class prediction on each point of the fused feature map to obtain the class and classification confidence of each point of the fused feature map, where the class includes a target class and a non-target class, the target class is the class to which the target belongs, and the non-target class is the class to which the non-target belongs; counting the number of points predicted as the target class in the fused feature map, and counting the number of points predicted as the non-target class in the fused feature map; calculating a first average classification confidence according to the number of points predicted as the target class and the classification confidence of each point predicted as the target class, and calculating a second average classification confidence according to the number of points predicted as the non-target class and the classification confidence of each point predicted as the non-target class; if the number of points predicted as the target class is greater than the number of points predicted as the non-target class, and the first average classification confidence is greater than the second average classification confidence, then determine that the candidate object is a target; otherwise, determine that the candidate object is a non-target.
2. The method according to claim 1, characterized in that, the fusing the RGB feature map and the IR feature map to obtain a fused feature map includes: mapping the first pixel value of each point of the IR feature map to between 0 and 1 to obtain the second pixel value of each point of the IR feature map; multiplying the second pixel value of each point of the IR feature map by the third pixel value of the corresponding position point of the RGB feature map to obtain a plurality of fourth pixel values; combining the plurality of fourth pixel values to obtain the fused feature map.
3. The method according to claim 1 or 2, characterized in that, if there is a target in the image acquisition area, the method further includes: extracting a face image from the RGB image to obtain a target face image; comparing the target face image with a plurality of pre-stored template face images; if the comparison is successful, obtaining the identity information corresponding to the target template face image, and performing a preset operation according to the identity information, where the target template face image is the template face image that is successfully compared with the target face image.
4. The method according to claim 3, characterized in that, If the comparison fails, the method further includes: Sending a prompt message to the handheld terminal, the prompt message including the RGB image, the prompt message being used to indicate processing of the target, and the handheld terminal being a terminal held by a supervisor.
5. A target detection device, characterized in that the device includes: An extraction unit, configured to extract features from the RGB image to obtain an RGB feature map, and extract features from the IR image to obtain an IR feature map, where the RGB image and the IR image are obtained by collecting images of the same image acquisition area; specifically, the extraction unit is configured to mark prediction boxes in the RGB image, and mark the prediction boxes in the IR image, where the prediction boxes are used to frame candidate objects; extract features from the area of the RGB image framed by the prediction boxes to obtain the RGB feature map, and extract features from the area of the IR image framed by the prediction boxes to obtain the IR feature map; A fusion unit, configured to fuse the RGB feature map and the IR feature map to obtain a fused feature map; A detection unit, configured to detect whether there is a target in the image acquisition area according to the fused feature map; including: performing class prediction on each point of the fused feature map to obtain the class and classification confidence of each point of the fused feature map, where the class includes a target class and a non-target class, the target class being the class to which the target belongs, and the non-target class being the class to which non-targets belong; counting the number of points predicted as the target class in the fused feature map, and counting the number of points predicted as the non-target class in the fused feature map; calculating a first average classification confidence according to the number of points predicted as the target class and the classification confidence of each point predicted as the target class, and calculating a second average classification confidence according to the number of points predicted as the non-target class and the classification confidence of each point predicted as the non-target class; if the number of points predicted as the target class is greater than the number of points predicted as the non-target class, and the first average classification confidence is greater than the second average classification confidence, then determining that the candidate object is a target; otherwise, determining that the candidate object is a non-target.
6. The device according to claim 5, characterized in that the fusion unit is specifically configured to: Map the first pixel value of each point of the IR feature map to between 0 and 1 to obtain the second pixel value of each point of the IR feature map; Multiply the second pixel value of each point of the IR feature map by the third pixel value of the corresponding position point of the RGB feature map to obtain a plurality of fourth pixel values; Combine the plurality of fourth pixel values to obtain the fused feature map.
7. An electronic device, characterized in that it includes a processor, a memory, a communication interface, and one or more programs, the one or more programs are stored in the memory and are configured to be executed by the processor, and the programs include instructions for performing the steps in the method according to any one of claims 1-4.
8. A computer-readable storage medium, characterized in that, it stores a computer program for electronic data exchange, wherein the computer program causes a computer to execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Visible light infrared image enhancement color fusion method based on vision attention features
CN106952246A
Multispectral pedestrian detection method based on feature fusion deep neural network
CN111898427A