Detection Method, System and Electronic Device for Illegal Luggage
By using deep learning networks in the vault to detect and monitor illegal bags in the surveillance images, the problems of management difficulty and workload caused by relying on manual detection in the existing technology are solved, automated detection and alarm are realized, and the security of the vault is improved.
Patent Information
- Application Number
- CN202210044078.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-01-14
AI Technical Summary
The existing vault relies on manual detection of illegal luggage, which leads to problems such as difficult management, high workload and prone to negligence.
By monitoring the object detection of the image, the luggage image is extracted and input into the deep learning network for detection, the relationship between the feature vector and the hyperspherical body is calculated, whether the luggage is an illegal luggage, and alarm is issued when the illegal luggage is detected.
It realizes automatic detection of illegal luggage, reduces the workload of administrators or security personnel, improves the accuracy and efficiency of inspections, and ensures the safety of the vault.
Smart Images

Figure CN114419550B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of security detection, and in particular to a method for detecting a violative luggage, a system for detecting a violative luggage, and an electronic device. Background Art
[0002] A vault is a cash business repository set up by each bank for handling cash business, and is a special repository for centrally storing cash and securities. To ensure the security of the vault, any other luggage except compliant luggage is prohibited from being carried into the bank vault, such as plastic bags, handbags, briefcases, etc. Carrying violative luggage into the vault not only poses a security risk to the vault, but also brings unnecessary trouble to individuals. Currently, the detection of luggage entering and leaving the vault still relies on manual inspection. Management personnel or security personnel need to conduct inspections through monitoring or in person at the entrance or exit to check whether the people entering and leaving carry violative luggage. This increases the workload of the administrators or security personnel, and is affected by the attention of personnel or the flow of people, and occasional negligence may also occur. How to strengthen the detection of luggage carrying is a major difficulty in vault management. Summary of the Invention
[0003] The purpose of the embodiments of the present invention is to provide a method, a system and an electronic device for detecting a violative luggage. Through object detection, a second image including a luggage image is intercepted from a monitoring image, a deep learning network performs detection on the second image to obtain a feature vector, and then determines whether the luggage in the second image is a violative luggage based on the feature vector, and alarms in time when a violative luggage is detected, reducing the workload of administrators or security personnel.
[0004] To achieve the above purpose, the first aspect of the present invention provides a method for detecting a violative luggage, the method comprising:
[0005] Preprocessing the obtained monitoring image to obtain a first image including at least one moving object;
[0006] Extracting a second image including a luggage image from the first image through object detection;
[0007] Inputting the second image into a deep learning network model to obtain a feature vector corresponding to the second image;
[0008] Calculating the relationship between the feature vector corresponding to the second image and a hypersphere to determine whether the luggage image in the second image includes a violative luggage image;
[0009] Wherein, the hypersphere is composed of feature vectors corresponding to samples including violative luggage images.
[0010] Further, the deep learning network model is trained through the following steps:
[0011] Obtain training sample data, where the training sample data includes positive samples and negative samples;
[0012] Input the positive samples into the deep learning network model after the previous round of iteration, and calculate the feature vectors corresponding to the positive samples;
[0013] Calculate the center point of the positive samples according to the feature vectors corresponding to the positive samples;
[0014] Calculate the corrected center point according to the position of the center point of the positive samples;
[0015] Update the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples and the positive sample loss function;
[0016] Input the negative samples and positive samples into the updated deep learning network model, and calculate the feature vectors corresponding to the negative samples and the feature vectors corresponding to the positive samples;
[0017] Update the parameters of the deep learning network model according to the feature vectors corresponding to the negative samples, the feature vectors corresponding to the positive samples, the corrected center point, and the negative sample loss function;
[0018] Repeat the above steps a preset number of times, or when the change of the negative sample loss function meets the preset conditions, end the iteration to obtain the trained deep learning network model.
[0019] Optionally, the updating of the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples and the positive sample loss function includes:
[0020] Calculate the center point of the positive samples according to the feature vectors corresponding to the positive samples;
[0021] Calculate the corrected center point according to the position of the center point of the positive samples;
[0022] Update the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples, the corrected center point, and the positive sample loss function.
[0023] Optionally, the hypersphere includes a center and a radius;
[0024] The center of the hypersphere is the position of the last corrected center point during the training process of the deep learning network model;
[0025] The radius of the hypersphere is determined according to the distances between the positive samples and negative samples and the center of the hypersphere respectively. The deep learning network model is trained with positive and negative samples. The training process is a process in which positive samples are pulled in and negative samples are pushed away. After training, a hypersphere is formed. Inside the hypersphere are positive samples and outside are negative samples.
[0026] Optionally, the calculation formula for the center point of the positive sample is as follows:
[0027]
[0028] where l is the number of iterations, and the center point of the initial iteration is denoted as c1, f i is the feature vector corresponding to the positive sample;
[0029] The corrected center point is calculated by the following formula:
[0030] c b = βc l-1 + (1 - β)c l ;
[0031] where c b is the corrected center point, and β takes a value of 0.98. Calculating the center point according to the feature vector corresponding to the positive sample can cluster the positive samples around the center point; using the method of exponential weighted average to calculate the corrected center point can make the center point change drastically in the early stage of model training and accelerate model convergence.
[0032] Optionally, updating the parameters of the deep learning network model according to the feature vector corresponding to the positive sample, the position of the corrected center point, and the positive sample loss function includes:
[0033] Calculating the Euclidean distance from the feature vector corresponding to each positive sample to the corrected center point, denoted as dis i ;
[0034] Using the gradient descent method, calculating and updating the parameters of the deep learning network model according to the Euclidean distance and the positive sample loss function;
[0035] The positive sample loss function is denoted as: where dis i is the Euclidean distance from the positive sample to the corrected center point, and e is the maximum value of the distance from the positive sample to the corrected center point. Through the iteration of the positive sample loss function, all positive samples are restricted within a hypersphere with a radius of e.
[0036] Optionally, updating the parameters of the deep learning network model according to the feature vector corresponding to the negative sample, the feature vector corresponding to the positive sample, the corrected center point, and the negative sample loss function includes:
[0037] Calculating the Euclidean distance from the feature vector corresponding to each negative sample to the corrected center point and the Euclidean distance from the feature vector corresponding to each positive sample to the corrected center point
[0038] Using the gradient descent method, calculate and update the parameters of the deep learning network model according to the above Euclidean distance and the negative sample loss function;
[0039] The negative sample loss function is denoted as: where is the distance from the negative sample closest to the distance correction center point to the correction center point, and g is the interval between the distance from the positive sample to the correction center point and the distance from the negative sample closest to the distance correction center point to the correction center point. By iterating the negative sample loss function, all negative samples are restricted outside the sphere region with the sum of the radius of the hypersphere formed by the positive samples and g as the radius. During the iteration process, the hard sample mining method is adopted, and the distance from the negative sample closest to the distance correction center point to the correction center point is used to iterate the loss function to improve the effect of feature extraction.
[0040] Optionally, the positive sample is a compliant luggage image, and the negative sample is a non-compliant luggage image;
[0041] The obtaining of the training sample data includes:
[0042] Obtain compliant luggage images and non-compliant luggage images;
[0043] Enhance the compliant luggage images and non-compliant luggage images by using one or more of occlusion, rotation, size normalization, and brightness change, so that the number of compliant luggage images is balanced with the number of non-compliant luggage images. By data augmentation, more positive samples are obtained to solve the problem that the number of positive sample images of compliant luggage with a fixed format is small and there may be occlusion effects.
[0044] Optionally, the preprocessing of the obtained images to obtain a first image including at least one moving target includes:
[0045] Obtain consecutive first-frame image, second-frame image, and third-frame image composed of pixels of any channel in the surveillance image;
[0046] Subtract the first-frame image from the second-frame image to obtain a first difference image, and subtract the second-frame image from the third-frame image to obtain a second difference image;
[0047] Binarize the first difference image and the second difference image to obtain a first binarized image and a second binarized image;
[0048] Perform a bitwise AND operation on the first binarized image and the second binarized image to obtain a mask image of the moving target;
[0049] Calculate the connected components of the mask image;
[0050] Filter out the connected components smaller than a preset size in the mask image to obtain the connected components of the moving target;
[0051] Calculate the circumscribed quadrilateral of the connected region of the moving target;
[0052] Enlarge the circumscribed quadrilateral by a preset ratio, and obtain the monitoring image area corresponding to the enlarged circumscribed quadrilateral as the first image of the moving target. Through the above preprocessing steps, it is possible to quickly determine the existence of a moving target in the monitoring image, and quickly locate and determine the luggage target that occupies a small area in the moving target. On the one hand, it reduces the interference of stationary objects similar to luggage in the warehouse. On the other hand, since carrying luggage is a low-probability event, preprocessing can effectively eliminate images that do not require luggage detection, saving computing resources.
[0053] Optionally, after determining whether the luggage image in the second image includes a violation luggage image, the method further includes:
[0054] Statistically calculate the ratio of the number of frames of monitoring images with violation luggage images within a preset time. When the ratio exceeds the preset ratio, it is determined that the monitoring image includes a violation luggage image. Statistically calculate the detection results in multiple frames of monitoring images within a preset time, and perform suppression according to the statistical results to reduce the interference of special situations and special angles, and enhance the accuracy of the detection results.
[0055] The second aspect of the present invention provides a detection system for violation luggage, and the system includes:
[0056] An image preprocessing module for preprocessing the acquired monitoring image to obtain a first image including at least one moving target;
[0057] A target detection module for extracting a second image including a luggage image from the first image through target detection;
[0058] A feature extraction module for inputting the second image into a deep learning network model to obtain a feature vector corresponding to the second image;
[0059] A comparison and judgment module for calculating the relationship between the feature vector corresponding to the second image and the hypersphere, and determining whether the luggage image in the second image includes a violation luggage image;
[0060] The hypersphere is composed of the feature vectors corresponding to the samples including violation luggage images. This system intercepts a second image including a luggage image from the monitoring image through target detection, trains a deep learning network to detect the second image to obtain a feature vector, and then determines whether the luggage in the second image is a violation luggage based on the feature vector, and alarms in time when a violation luggage is detected, reducing the workload of administrators or security personnel.
[0061] Optionally, the deep learning network model is trained through the following steps:
[0062] Obtain training sample data, where the training sample data includes positive samples and negative samples;
[0063] Input the positive samples into the deep learning network model after the previous round of iteration, and calculate the feature vectors corresponding to n positive samples;
[0064] Update the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples and the positive sample loss function;
[0065] Input the negative samples and the positive samples into the updated deep learning network model, and calculate the feature vectors corresponding to the negative samples and the feature vectors corresponding to the positive samples;
[0066] Update the parameters of the deep learning network model according to the feature vectors corresponding to the negative samples, the feature vectors corresponding to the positive samples, the corrected center point, and the negative sample loss function;
[0067] Repeat the above steps for a preset number of times, or when the change of the negative sample loss function meets the preset conditions, end the iteration to obtain a trained deep learning network model.
[0068] Optionally, the updating of the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples and the positive sample loss function includes:
[0069] Calculate the center point of the positive samples according to the feature vectors corresponding to the positive samples;
[0070] Calculate the corrected center point according to the position of the center point of the positive samples;
[0071] Update the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples, the corrected center point, and the positive sample loss function.
[0072] Optionally, the hypersphere includes a center and a radius;
[0073] The center of the hypersphere is the position of the last corrected center point during the training process of the deep learning network model;
[0074] The radius of the hypersphere is determined according to the distances between the positive samples and the negative samples and the center of the hypersphere respectively. The deep learning network model is trained using positive and negative samples. The training process is a process in which positive samples are pulled in and negative samples are pushed away. After training, a hypersphere is formed. Inside the hypersphere are positive samples, and outside are negative samples.
[0075] Optionally, the calculation formula for the center point of the positive samples is as follows:
[0076]
[0077] where l is the number of iterations, the center point of the first iteration is denoted as c1, and f i is the feature vector corresponding to the positive sample;
[0078] The corrected center point is calculated by the following formula:
[0079] c b = βc l-1 + (1 - β)c l ;
[0080] where c b is the corrected center point, and β takes a value of 0.98. Calculating the center point according to the feature vector corresponding to the positive sample can achieve clustering the positive samples around the center point; using the method of exponential weighted average to calculate the corrected center point can enable the center point to change drastically in the early stage of model training and accelerate model convergence.
[0081] Optionally, updating the parameters of the deep learning network model according to the feature vector corresponding to the positive sample, the position of the corrected center point, and the positive sample loss function includes:
[0082] Calculating the Euclidean distance from the feature vector corresponding to each positive sample to the corrected center point, denoted as dis i ;
[0083] Using the gradient descent method, calculate and update the parameters of the deep learning network model according to the Euclidean distance and the positive sample loss function;
[0084] The positive sample loss function is denoted as: where dis i is the Euclidean distance from the positive sample to the corrected center point, and e is the maximum value of the distance from the positive sample to the corrected center point. Through the iteration of the positive sample loss function, all positive samples are restricted within a hypersphere with a radius of e.
[0085] Optionally, updating the parameters of the deep learning network model according to the feature vector corresponding to the negative sample, the feature vector corresponding to the positive sample, the corrected center point, and the negative sample loss function includes:
[0086] Calculating the Euclidean distance from the feature vector corresponding to each negative sample to the corrected center point and the Euclidean distance from the feature vector corresponding to each positive sample to the corrected center point
[0087] Using the gradient descent method, calculate and update the parameters of the deep learning network model according to the Euclidean distance and the negative sample loss function;
[0088] The negative sample loss function is denoted as: where Let \(d\) be the distance from the negative sample closest to the correction center point to the correction center point, and \(g\) be the interval between the distance from the positive sample to the correction center point and the distance from the negative sample closest to the correction center point to the correction center point. During the iteration process, the hard sample mining method is adopted, and the distance from the negative sample closest to the correction center point to the correction center point is used to iterate the loss function to improve the effect of feature extraction.
[0089] Optionally, the image preprocessing module includes:
[0090] A monitoring image acquisition module, configured to acquire consecutive first-frame, second-frame, and third-frame images composed of pixels in any channel of the monitoring image;
[0091] A difference calculation module, configured to calculate the difference between the first-frame image and the second-frame image to obtain a first difference image, and calculate the difference between the second-frame image and the third-frame image to obtain a second difference image;
[0092] A binarization module, configured to binarize the first difference image and the second difference image to obtain a first binarized image and a second binarized image;
[0093] A bitwise AND calculation module, configured to perform a bitwise AND operation on the first binarized image and the second binarized image to obtain a mask image of the moving target;
[0094] A connected component calculation module, configured to calculate the connected components of the mask image;
[0095] A connected component filtering module, configured to filter the connected components smaller than a preset size in the mask image to obtain the connected components of the moving target;
[0096] An enclosing quadrilateral calculation module, configured to calculate the enclosing quadrilateral of the connected components of the moving target;
[0097] An enlargement module, configured to enlarge the enclosing quadrilateral by a preset ratio, and obtain the monitoring image area corresponding to the enlarged enclosing quadrilateral as the first image of the moving target. Each sub-module in the image preprocessing module cooperates with each other to quickly determine the existence of a moving target in the monitoring image, and can quickly locate and determine the luggage target that occupies a small area of the image. On the one hand, it reduces the interference of stationary objects similar to luggage in the warehouse. On the other hand, since carrying luggage is a low-probability event, the preprocessing can effectively eliminate the images that do not need to be detected for luggage, saving computing resources.
[0098] Optionally, the system further includes:
[0099] The result suppression module is used to count the ratio of the images of illegal luggage cases in multiple frames of surveillance images within a preset time. When the ratio exceeds the preset ratio, it is determined that the surveillance images include images of illegal luggage cases. The result suppression module counts the detection results in multiple frames of surveillance images within a preset time, performs suppression according to the statistical results, reduces the interference of special situations and special angles, and enhances the accuracy of the detection results.
[0100] The third aspect of the present invention provides an electronic device, including:
[0101] A memory for storing a computer program;
[0102] A processor for running the computer program to execute the detection method for illegal luggage cases.
[0103] The fourth aspect of the present invention provides a machine-readable storage medium, on which instructions are stored, and the instructions are used to make a machine execute the detection method for illegal luggage cases.
[0104] On the other hand, the present invention provides a computer program product, including a computer program, characterized in that the computer program realizes the detection method for illegal luggage cases when executed by a processor.
[0105] Through the above technical solutions, the present invention can use the surveillance images collected by existing surveillance devices, extract the second images of luggage cases from the surveillance images through target detection, use the trained deep learning network to detect the second images, determine whether the luggage cases in the second images are illegal luggage cases, count the detection results in multiple frames of surveillance images within a preset time, perform suppression according to the statistical results, reduce the interference of special situations and special angles, and enhance the accuracy of the detection results. It realizes the automatic detection of whether there are illegal luggage cases, reduces the workload of administrators or security personnel, and ensures the safety of the vault.
[0106] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific embodiments section. BRIEF DESCRIPTION OF THE DRAWINGS
[0107] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification, and are used to explain the embodiments of the present invention together with the following specific embodiments, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0108] Figure 1 is a flowchart of the detection method for illegal luggage cases provided by an embodiment of the present invention;
[0109] Figure 2 is a flowchart of image preprocessing for the detection method for illegal luggage cases provided by an embodiment of the present invention;
[0110] Figure 3 It is a flow chart for training a deep learning network model for the detection method of illegal luggage provided by an embodiment of the present invention;
[0111] Figure 4 It is a schematic diagram of the global average pooling operation provided by an embodiment of the present invention;
[0112] Figure 5 It is a flow chart of the detection method of illegal luggage provided by another embodiment of the present invention;
[0113] Figure 6 It is a block diagram of the detection system of illegal luggage provided by an embodiment of the present invention;
[0114] Figure 7 It is a block diagram of the detection system of illegal luggage provided by another embodiment of the present invention. Specific Embodiments
[0115] The following further elaborates on the specific embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.
[0116] Embodiment 1
[0117] Figure 1 It is a flow chart of the detection method of illegal luggage provided by an embodiment of the present invention. As Figure 1 shown, the method includes:
[0118] Step 1: Preprocess the acquired surveillance image to obtain a first image including at least one moving target. As Figure 2 shown, it specifically includes:
[0119] S101: Obtain consecutive first-frame image, second-frame image, and third-frame image composed of pixels of any channel in the surveillance image.
[0120] The surveillance images collected by the surveillance device are generally color images. Any color image is composed of multiple color pixels, and any color pixel includes three channels: red (R), green (G), and blue (B). To reduce the amount of calculation, three consecutive frames of images are extracted from the surveillance image composed of pixels of any one of the channels, and are respectively denoted as the first-frame image, the second-frame image, and the third-frame image.
[0121] S102: Subtract the first-frame image from the second-frame image to obtain a first difference image, and subtract the second-frame image from the third-frame image to obtain a second difference image. Denote the first difference image as P1 and the second difference image as P2.
[0122] S103: Binarize the first difference image and the second difference image to obtain a first binarized image and a second binarized image. Denote the first binarized image as T1 and the second binarized image as T2.
[0123] S104: Perform a bitwise AND operation on the first binarized image and the second binarized image to obtain a mask image of the moving target. For example, for two binarized images, assuming that a certain point is 1 and 0 respectively, the result of the AND operation is 0; if a certain point is 1 for both, the result of the bitwise AND operation is 1, thus obtaining the mask image.
[0124] S105: Calculate the connected components of the mask image.
[0125] S106: Filter out the connected components smaller than a preset size in the mask image to obtain the connected components of the moving target. The preset value of the connected components is set according to empirical values. If all the connected components are filtered out, it can be understood that there is no moving target in the image. If there are larger connected components, each connected component is a potential moving target.
[0126] S107: Calculate the circumscribed quadrilateral of the connected components of the moving target. Assume that the coordinates of the upper, lower, left, and right limit positions of the pixels of the connected component coordinates are (y1, y2, x1, x2) respectively, then the coordinates of the upper left corner and the lower right corner of the quadrilateral are (x1, y1) and (x2, y2).
[0127] S108: Enlarge the circumscribed quadrilateral by a preset ratio, and obtain the monitoring image area corresponding to the enlarged circumscribed quadrilateral as the first image of the moving target. To avoid the circumscribed quadrilateral being unable to extract all the moving targets, the circumscribed quadrilateral is enlarged by a preset ratio to ensure that the obtained first image contains the complete moving target. In some embodiments, the circumscribed quadrilateral is enlarged to 1.2 times its original size, which is equivalent to an outward expansion of 20% of the circumscribed quadrilateral.
[0128] It should be noted that if all the connected components are filtered out, it can be understood that there is no moving target in the image. If there is no moving target, there is no need to process the obtained image, and this round of processing ends. Directly re - obtain the monitoring image for pre - processing.
[0129] Through the above pre - processing steps, it is possible to quickly determine the existence of a moving target in the monitoring image, and quickly locate and determine the luggage target that occupies a small area in the moving target. On the one hand, it reduces the interference of stationary objects similar to luggage in the warehouse. On the other hand, since carrying luggage is a low - probability event, pre - processing can effectively eliminate the images that do not require luggage detection, saving computing resources.
[0130] Step 2: Through target detection, extract a second image including the luggage image from the first image.
[0131] In this application, the object detection algorithm based on YOLOv5 is used to extract the second image of the luggage from the first image. The first image is input into the YOLOv5 network to extract the feature map, which is the position map of the luggage in the first image. Similarly, the target box of the feature map is expanded by 20% to obtain the second image of the luggage. The YOLOv5 (You Only Look Once v5) object detection algorithm can more accurately identify the luggage in the image and improve the detection accuracy.
[0132] Step 3: Input the second image into the deep learning network model to obtain the feature vector corresponding to the second image.
[0133] In this embodiment, as Figure 3 shown, the deep learning network model is trained through the following steps:
[0134] 1) Obtain training sample data, which includes positive samples and negative samples. The positive samples are compliant luggage images, and the negative samples are non-compliant luggage images. Since compliant luggage has a fixed format and the sample size is small, after obtaining the compliant luggage images and non-compliant luggage images, it is necessary to use data augmentation methods to obtain a balanced number of compliant luggage images and non-compliant luggage images. In some embodiments, the compliant luggage is a vault luggage, and the non-compliant luggage includes personal handbags, personal suitcases, etc.
[0135] In some embodiments, one or more of occlusion, rotation, size normalization, or brightness change are used to augment the compliant luggage images. Occlusion is to use randomly colored, randomly sized, and randomly shaped color blocks to occlude the target, with a maximum of no more than one-third of the luggage area. This augmentation also provides sample data for possible occlusions during the handling of compliant luggage. Rotation refers to randomly rotating the angle. Size normalization and brightness change are both conventional means of image processing and will not be elaborated here. In some embodiments, one or more of the above methods are randomly rotated for augmentation. For example, image A may be augmented by occlusion, size normalization, and rotation, while image B may be augmented by brightness change and occlusion. To meet the data volume requirements, a compliant luggage image is usually augmented to 4 - 6 images.
[0136] For non-compliant luggage images, since their types and styles are sufficient and the data volume is large; if the data volume is sufficient, data augmentation may not be necessary, and in the case of insufficient data volume, random augmentation can be performed. By data augmentation, more positive samples are obtained, solving the problem of fewer positive sample images of compliant luggage with a fixed format and possible occlusion effects.
[0137] 2) Input the positive samples into the deep learning network model after the previous round of iteration, and calculate the feature vectors corresponding to the positive samples. In this embodiment, n samples are input, and n m-dimensional vectors f are obtained.
[0138] The final output value of the traditional model is generally a feature map. For example, for an image with a size of W x H, after passing through the convolutional layer, it becomes W x H x C, where C is the number of channels. After passing through the pooling layer, it becomes W / 2 x H / 2 x C. Conventional deep learning networks are all composed of convolutional layers, pooling layers, or their variants. Assuming that the final output result of a conventional deep learning network is a feature map of M x N x C, the last layer uses global average pooling (GAP), that is, each M x N map becomes the average value of all the data in this feature map, and the final output result is 1 x C. Then, through a fully connected layer, a 1-dimensional feature map is finally obtained. As Figure 4 shown, a 3x3 feature map passes through global average pooling (i.e., the average value of 9 points) to obtain an average value as the final output. Then, the final output of a 3x3x4 feature map is a 1x4 vector.
[0139] 3) Update the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples and the positive sample loss function, specifically including:
[0140] 3-1) Calculate the center point of the positive samples according to the feature vectors corresponding to the positive samples. The calculation formula for the center point of the positive samples is as follows:
[0141]
[0142] where l is the number of iterations, and the center point in the initial iteration is denoted as c1, and f i is the feature vector corresponding to the positive sample. Calculating the center point according to the feature vector corresponding to the positive sample realizes clustering the positive samples around the center point.
[0143] 3-2) Calculate the corrected center point according to the position of the center point of the positive samples. The corrected center point is calculated by the following formula:
[0144] c b = βc l-1 + (1 - β)c l ;
[0145] where c b is the corrected center point, and β takes a value of 0.98; calculating the corrected center point by the method of exponential weighted average can realize that the center point can change violently in the early stage of model training, accelerating the convergence of the model.
[0146] 3-3) Update the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples, the corrected center point, and the positive sample loss function, specifically including:
[0147] Calculate the Euclidean distance from the feature vector corresponding to each positive sample to the corrected center point, denoted as dis i ;
[0148] Adopt the gradient descent method to calculate and update the parameters of the deep learning network model according to the Euclidean distance and the positive sample loss function;
[0149] The positive sample loss function is denoted as: where dis i is the Euclidean distance from the positive sample to the corrected center point, and e is the maximum value of the distance from the positive sample to the corrected center point. Through the iteration of the positive sample loss function, all positive samples are restricted within the hypersphere with a radius of e.
[0150] 4) Input the negative samples and positive samples into the updated deep learning network model, and calculate the feature vectors corresponding to the negative samples and the feature vectors corresponding to the positive samples.
[0151] 5) Update the parameters of the deep learning network model according to the feature vectors corresponding to the negative samples, the feature vectors corresponding to the positive samples, the corrected center point, and the negative sample loss function, specifically including:
[0152] Calculate the Euclidean distance from the feature vector corresponding to each negative sample to the corrected center point and the Euclidean distance from the feature vector corresponding to each positive sample to the corrected center point
[0153] Adopt the gradient descent method to calculate and update the parameters of the deep learning network model according to the above Euclidean distance and the negative sample loss function;
[0154] The negative sample loss function is denoted as: where is the distance from the negative sample closest to the corrected center point to the corrected center point, and g is the interval between the distance from the positive sample to the corrected center point and the distance from the negative sample closest to the corrected center point to the corrected center point. Through the iteration of the negative sample loss function, all negative samples are restricted outside the sphere region with the sum of the radius of the hypersphere formed by the positive samples and g as the radius, that is, all negative samples are restricted outside the sphere region with a radius of e + g. In the iteration process, the hard sample mining method is adopted, and the distance from the negative sample closest to the corrected center point to the corrected center point is used to iterate the loss function to improve the effect of feature extraction.
[0155] , 6) Repeat the above steps for a preset number of times, or until the change in the loss function of the negative samples meets the preset conditions, then end the iteration to obtain the trained deep learning network model.
[0156] The deep learning network model is trained using positive and negative samples. The training process is one where positive samples are pulled in and negative samples are pushed away. After training, a hypersphere is formed. Inside the hypersphere are positive samples, and outside are negative samples.
[0157] In one embodiment, when the number of loops reaches 200 times, end the training process and record the final center point and the parameters of the deep learning network model. Or when the loss function of the negative samples does not decrease for 5 consecutive rounds, also end the training process and record the final center point and the parameters of the deep learning network model.
[0158] During the above model training process, the parameter e represents the radius of the hypersphere where positive samples gather, and g represents the minimum distance between positive and negative samples. For business flexibility, during the training process, it can be adjusted according to the training results. For example, if the lighting is abnormal in some scenarios, it may cause some positive samples to exceed the range of the hypersphere. At this time, the threshold can be appropriately adjusted to adapt to this scenario, such as increasing the value of e. That is to say, e and g are used to constrain features during the training process.
[0159] In this embodiment, the hypersphere for judgment includes a center point and a radius;
[0160] The center point of the hypersphere is the position of the last corrected center point during the training process of the deep learning network model;
[0161] The radius of the hypersphere is determined according to the distances between the positive samples and the negative samples respectively and the center point of the hypersphere. In this embodiment, the radius of the hypersphere is confirmed according to the longest distance L2 between the positive samples and the center point of the hypersphere and the shortest distance L1 between the negative samples and the center point of the hypersphere. Denote the radius of the hypersphere as thr, then e = thr = min(L1, L2). The radius can be adjusted according to business requirements to increase or decrease the threshold.
[0162] Step 4: Calculate the relationship between the feature vector corresponding to the second image and the hypersphere, and determine whether the luggage image in the second image includes a violation luggage image. In this embodiment, the hypersphere is composed of the feature vectors corresponding to the samples including violation luggage images.
[0163] Calculating the relationship between the feature vector corresponding to the second image and the hypersphere is essentially calculating the distance from the feature vector corresponding to the second image to the center point of the trained hypersphere. This distance can be calculated using the Euclidean distance calculation method, which will not be elaborated here.
[0164] When the distance is less than or equal to the hypersphere radius, it indicates that the luggage image in the second image is a compliant luggage. When the distance is greater than the hypersphere radius, it indicates that the luggage image in the second image is a non-compliant luggage. In the case of detecting a non-compliant luggage, an alarm is given. While giving the alarm, the surveillance image is recorded as evidence, which is also convenient for optimizing the deep learning model subsequently. In the case of not detecting a non-compliant luggage, the next round of detection is carried out.
[0165] Embodiment 2
[0166] As Figure 4 shown is a flowchart of a method for detecting non-compliant luggage provided by another embodiment of the present invention. As Figure 4 shown, the method includes:
[0167] Step 1: Preprocess the acquired surveillance image to obtain a first image including at least one moving target;
[0168] Step 2: Extract a second image including a luggage image from the first image through target detection;
[0169] Step 3: Input the second image into a deep learning network model to obtain a feature vector corresponding to the second image;
[0170] Step 4: Calculate the relationship between the feature vector corresponding to the second image and the hypersphere to determine whether the luggage image in the second image includes a non-compliant luggage image;
[0171] Step 5: Statistically calculate the ratio of frames with non-compliant luggage images in multiple frames of surveillance images within a preset time. In the case where the ratio exceeds a preset ratio, it is determined that the surveillance image includes a non-compliant luggage image.
[0172] In this embodiment, by statistically calculating the detection results in multiple frames of surveillance images within a preset time, suppression is carried out according to the statistical results to reduce the interference of special situations and special angles and enhance the accuracy of the detection results. The frames with the detection result of a non-compliant luggage in a single frame image are marked as 1, and the frames with the detection result of a compliant luggage are marked as 0. Calculate that within a preset time, if more than a preset ratio of frames are marked as 1, an alarm is issued to prompt the administrator that there is someone carrying a non-compliant luggage in the warehouse, and the pictures of this time period are recorded. In some embodiments, the preset time is 2 - 5 s, and the preset ratio is 50% - 80%.
[0173] Embodiment 3
[0174] Figure 5 is a block diagram of a system for detecting non-compliant luggage provided by an embodiment of the present invention. As Figure 5 shown, the system includes:
[0175] An image preprocessing module, configured to preprocess the acquired surveillance image to obtain a first image including at least one moving target;
[0176] A target detection module, configured to extract a second image including a luggage image from the first image through target detection;
[0177] A feature extraction module, configured to input the second image into a deep learning network model to obtain a feature vector corresponding to the second image;
[0178] A comparison and judgment module, configured to calculate the relationship between the feature vector corresponding to the second image and a hypersphere, and judge whether the luggage image in the second image includes a non-compliant luggage image;
[0179] The hypersphere is composed of feature vectors corresponding to samples including non-compliant luggage images. This system intercepts the second image of the luggage from the surveillance image through target detection, and the trained deep learning network detects the second image to judge whether the luggage in the second image is a non-compliant luggage, and alarms in time when a non-compliant luggage is detected, reducing the workload of administrators or security personnel.
[0180] In this embodiment, the deep learning network model is trained through the following steps:
[0181] 1) Obtain training sample data, where the training sample data includes positive samples and negative samples. The positive samples are compliant luggage images, and the negative samples are non-compliant luggage images. Since compliant luggage has a fixed format and the sample size is small, after obtaining compliant luggage images and non-compliant luggage images, it is necessary to use data augmentation methods to obtain a balanced number of compliant luggage images and non-compliant luggage images.
[0182] In some embodiments, one or more of occlusion, rotation, size normalization, or brightness change are used to augment the compliant luggage images. Occlusion is to occlude the target with randomly colored, sized, and shaped color blocks, with a maximum of no more than one-third of the luggage area. This augmentation also provides sample data for possible occlusions during the handling of compliant luggage. Rotation refers to randomly rotating the angle. Size normalization and brightness change are both conventional means of image processing and will not be elaborated here. In some embodiments, one or more of the above methods are randomly rotated for augmentation. For example, image A may undergo occlusion, size normalization, and rotation, while image B may undergo brightness change and occlusion. To meet the data volume requirements, a compliant luggage image is usually augmented to 4-6 images.
[0183] For the images of illegal luggage, due to the large variety and styles of the luggage itself, the data volume is large. If the data volume is sufficient, data augmentation may not be necessary. In the case of insufficient data volume, augmentation can be performed randomly. By performing data augmentation, more positive samples can be obtained to solve the problem of fewer positive sample images of compliant luggage with fixed formats and the possible occlusion effects.
[0184] 2) Input the positive samples into the deep learning network model after the previous round of iteration, and calculate the feature vectors corresponding to the positive samples. In this embodiment, when n samples are input, n m-dimensional vectors f are obtained.
[0185] The final output value of a traditional model is generally a feature map. For example, for an image with a size of W x H, after passing through the convolutional layer, it becomes W x H x C, where C is the number of channels. After passing through the pooling layer, it becomes W / 2 x H / 2 x C. Conventional deep learning networks are all composed of convolutional layers, pooling layers, or their variants. Assuming that the final output result of a conventional deep learning network is a feature map of M x N x C, the global average pooling (GAP) is used for the last layer, that is, each M x N map becomes the average value of all the data in the feature map, and the final output result is 1 x C. Then, through a fully connected layer, a 1-dimensional feature map is finally obtained. As Figure 4 shown, a 3x3 feature map passes through global average pooling (i.e., the average value of 9 points) to obtain an average value as the final output. Then, the final output of a 3x3x4 feature map is a 1x4 vector.
[0186] 3) Update the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples and the positive sample loss function, specifically including:
[0187] 3-1) Calculate the center point of the positive sample according to the feature vector corresponding to the positive sample. The calculation formula for the center point of the positive sample is as follows:
[0188]
[0189] where l is the number of iterations, and the center point at the initial iteration is denoted as c1, and f i is the feature vector corresponding to the positive sample. Calculating the center point according to the feature vector corresponding to the positive sample realizes clustering the positive samples around the center point.
[0190] 3-2) Calculate the corrected center point according to the position of the center point of the positive sample. The corrected center point is calculated by the following formula:
[0191] c b = βc l-1 + (1 - β)c l ;
[0192] Among them, c b is the correction center point, and β takes a value of 0.98. The correction center point is calculated by the method of exponential weighted average, which can achieve that the center point changes violently in the early stage of model training and accelerate model convergence.
[0193] 3-3) Update the parameters of the deep learning network model according to the feature vector corresponding to the positive sample, the correction center point, and the positive sample loss function, specifically including:
[0194] Calculate the Euclidean distance from the feature vector corresponding to each positive sample to the correction center point, denoted as dis i ;
[0195] Adopt the gradient descent method to calculate and update the parameters of the deep learning network model according to the Euclidean distance and the positive sample loss function;
[0196] The positive sample loss function is denoted as: Among them, dis i is the Euclidean distance from the positive sample to the correction center point, and e is the maximum value of the distance from the positive sample to the correction center point. Through the iteration of the positive sample loss function, all positive samples are restricted within the hypersphere with a radius of e.
[0197] 4) Input the negative sample and the positive sample into the updated deep learning network model, and calculate the feature vectors corresponding to n negative samples and the feature vector corresponding to the positive sample.
[0198] 5) Update the parameters of the deep learning network model according to the feature vector corresponding to the negative sample, the feature vector corresponding to the positive sample, the correction center point, and the negative sample loss function, specifically including:
[0199] Calculate the Euclidean distance from the feature vector corresponding to each negative sample to the correction center point and the Euclidean distance from the feature vector corresponding to each positive sample to the correction center point
[0200] Adopt the gradient descent method to calculate and update the parameters of the deep learning network model according to the above Euclidean distance and the negative sample loss function;
[0201] The negative sample loss function is denoted as: Among them, Let \(d\) be the distance from the negative sample closest to the center point of correction to the center point of correction, and \(g\) be the interval between the distance from the positive sample to the center point of correction and the distance from the negative sample closest to the center point of correction to the center point of correction. Through the iteration of the negative sample loss function, all negative samples are restricted outside the sphere region with the sum of the radius of the hypersphere formed by positive samples and \(g\) as the radius, that is, all negative samples are restricted outside the sphere region with a radius of \(e + g\). During the iteration process, the hard sample mining method is adopted, and the distance from the negative sample closest to the center point of correction to the center point of correction is used to iterate the loss function to improve the effect of feature extraction.
[0202] 6) Repeat the above steps for a preset number of times, or when the change of the negative sample loss function meets the preset conditions, end the iteration to obtain the trained deep learning network model.
[0203] The deep learning network model is trained with positive and negative samples. The training process is a process in which positive samples are pulled in and negative samples are pushed away. After training, a hypersphere is formed. Inside the hypersphere are positive samples, and outside are negative samples.
[0204] In one embodiment, when the number of loops reaches 200 times, end the training process and record the final center point and the parameters of the deep learning network model. Or when the negative sample loss function does not decrease for 5 consecutive rounds, also end the training process and record the final center point and the parameters of the deep learning network model.
[0205] In this embodiment, the hypersphere used for judgment includes the center of the sphere and the radius;
[0206] The center of the sphere of the hypersphere is the position of the center point of correction in the last correction during the training process of the deep learning network model;
[0207] The radius of the hypersphere is determined according to the distances between the positive samples and the negative samples and the center of the sphere of the hypersphere respectively. In this embodiment, the radius of the hypersphere is confirmed according to the longest distance \(L2\) between the positive sample and the center of the sphere of the hypersphere and the shortest distance \(L1\) between the negative sample and the center of the sphere of the hypersphere. Denote the radius of the hypersphere as \(thr\), then \(e = thr = min(L1, L2)\). The radius can be adjusted according to business requirements to increase or decrease the threshold.
[0208] In this embodiment, the image preprocessing module includes:
[0209] A monitoring image acquisition module, configured to acquire consecutive first-frame images, second-frame images, and third-frame images composed of pixels of any channel in the monitoring image;
[0210] A difference calculation module, configured to calculate the difference between the first-frame image and the second-frame image to obtain a first difference image, and calculate the difference between the second-frame image and the third-frame image to obtain a second difference image;
[0211] A binarization module, configured to binarize the first difference image and the second difference image to obtain a first binarized image and a second binarized image;
[0212] A bitwise AND calculation module, configured to perform a bitwise AND operation on the first binarized image and the second binarized image to obtain a mask image of the moving target;
[0213] A connected component calculation module, configured to calculate the connected components of the mask image;
[0214] A connected component filtering module, configured to filter out the connected components smaller than a preset size in the mask image to obtain the connected components of the moving target;
[0215] An enclosing quadrilateral calculation module, configured to calculate the enclosing quadrilateral of the connected components of the moving target;
[0216] An enlargement module, configured to enlarge the enclosing quadrilateral by a preset ratio, and obtain the monitoring image area corresponding to the enlarged enclosing quadrilateral as the first image of the moving target. Each sub-module in the image preprocessing module can cooperate with each other to quickly determine the existence of a moving target in the monitoring image, and can quickly locate and determine a luggage target that occupies a small area in the moving target. On the one hand, it reduces the interference of stationary objects similar to luggage in the warehouse. On the other hand, since carrying luggage is a low-probability event, preprocessing can effectively eliminate images that do not require luggage detection, saving computing resources.
[0217] Embodiment 4
[0218] Figure 6 It is a block diagram of a detection system for illegal luggage provided by an embodiment of the present invention. As Figure 6 shown, the system includes:
[0219] An image preprocessing module, configured to preprocess the acquired monitoring image to obtain a first image including at least one moving target;
[0220] A target detection module, configured to extract a second image including a luggage image from the first image through target detection;
[0221] A feature extraction module, configured to input the second image into a deep learning network model to obtain a feature vector corresponding to the second image;
[0222] A comparison and judgment module, configured to calculate the relationship between the feature vector corresponding to the second image and the hypersphere, and judge whether the luggage image in the second image includes an illegal luggage image;
[0223] The hypersphere is composed of the feature vectors corresponding to the samples including illegal luggage images;
[0224] A result suppression module is configured to count the ratio of the number of monitoring images with illegal luggage images within a preset time. When the ratio exceeds a preset ratio, it is determined that the monitoring images include illegal luggage images. The system extracts a second image of the luggage from the monitoring images through object detection, and a trained deep learning network detects the second image to determine whether the luggage in the second image is an illegal luggage, and alarms in time when an illegal luggage is detected, reducing the workload of administrators or security personnel.
[0225] In this embodiment, by counting the detection results in multiple frames of monitoring images within a preset time and suppressing according to the statistical results, the interference of special situations and special angles is reduced, and the accuracy of the detection results is enhanced.
[0226] The present invention further provides an electronic device, including:
[0227] A memory for storing a computer program;
[0228] A processor for running the computer program to execute the method for detecting illegal luggage.
[0229] The present invention further provides a machine-readable storage medium, on which instructions are stored, and the instructions are used to enable the machine to execute the method for detecting illegal luggage.
[0230] On the other hand, the present invention provides a computer program product, including a computer program, characterized in that the computer program implements the method for detecting illegal luggage when executed by a processor.
[0231] Through the above technical solutions, the present invention can utilize the monitoring images collected by existing monitoring devices, extract a second image of the luggage from the monitoring images through object detection, use a trained deep learning network to detect the second image, determine whether the luggage in the second image is an illegal luggage, count the detection results in multiple frames of monitoring images within a preset time, and suppress according to the statistical results, reducing the interference of special situations and special angles, and enhancing the accuracy of the detection results. It realizes automatic detection of whether there is an illegal luggage carried, reduces the workload of administrators or security personnel, and ensures the safety of the vault.
[0232] Those skilled in the art can understand that all or part of the steps in the methods of the above-described embodiments can be completed by instructing relevant hardware through a program, and this program is stored in a storage medium, including several instructions to enable a single-chip microcomputer, a chip, or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0233] The optional embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the technical concept scope of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention. Additionally, it should be noted that, among the various specific technical features described in the above specific embodiments, they can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not separately describe various possible combination methods.
[0234] In addition, any combination can be made between various different embodiments of the present invention as long as it does not violate the idea of the embodiments of the present invention, and it should also be regarded as the content disclosed in the embodiments of the present invention.
Claims
1. A detection method for illegal luggage, characterized in that, The method includes: Preprocessing the acquired monitoring image to obtain a first image including at least one moving target; Extracting a second image including a luggage image from the first image through target detection; Inputting the second image into a deep learning network model to obtain a feature vector corresponding to the second image; Calculating the relationship between the feature vector corresponding to the second image and a hypersphere, and determining whether the luggage image in the second image includes a non-compliant luggage image; Wherein, the hypersphere is composed of the feature vectors corresponding to the samples including non-compliant luggage images; The deep learning network model is trained through the following steps: Obtaining training sample data, where the training sample data includes positive samples and negative samples; Inputting the positive samples into the deep learning network model after the previous round of iteration, and calculating the feature vectors corresponding to the positive samples; Updating the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples and the positive sample loss function, including: Calculating the center point of the positive samples according to the feature vectors corresponding to the positive samples; Calculating a corrected center point according to the position of the center point of the positive samples; Updating the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples, the corrected center point, and the positive sample loss function; Inputting the negative samples and the positive samples into the updated deep learning network model, and calculating the feature vectors corresponding to the negative samples and the feature vectors corresponding to the positive samples; Updating the parameters of the deep learning network model according to the feature vectors corresponding to the negative samples, the feature vectors corresponding to the positive samples, the corrected center point, and the negative sample loss function; Repeating the above steps a preset number of times, or ending the iteration when the change of the negative sample loss function meets the preset conditions, to obtain a trained deep learning network model.
2. The detection method according to claim 1, wherein The hypersphere includes a center and a radius; The center of the hypersphere is the position of the last corrected center point during the training process of the deep learning network model; The radius of the hypersphere is determined according to the distances between the positive samples and the negative samples and the center of the hypersphere respectively; 3. The detection method according to claim 1, wherein The calculation formula for the center point of the positive samples is as follows: where l is the number of iterations, the center point of the first iteration is denoted as c1, and f i is the feature vector corresponding to the positive sample; The corrected center point is calculated through the following formula: c b = βc l-1 + (1 - β)c l ; Among them, c b is the correction center point, and the value of β is 0.
98.
4. The detection method according to claim 1, wherein Updating the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples, the position of the corrected center point, and the positive sample loss function, including: Calculate the Euclidean distance from the feature vector corresponding to each positive sample to the corrected center point, denoted as dis i ; Adopting the gradient descent method, calculating and updating the parameters of the deep learning network model according to the Euclidean distance and the positive sample loss function; The positive sample loss function is denoted as: where dis i is the Euclidean distance from the positive sample to the corrected center point, and e is the maximum value of the distance from the positive sample to the corrected center point.
5. The detection method according to claim 1, characterized in that, Updating the parameters of the deep learning network model according to the feature vectors corresponding to the negative samples, the feature vectors corresponding to the positive samples, the corrected center point, and the negative sample loss function, including: Calculate the Euclidean distance from the feature vector corresponding to each negative sample to the corrected center point and the Euclidean distance from the feature vector corresponding to each positive sample to the corrected center point Adopting the gradient descent method, calculating and updating the parameters of the deep learning network model according to the above Euclidean distance and the negative sample loss function; The negative sample loss function is denoted as: where is the distance from the negative sample closest to the correction center point to the correction center point, and g is the interval between the distance from the positive sample to the correction center point and the distance from the negative sample closest to the correction center point to the correction center point.
6. The detection method according to claim 1, wherein The positive samples are compliant luggage images, and the negative samples are non-compliant luggage images; Obtaining the training sample data includes: Obtaining compliant luggage images and non-compliant luggage images; Enhance the compliant luggage images and non-compliant luggage images by one or more of occlusion, rotation, size normalization, and brightness variation, so that the number of compliant luggage images is balanced with the number of non-compliant luggage images.
7. The detection method according to claim 1, characterized in that, The preprocessing of the acquired images to obtain a first image including at least one moving object includes: Obtain consecutive first-frame, second-frame, and third-frame images composed of pixels in any channel of the surveillance image; Subtract the first-frame image from the second-frame image to obtain a first difference image, and subtract the second-frame image from the third-frame image to obtain a second difference image; Binarize the first difference image and the second difference image to obtain a first binarized image and a second binarized image; Perform a bitwise AND operation on the first binarized image and the second binarized image to obtain a mask image of the moving object; Calculate the connected components of the mask image; Filter out the connected components smaller than a preset size in the mask image to obtain the connected components of the moving object; Calculate the circumscribed quadrilateral of the connected components of the moving object; Enlarge the circumscribed quadrilateral by a preset ratio, and obtain the surveillance image area corresponding to the enlarged circumscribed quadrilateral as the first image of the moving object.
8. The detection method according to claim 1, wherein After determining whether the luggage images in the second image include non-compliant luggage images, the method further includes: Statistically calculate the ratio of non-compliant luggage images in multiple frames of surveillance images within a preset time. If the ratio exceeds the preset ratio, it is determined that the surveillance images include non-compliant luggage images.
9. A detection system for illegal luggage, characterized in that, The system includes: An image preprocessing module for preprocessing the acquired surveillance images to obtain a first image including at least one moving object; A target detection module for extracting a second image including luggage images from the first image through target detection; A feature extraction module for inputting the second image into a deep learning network model to obtain a feature vector corresponding to the second image; A comparison and judgment module for calculating the relationship between the feature vector corresponding to the second image and a hypersphere, and determining whether the luggage images in the second image include non-compliant luggage images; Wherein, the hypersphere is composed of the feature vectors corresponding to the samples including non-compliant luggage images; The deep learning network model is trained through the following steps: Obtain training sample data, where the training sample data includes positive samples and negative samples; Input the positive samples into the deep learning network model after the previous round of iteration, and calculate the feature vectors corresponding to the positive samples; Update the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples and the positive sample loss function, including: Calculate the center point of the positive samples according to the feature vectors corresponding to the positive samples; Calculate the corrected center point according to the position of the center point of the positive samples; Update the parameters of the deep learning network model according to the feature vectors corresponding to the positive samples, the corrected center point, and the positive sample loss function; Input the negative samples and the positive samples into the updated deep learning network model, and calculate the feature vectors corresponding to the negative samples and the feature vectors corresponding to the positive samples; Update the parameters of the deep learning network model according to the feature vector corresponding to the negative sample, the feature vector corresponding to the positive sample, the corrected center point, and the negative sample loss function; Repeat the above steps for a preset number of times, or when the change of the negative sample loss function meets the preset conditions, end the iteration to obtain a trained deep learning network model.
10. The detection system according to claim 9, characterized in that, The hypersphere includes a center and a radius; The center of the hypersphere is the position of the last corrected center point during the training process of the deep learning network model; The radius of the hypersphere is determined according to the distances between the positive samples and the negative samples and the center of the hypersphere respectively.
11. The detection system according to claim 9, characterized in that, The calculation formula for the center point of the positive sample is as follows: where l is the number of iterations, the center point of the first iteration is denoted as c1, and f i is the feature vector corresponding to the positive sample; The corrected center point is calculated by the following formula: c b = βc l-1 + (1 - β)c l ; Among them, c b is the correction center point, and the value of β is 0.
98.
12. The detection system according to claim 9, wherein The updating of the parameters of the deep learning network model according to the feature vector corresponding to the positive sample, the position of the corrected center point, and the positive sample loss function includes: Calculate the Euclidean distance from the feature vector corresponding to each positive sample to the corrected center point, denoted as dis i ; Using the gradient descent method, calculate and update the parameters of the deep learning network model according to the Euclidean distance and the positive sample loss function; The positive sample loss function is denoted as: where dis i is the Euclidean distance from the positive sample to the corrected center point, and e is the maximum value of the distance from the positive sample to the corrected center point.
13. The detection system according to claim 9, characterized in that, The updating of the parameters of the deep learning network model according to the feature vector corresponding to the negative sample, the feature vector corresponding to the positive sample, the corrected center point, and the negative sample loss function includes: Calculate the Euclidean distances from the feature vectors corresponding to each negative sample to the corrected center point and the Euclidean distances from the feature vectors corresponding to each positive sample to the corrected center point Using the gradient descent method, calculate and update the parameters of the deep learning network model according to the above Euclidean distance and the negative sample loss function; The negative sample loss function is denoted as: where is the distance from the negative sample closest to the correction center point to the correction center point, and g is the interval between the distance from the positive sample to the correction center point and the distance from the negative sample closest to the correction center point to the correction center point.
14. The detection system according to claim 9, wherein The image preprocessing module includes: A monitoring image acquisition module for acquiring consecutive first, second, and third frame images composed of pixels in any channel of the monitoring image; A difference calculation module for calculating the difference between the first frame image and the second frame image to obtain a first difference image, and calculating the difference between the second frame image and the third frame image to obtain a second difference image; A binarization module for binarizing the first difference image and the second difference image to obtain a first binarized image and a second binarized image; A bitwise AND calculation module for performing a bitwise AND operation on the first binarized image and the second binarized image to obtain a mask image of the moving target; A connected component calculation module for calculating the connected components of the mask image; A connected component filtering module for filtering the connected components smaller than a preset size in the mask image to obtain the connected components of the moving target; An enclosing quadrilateral calculation module for calculating the enclosing quadrilateral of the connected components of the moving target; An enlargement module for enlarging the enclosing quadrilateral by a preset ratio and acquiring the monitoring image area corresponding to the enlarged enclosing quadrilateral as the first image of the moving target.
15. The detection system according to claim 9, wherein The system further includes: A result suppression module for counting the ratio of the number of monitoring images containing illegal luggage images within a preset time, and determining that the monitoring image includes an illegal luggage image when the ratio exceeds the preset ratio.
16. An electronic device, characterized in that, including: A memory for storing a computer program; A processor for running the computer program to execute the detection method for illegal luggage according to any one of claims 1-8.
17. A machine-readable storage medium having instructions stored thereon for causing a machine to execute the detection method for illegal luggage according to any one of claims 1-8.
18. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method for detecting a non-compliant luggage case according to any one of claims 1-8.
Citation Information
Patent Citations
Target object recognition method and device, storage medium and electronic device
CN112633297A