A method and system for detecting leftovers in an industrial scenario
The twin neural network-based method with ResNet50 backbone and negative correlation operator effectively addresses high false and missed detection rates in industrial abandoned object detection by enhancing feature extraction and classification accuracy.
Patent Information
- Application Number
- CN202210532287.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-05-09
AI Technical Summary
The existing residual detection methods in industrial scenarios have high misjudgment and missed detection rates under light changes, motion occlusion and complex backgrounds, and lack data sets that meet industrial needs, making them difficult to apply to actual scenarios.
A twin neural network is used to extract the current frame and background frame feature maps of the monitoring video, and a negative correlation operator is used to calculate the correlation, combining the target detection network and the classification network to reduce background interference and improve detection accuracy.
It effectively reduces the misjudgment rate and missed detection rate of legacy detection, meets the real-time detection requirements in industrial scenarios, and adapts to complex backgrounds and light changes.
Smart Images

Figure CN114926764B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent detection of industrial monitoring videos, and more specifically, relates to a method and system for detecting left objects in an industrial scenario. Background Art
[0002] The detection of left objects is an important research direction in the field of intelligent monitoring videos. Especially in industrial production areas, the environmental safety requirements are extremely strict. The presence of unknown objects may affect the operation of the unit, pose potential safety hazards, and thus cause significant losses. This algorithm can detect the objects left by personnel in the industrial production area and send alarm information to the monitoring center in a timely manner to ensure production safety.
[0003] In recent years, many researchers have proposed some algorithms for detecting left objects. These algorithms perform well on public datasets, but mostly adopt traditional image processing methods. However, in the actual environment, there are challenges such as light changes, motion occlusion, target scale changes, and background interference, which make it difficult to apply the current methods to actual scenarios. There are still some problems in the detection of left objects in industrial scenarios, mainly reflected in the following points:
[0004] 1. Existing research solutions lack a left object dataset that meets industrial requirements. On the one hand, the scenarios of existing public datasets are mostly airports and railway station platforms, with a single scene style and simple background, and do not define and distinguish various left behaviors of personnel, and cannot detect and evaluate the relatively complex and diverse left behaviors of personnel in industrial environments. On the other hand, the objects in existing left object datasets are mostly suitcases, backpacks, etc., while in industrial environments, personnel often leave sundries, tools, etc. The large difference in the categories of left objects makes it impossible for deep learning methods to play their detection advantages; in addition, although there are many datasets in the field of object detection, the objects in them are large objects in daily scenarios, while the objects in monitoring videos are mostly small and medium-sized objects. The target types and scale sizes of existing datasets do not meet the industrial detection requirements.
[0005] 2. Existing research solutions have high omission rates and misjudgment rates. Existing algorithms mainly adopt traditional image processing methods and are very likely to misjudge stationary targets such as cars and people as left objects. Moreover, the frequent movement and stay of workers in the factory, limb swings, motion occlusion, and similar background interference will all cause serious omissions.
[0006] 3. Existing research solutions have poor robustness. Existing algorithms rely on artificially designed feature extractors and often work only in specific backgrounds. When the background changes, they fail. Especially in the case of a relatively high complexity of the industrial background, traditional methods will cause great interference, resulting in unstable detection. Summary of the Invention
[0007] In view of the deficiencies and improvement requirements of the prior art, the present invention provides a method and system for detecting leftovers in an industrial scenario, aiming to reduce the false positive rate and missed detection rate of leftover detection in the industrial scenario.
[0008] To achieve the above object, according to the first aspect of the present invention, there is provided a method for detecting leftovers in an industrial scenario, the method comprising:
[0009] Extract the feature map of the current frame of the surveillance video and the feature map of the background frame respectively. The background frame corresponds to an industrial surveillance camera one by one, and is the industrial production area monitored by the camera without any leftovers. Once there is a change, it needs to be recalibrated;
[0010] Calculate the correlation between the feature map of the current frame and the feature map of the background frame using a negatively correlated correlation operator to obtain a correlation feature map;
[0011] Detect the regression boxes and positions of all foreground objects in the current frame that do not belong to the background from the correlation feature map;
[0012] Classify each foreground object to obtain the category of the foreground object;
[0013] According to the category and position of the foreground object in the current frame, combined with the characteristics of the industrial scenario, output the leftover detection result.
[0014] Preferably, the calculating the correlation between the feature map of the current frame and the feature map of the background frame using a negatively correlated correlation operator to obtain a correlation feature is specifically as follows:
[0015]
[0016] where F(a, b) is the correlation feature map of the current frame and the background frame, a is the feature map of the current frame, b is the feature map of the background frame, a ij is the feature of the i-th row and j-th column of the feature map of the previous frame, b ij is the feature of the i-th row and j-th column of the feature map of the background frame, i = 1,..., m, j = 1,..., n, and m and n are the length and width of the feature map respectively.
[0017] Advantageous effects: Aiming at the problem that the difference between the depth features of different leftovers and the depth features of the background is large, the present invention directly uses the feature map difference and then passes through the softmax function as the new feature map of the current frame, so that the interpolation of different leftovers is in the same range and can be treated with the same priority in detection.
[0018] Preferably, a siamese neural network is used to extract the feature map of the current frame of the surveillance video and the feature map of the background frame respectively. The two branch network structures of the siamese neural network are the same, and the network backbone structure is composed of the first four-layer network structure of ResNet50 in a cascade structure, and the parameters of the two branch networks are different.
[0019] Beneficial effects: By adopting the feature extraction network with the above preferred structure in the present invention, the cascaded features of the background frame and the current frame can be effectively extracted, and the relevant features composed of the cascaded features of the two networks can well represent the foreground objects in the current frame that do not belong to the background frame; in terms of computing speed, the current structure can meet the requirements of real-time detection when running on a GPU.
[0020] Preferably, a target detection network is used to detect the regression boxes and positions of all foreground objects in the current frame that do not belong to the background from the relevant feature maps, and the target detection network is a cascaded RPN head and ROI head.
[0021] Beneficial effects: By adopting the target detection network with the above preferred structure in the present invention, this structure can effectively extract the foreground objects that do not belong to the background frame, and both the missed detection rate and the misjudgment rate meet the usage requirements in industrial scenarios; in terms of computing speed, the current structure can meet the requirements of real-time detection when running on a GPU.
[0022] Preferably, if the same item foreground object is detected in consecutive M frames and no person is detected, it is determined that the foreground object is a left-behind object, and it is marked and an alarm message is sent.
[0023] Preferably, for each newly emerged item foreground object, the person closest to it in distance is matched; if the person matched by the item foreground object disappears for more than a certain time, it is marked as a left-behind object, and it is marked and an alarm message is sent.
[0024] Preferably, the number of frames of its stay is counted and converted into the stay time according to the video frame rate, the position and stay time of the left-behind object are marked on the screen, and an alarm message is pushed to the monitoring center.
[0025] To achieve the above object, according to the second aspect of the present invention, a left-behind object detection system in an industrial scenario is provided, including: a computer-readable storage medium and a processor;
[0026] The computer-readable storage medium is used to store executable instructions;
[0027] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the left-behind object detection method in the industrial scenario described in the first aspect.
[0028] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0029] Aiming at the problems of high false positive rate and missed detection rate in the detection of leftovers in existing industrial scenarios, the present invention utilizes the characteristic that the background of the monitoring video in the industrial scenario changes little, and proposes to construct a feature map that eliminates the influence of the background. However, in the industrial scenario, the background is affected by factors such as light changes, camera jitter, and noise, and the background changes greatly at the pixel level. The traditional background difference method cannot be directly used. The present invention uses a negative correlation operator to calculate the correlation features between the current frame feature map and the background frame feature map, which can accurately identify all foreground targets outside the background, eliminate the interference effects such as light changes, and reduce the false positive rate and missed detection rate of leftover detection. Brief Description of the Drawings
[0030] Figure 1 It is a flowchart of a leftover detection method in an industrial scenario provided by the present invention.
[0031] Figure 2 It is a schematic diagram of constructing a new feature map provided by the present invention.
[0032] Figure 3 It is a schematic diagram of ten target categories of the AOD dataset.
[0033] Figure 4 It is a comparison of the detection effects of three algorithms with leftovers. (a) is Algorithm 1, (b) is Algorithm 2, and (c) is the present invention.
[0034] Figure 5 It is a comparison of the detection effects of three algorithms without leftovers. (a) is Algorithm 1, (b) is Algorithm 2, and (c) is the present invention. Detailed Embodiments
[0035] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0036] As Figure 1 shown, the present invention discloses a leftover detection method in an industrial scenario, including the following steps:
[0037] Extract the feature map of the current frame and the feature map of the background frame of the monitoring video respectively. The background frame corresponds to an industrial monitoring camera one by one, and is the industrial production area monitored by the camera without any leftovers. Once there is a change, it needs to be recalibrated.
[0038] As Figure 2As shown in the figure, a Siamese neural network is used to extract the feature map of the current frame and the feature map of the background frame of the surveillance video respectively. The two branch network structures of the Siamese neural network are the same, and the backbone network structure is composed of the first four network structures of ResNet50 to form a cascaded structure. The parameters of the two branch networks are different and are respectively used to extract the cascaded features of the current frame and the background frame. The weights are trained in the legacy object dataset in the industrial scenario and are respectively used to extract the cascaded features of the current frame and the background frame.
[0039] The correlation between the feature map of the current frame and the feature map of the background frame is calculated using a negatively correlated correlation operator to obtain a correlation feature map.
[0040] The calculation of the correlation between the feature map of the current frame and the feature map of the background frame using a negatively correlated correlation operator to obtain the correlation features is specifically as follows:
[0041]
[0042] Among them, F(a, b) is the correlation feature map of the current frame and the background frame, a is the feature map of the current frame, and b is the feature map of the background frame. is the chessboard distance between the feature map of the current frame and the feature map of the background, and a ij is the feature at the i-th row and j-th column of the feature map of the previous frame, and b ij is the feature at the i-th row and j-th column of the feature map of the background frame, where i = 1, …, m and j = 1, …, n, and m and n are the length and width of the feature map respectively.
[0043] The above is the preferred operator, and it can also be the following operator:
[0044]
[0045]
[0046] The regression boxes and positions of all foreground objects in the current frame that do not belong to the background are detected from the correlation feature map.
[0047] The regression boxes and positions of all foreground objects in the current frame that do not belong to the background are detected from the correlation feature map using an object detection network. The object detection network is a cascaded RPN head and ROI head. Among them, the input of the RPN head is a large number of correlation features in the previous step, and the output is proposal regions. The input of the ROI head is proposal regions, and the output is regions of interest, that is, the regression boxes of the objects.
[0048] The present invention adopts an end-to-end training method. The dataset is the historical data of one year monitored by industrial monitoring cameras. 80% of the data is used as training data and 20% is used as test data. The positions and categories of the leftovers in each video frame are labeled as tags. In this embodiment, there are 10 categories, including "person".
[0049] Classify each foreground object to obtain the category of the foreground object.
[0050] In this embodiment, the backbone structure of the ResNet18 network is adopted for the image classification algorithm. The output layer is adjusted according to the industrial scenario leftover dataset, and the weights in the network are trained using the industrial scenario leftover dataset. The classification network classifies the objects in the region of interest obtained in the previous step.
[0051] According to the category and position of the foreground object in the current frame, combined with the characteristics of the industrial scenario, the leftover detection result is output.
[0052] The first leftover detection mechanism: If the same object foreground is detected in consecutive M frames and no person is detected, determine that the foreground object is a leftover, mark it and send an alarm message.
[0053] The second leftover detection mechanism: For each newly emerged object foreground, match the person closest to it; if the person matched by the object foreground disappears for more than a certain time, mark it as a leftover, mark it and send an alarm message.
[0054] Count the number of frames it has been left and convert it to the leftover time according to the video frame rate. Mark the position and leftover time of the leftover on the screen and push an alarm message to the monitoring center.
[0055] In this embodiment, a leftover dataset (Abandoned Objects Dataset, hereinafter referred to as "AOD") in the industrial scenario is constructed. AOD is mainly used to evaluate the detection model of deep learning and is divided into a training set, a validation set, and a test set. Sampling of AOD requires placing various leftovers at different positions in the scenario. The leftover items need to comprehensively consider the items that may appear in industrial production. After multiple observations and on-site investigations, nine types of items such as toolboxes and backpacks are selected, plus ten types including the people in the factory as the objects to be labeled. The AOD targets are as follows Figure 3As shown. At the same time, it is necessary to ensure that the targets in the dataset have different angles, sizes, brightnesses, and different degrees of occlusion. A test set is set up to compare the detection accuracies of different models in the follow-up. The AOD dataset contains a total of 3,600 images and 10,029 target instances. Among them, the training set contains 1,900 images and 5,159 target instances, the validation set contains 300 images and 477 target instances, and the test set contains 1,400 images and 4,393 target instances. Due to the large number of personnel activities, the number of personnel target instances is the largest, and the number of other item target instances is balanced. The statistical information of the number of various targets in AOD is shown in Table 1. Small targets are defined as targets with widths and heights less than 10% of the width and height of the video frame. Among them, there are 3,968 small target instances, accounting for about 40%, and 8,566 medium-sized target instances, accounting for about 85%. AOD contains targets of different scales and different states. The large number of small and medium-sized targets and their indistinguishable features in industrial scenarios greatly increase the detection difficulty of the model.
[0056] Table 1 Statistical information on the number of AOD instances
[0057]
[0058] The method of the present invention is respectively compared with Algorithm 1 in "Abandoned Object Detection using Frame Differencing and Background Subtraction" and Algorithm 2 in "Application of YOLO Deep Learning Model for Real Time Abandoned Baggage Detection", and the detection accuracy of the algorithm is evaluated by the missed detection rate and false judgment rate indicators.
[0059] The missed detection rate is equal to the number of remaining objects not detected divided by the total number of actual remaining objects; the false judgment rate is equal to the number of objects in which the remaining objects are misclassified divided by the total number of detected remaining objects. The statistical results of the detection of various events are shown in Table 2.
[0060] Table 2 Comparison of different methods for detecting remaining objects
[0061] Miss detection rate False positive rate Algorithm 1 60.00% 65.46% Algorithm 2 34.82% 34.33% This method 9.63% 4.50%
[0062] From Table 2 and Figure 4 - Figure 5 it can be seen that the missed detection rate and false judgment rate of the method of the present invention for industrial monitoring videos are much lower than those of other methods.
[0063] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting leftovers in an industrial scenario, characterized in that, The method includes: Using a Siamese neural network to extract the feature map of the current frame of the surveillance video and the feature map of the background frame respectively. The background frame corresponds to an industrial surveillance camera one by one, and is the industrial production area monitored by the camera without any leftovers. Once there is a change, it needs to be recalibrated. Among them, the two branch network structures of the Siamese neural network are the same, and the network backbone structure is composed of the first four network structures of ResNet50 in a cascaded structure, and the parameters of the two branch networks are different; Using a negatively correlated correlation operator to calculate the correlation between the current frame feature map and the background frame feature map to obtain a correlation feature map; Using an object detection network to detect the regression boxes and positions of all foreground objects in the current frame that do not belong to the background from the correlation feature map. The object detection network is a cascaded RPN head and ROI head; classifying each foreground object to obtain the category of the foreground object; According to the category and position of the foreground object in the current frame, combined with the characteristics of the industrial scenario, output the detection result of the leftover; 2. The method according to claim 1, wherein The use of a negatively correlated correlation operator to calculate the correlation between the current frame feature map and the background frame feature map to obtain a correlation feature is specifically as follows: Among them, F(a, b) is the correlation feature map between the current frame and the background frame, a is the feature map of the current frame, b is the feature map of the background frame, and a ij is the feature at the i-th row and j-th column of the feature map of the previous frame, and b ij is the feature at the i-th row and j-th column of the feature map of the background frame, where i = 1, …, m, j = 1, …, n, and m and n are the length and width of the feature map respectively.
3. The method according to any one of claims 1 to 2, characterized in that If the same item foreground object is detected in M consecutive frames and no person is detected, determine that the foreground object is a leftover, mark it and send an alarm message; 4. The method according to any one of claims 1 to 2, characterized in that, For each newly appeared item foreground object, match the person closest to it in distance; if the person matched by the item foreground object disappears for more than a certain time, mark it as a leftover, mark it and send an alarm message; 5. The method according to claim 3 or 4, characterized in that, Count the number of frames it has been left and convert it to the leftover time according to the video frame rate, mark the position and leftover time of the leftover on the screen, and push an alarm message to the monitoring center; 6. A legacy detection system in an industrial scenario, characterized in that, It includes: A computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the leftover detection method in the industrial scenario according to any one of claims 1 to 5.
Citation Information
Patent Citations
Remnant detection and tracking method
CN103605983A