Personnel security detection method

The combination of RetinaNet, OpenCV, and DeepSort with Kalman filtering addresses inefficiencies in existing personnel detection systems, enhancing precision and robustness for real-time tracking in hazardous areas.

CN120279068APending Publication Date: 2025-07-08CHENGDU ANMUSEN INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510353801.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional personnel entry detection methods are inefficient and have poor reliability in complex and changing industrial environments, making it difficult to achieve all-weather-free monitoring, and are not intelligent enough to meet the high standards of modern industrial safety management.

Method used

The target detection based on Ret i naNet and opencv template matching algorithm combined with Kalman filter prediction are used to track people through the DeepSort algorithm to realize multi-scale feature fusion and real-time tracking, and optimize the tracking results using cascading matching and interleaving and matching matching strategies.

Benefits of technology

It improves the accuracy and robustness of personnel detection, realizes real-time tracking and identification of people entering dangerous areas, adapts to complex environments, and has strong versatility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279068A_ABST
    Figure CN120279068A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of personnel security and protection detection, and discloses a personnel security and protection detection method, which specifically comprises the following steps of: 1, performing target detection based on RetinaNet for personnel detection: extracting three layers of feature maps of c3, c4 and c5 by taking ResNet as a backbone network; through a deep learning-based Retinanet target detection algorithm, a personnel target in an image can be accurately detected, through combination of a DeepSort multi-target tracking algorithm and an opencv correlation coefficient template matching algorithm, the detection precision and robustness are further improved, and through combination of the multi-target tracking algorithm and the template matching algorithm, the detection precision and robustness are improved. The method can realize real-time tracking and identification of people entering a dangerous area, multi-scale feature fusion supports detection of targets of different sizes, a template matching threshold is adjustable to adapt to complex illumination and shielding scenes, covariance matrix dynamic update and feature set management enhance the generalization ability of an algorithm to a dynamic environment, and the algorithm has high robustness. The method can adapt to different industrial environments and dangerous areas, and has high universality and expandability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of personnel security detection, and specifically relates to a personnel security detection method. Background Art

[0002] In a complex and changeable industrial production environment, various dangerous areas widely exist, such as high-speed rotating mechanical operation areas, high-voltage areas, and chemical treatment areas storing harmful substances. The safety management of these areas is not only the cornerstone of enterprise operation, but also the top priority of ensuring the safety of employees' lives.

[0003] When managing these dangerous areas, it is necessary to detect the entry of personnel. Although traditional personnel entry detection methods can perform monitoring tasks to a certain extent, their inherent limitations are becoming increasingly prominent and are difficult to meet the high standards of modern industrial safety management. First of all, although the method of manual monitoring is intuitive and has a certain degree of flexibility, the problem of its low efficiency cannot be ignored. And manual monitoring depends on the personal alertness and reaction speed of the monitoring personnel, and these factors are extremely vulnerable to interference by human factors such as physiological fatigue and mental inattention. Especially in the case of continuous monitoring for a long time, the reliability of manual monitoring will be greatly reduced, and it is difficult to achieve effective all-weather and dead-angle-free monitoring, which undoubtedly increases the risk of safety accidents. Secondly, although common devices such as traditional infrared sensors and laser sensors can detect the entry of personnel to a certain extent, their performance is often limited by environmental factors. Temperature changes, dust accumulation, light interference, etc. may all cause the sensors to give false alarms or missed alarms, reducing the accuracy and reliability of safety monitoring. This technical limitation not only increases the difficulty of safety management, but may also interfere with the normal production process due to misoperations caused by false alarms. Moreover, the lack of intelligence in traditional detection methods is also a major shortcoming. In the current era of intelligence, industrial safety management is developing towards a more automated and intelligent direction. However, traditional manual monitoring and simple sensor detection methods cannot achieve real-time tracking and accurate identification of personnel entering dangerous areas. This lack of intelligent monitoring method not only limits the depth and breadth of safety management, but also is difficult to meet the needs of future industrial development. Therefore, it needs to be improved. Summary of the Invention

[0004] The purpose of the present invention is to provide a personnel security detection method to solve the problems raised in the above background art.

[0005] In order to achieve the above purpose, the present invention provides the following technical solution: A personnel security detection method, the specific steps are as follows:

[0006] Step 1: Personnel detection

[0007] Object Detection Based on RetinaNet: The feature maps of three layers, c3, c4, and c5, are extracted through ResNet as the backbone network. The sizes of the feature maps are 1 / 8, 1 / 16, and 1 / 32 of the original image size respectively. After c3, c4, and c5 pass through the FPN (Feature Pyramid Network) structure to fuse multi-scale features, the feature maps p3, p4, p5, p6, and p7 are obtained. Multiple anchor boxes are preset on the feature maps from p3 to p7, and two sub-networks are used for the feature maps from p3 to p7 respectively, namely the classification network and the regression of the offset of the detection box position. Predictions of the target category and position offset are generated accordingly for each layer. Combining with the preset anchor boxes, the coordinate predictions on the multi-scale feature layers are obtained. After post-processing with NMS (Non-Maximum Suppression), the final detection categories and coordinate results are obtained;

[0008] Object Detection Based on OpenCV: First, a template image is established. The template is moved from left to right and top to bottom on the target image. At each position, the similarity between the template image and the original image area is calculated. The similarity matrix is calculated by the correlation coefficient matching method, and the position with the maximum similarity and the maximum similarity value are found. By comparing the maximum similarity value with the threshold, it is judged whether the matching is successful. If the maximum similarity value exceeds the threshold and the detection area meets the requirements, the matching result is marked with a rectangular box. By comparing the overlap degree, the results of template matching and object detection are compared, and the object detection result is taken as the main one to obtain the comprehensive detection result;

[0009] Step 2: Kalman Filter Prediction

[0010] Use the RetinaNet and the OpenCV correlation coefficient template matching algorithm in Step 1 to complete object detection. Person detection has been completed before tracking, and the person coordinate bounding box and feature set are obtained. After inputting them into the DeepSort algorithm, the Kalman filter first judges whether there is a tracking value. If there is, the prior probability prediction is made on its position information to obtain the prior prediction;

[0011] Step 3: Matching of Tracking Value and Observation Value

[0012] After the prior prediction is obtained, it is necessary to match the tracking value and the observation value. DeepSort uses a cascaded matching plus intersection over union (IoU) matching series strategy. The tracking value is classified into types that match the tracking prediction and the observation value, pending, and deleted according to the state; the cascaded matching is only performed on the tracking values that match the observation value successfully, and they are batch-matched according to the time distance from the previous match. The cost matrix used for matching is constructed by the cosine similarity distance and the Mahalanobis distance of the features. The Hungarian algorithm is used to calculate the cost matrix to obtain the matched and unmatched ones; after the tracking values that match the observation value successfully are matched using the Hungarian algorithm, the unmatched and pending tracking values are combined into a new set for IoU matching; the IoU matching directly constructs an IoU cost matrix with all tracking values and observation values as elements, and the Hungarian algorithm is used for matching. The matching method is the same as that of the cascaded matching.

[0013] Step 4: Kalman filter update

[0014] After matching, the matched, unmatched tracking values and unmatched observation values are obtained. The matched tracking values are corrected in turn, the unmatched tracking values are updated in state, the unmatched observation values are converted into tracking values, and the feature sets of the matched tracking values are updated; after the Kalman update is completed, the core operations of DeepSort are all completed. The subsequent operations are to update the state of each tracking value, delete the dead tracking values, and update the feature sets of the matched tracking values. After all the updates are completed, it enters the next frame of detection and tracking.

[0015] Preferably, if the maximum similarity value in Step 1 exceeds the threshold and the detection area meets the requirements, the matching result is marked by a rectangular box, and the offset of the label is calculated by comparing with the standard position; conversely, if no target is matched, it means the label does not exist.

[0016] Preferably, the similarity matrix is calculated by the correlation coefficient matching method in Step 1, and the formula is:

[0017]

[0018] where R is the similarity result matrix; R(x, y) represents the similarity between the area at x, y and the template; T is the template image matrix; I is the target image matrix; T' is the mean-subtracted matrix of the template image; I' is the mean-subtracted matrix of the target image; w and h represent the width and height of the template image and the target image in their respective formulas; x and y represent the coordinates of the upper-left corner element of the current search box in the target image matrix; x' and y' represent the relative coordinates of the elements within the search box in the target image matrix; represent the coordinates of the template elements in the template image matrix; x” and y” represent the element coordinates of the template image matrix and the target image matrix in their respective formulas.

[0019] Preferably, when performing Kalman filter prediction in step 2, a prior prediction is made on the tracking value at time t-1, and the Kalman filter formula is used as follows:

[0020]

[0021] is the prior prediction; x t-1 is the state information matrix of the tracking value at time t-1, which is an 8-dimensional long vector [cx, cy, w, h, vx, vy, vw, vh], representing the position information and the corresponding speed information; F is the state transition matrix from time t-1 to time t; dt is the time interval between consecutive frames; P t-1 and are respectively the 8*8 covariance matrix at time t-1 and its prior prediction at time t. When the tracking value is initialized, its covariance matrix and mean matrix are generated from the height of the target box and the coordinate length and width information of the target box; Q is the motion estimation error of the Kalman filter, representing the degree of uncertainty.

[0022] Preferably, for the cost matrix in step 3, first calculate the cosine distance according to the features corresponding to each tracking value and observation value. Cosine distance = 1 - cosine similarity, where each element in the cost matrix is the cosine distance; after obtaining the first cost matrix, adjust it through the Mahalanobis distance. If the Mahalanobis distance of a certain element in the cost matrix is greater than the threshold, modify the value; after the modification is completed, it is the final cost matrix for cascade matching.

[0023] Preferably, the cosine similarity formula is as follows:

[0024]

[0025] where A and B are respectively the features corresponding to the observation value and the tracking value, which is a vector with a length of 128. After normalization, ||A||2 and ||B||2 are 1; i is the index value of the feature vector; n is the maximum value of the feature vector index value, which is 128.

[0026] Preferably, the Kalman update formula in step 4 is as follows:

[0027]

[0028] where, is the covariance matrix at the predicted time t; C is the measurement matrix; C T is the transpose of the measurement matrix; R is the noise matrix, which is a 4×4 diagonal matrix; is the state information matrix of the tracking value at the predicted time t; y kDetect the coordinate information (d_cx, d_cy, d_r, d_h) of the target observation value for the current frame.

[0029] Preferably, when performing personnel tracking detection in step four, if the Retinanet template matching DeepSort is used for personnel target tracking detection, and the target is detected in the first two frames and the personnel position remains unchanged in the third frame, it can be determined that the interference is excluded.

[0030] Preferably, when updating the feature set in step four, if a certain tracking target is not successfully matched for multiple consecutive frames, its historical features are automatically reduced in priority in the feature set and gradually replaced by the latest successfully matched observation features.

[0031] The beneficial effects of the present invention are as follows:

[0032] Through the Retinanet target detection algorithm based on deep learning, the personnel targets in the image can be accurately detected. Combining the DeepSort multi-target tracking algorithm and the opencv correlation coefficient template matching algorithm further improves the detection accuracy and robustness. And through the combination of the multi-target tracking algorithm and the template matching algorithm, the real-time tracking and recognition of personnel entering the dangerous area can be realized. The multi-scale feature fusion supports the detection of targets of different sizes, the template matching threshold is adjustable to adapt to complex lighting and occlusion scenarios, and the dynamic update of the covariance matrix and the feature set management enhance the generalization ability of the algorithm to the dynamic environment, enabling it to adapt to different industrial environments and dangerous areas, and having strong versatility and scalability. Brief Description of the Drawings

[0033] Figure 1 It is the network structure of the Retinanet of the present invention. Detailed Embodiments

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] As Figure 1 shown, the embodiment of the present invention provides a personnel security detection method, and the specific steps are as follows:

[0036] Step 1: Personnel Detection

[0037] Object Detection Based on RetinaNet: The feature maps of three layers, namely c3, c4, and c5, are extracted through ResNet as the backbone network. The sizes of the feature maps are 1 / 8, 1 / 16, and 1 / 32 of the original image size respectively. After c3, c4, and c5 pass through the FPN (Feature Pyramid Network) structure to fuse multi-scale features, the feature maps p3, p4, p5, p6, and p7 are obtained. Multiple anchor boxes are preset on the feature maps from p3 to p7, and two sub-networks are used for the feature maps from p3 to p7 respectively, the classification network and the regression of the offset of the detection box position. Predictions of the target category and position offset are generated for each layer accordingly. Combining with the preset anchor boxes, the coordinate predictions on the multi-scale feature layers are obtained. After post-processing with NMS (Non-Maximum Suppression), the final detection categories and coordinate results are obtained;

[0038] Object Detection Based on OpenCV: First, a template image is established. The template is moved from left to right and top to bottom on the target image. At each position, the similarity between the template image and the original image area is calculated. The similarity matrix is calculated by the correlation coefficient matching method, and the position with the maximum similarity and the maximum similarity value are found. By comparing the maximum similarity value with the threshold, it is judged whether the matching is successful. If the maximum similarity value exceeds the threshold and the detection area meets the requirements, the matching result is marked with a rectangular box. By comparing the overlap degree, the results of template matching and object detection are compared, and the object detection result is taken as the main one to obtain the comprehensive detection result;

[0039] Step 2: Kalman Filter Prediction

[0040] Use the RetinaNet and the OpenCV correlation coefficient template matching algorithm in Step 1 to complete object detection. Person detection has been completed before tracking, and the person coordinate bounding box and feature set are obtained. After inputting them into the DeepSort algorithm, the Kalman filter first judges whether there is a tracking value. If there is, the prior probability prediction of its position information is carried out to obtain the prior prediction;

[0041] Step 3: Matching of Tracking Value and Observation Value

[0042] After obtaining the prior prediction, it is necessary to match the tracking value and the observation value. DeepSort uses a cascaded matching plus intersection over union (IoU) matching strategy in series. The tracking value is classified into types that match the tracking prediction and the observation value, pending, and deleted according to the state; the cascaded matching is only performed on the tracking values that match the observation value successfully, and they are batch-matched according to the time distance from the previous match. The cost matrix used for matching is constructed by the cosine similarity distance and the Mahalanobis distance of the features, and the Hungarian algorithm is used to calculate the cost matrix to obtain the matched and unmatched ones; after the tracking values that match the observation value successfully are matched using the Hungarian algorithm, the unmatched and pending tracking values are combined into a new set for IoU matching; the IoU matching directly constructs an IoU cost matrix with all tracking values and observation values as elements, and the Hungarian algorithm is used for matching, and the matching method is the same as that of the cascaded matching;

[0043] Step 4: Kalman filter update

[0044] After matching, the matched, unmatched tracking values and unmatched observation values are obtained. The matched tracking values are corrected in turn, the state of the unmatched tracking values is updated, the unmatched observation values are converted into tracking values, and the feature set of the matched tracking values is updated; after the Kalman update is completed, the core operations of DeepSort are all completed. The subsequent operations are to update the state of each tracking value, delete the dead tracking values, and update the feature set of the matched tracking values, and all updates are completed and enter the next frame for detection and tracking.

[0045] The Retinanet deep learning algorithm is used to detect the positions of personnel and generate coordinate boxes. At the same time, the OpenCV template matching algorithm is combined to compare with the preset template to optimize the detection accuracy. The detected personnel coordinates and features are input into the DeepSort algorithm. The Kalman filter is used to predict the motion trajectory and match the targets in the front and rear frames, and the cascaded matching and IoU matching strategies are adopted. Combining the feature similarity and the motion trajectory to associate and track the targets, the target position and feature information are dynamically updated according to the matching results, and the invalid tracking data is automatically cleared, and finally continuous and accurate warnings for personnel entering the dangerous area are realized.

[0046] Among them, in Step 1, if the maximum similarity exceeds the threshold and the detection area meets the requirements, the matching result will be marked by a rectangular box.

[0047] By the dual conditions of the similarity threshold and the detection area to screen the matching results, the false matching probability is significantly reduced.

[0048] Among them, in Step 1, the correlation coefficient matching method is used to calculate the similarity matrix, and the formula is:

[0049]

[0050] Among them, R is the similarity result matrix; R(x, y) represents the similarity between the regions at x and y and the template; T is the template image matrix; I is the target image matrix; T' is the mean-subtracted matrix of the template image; I' is the mean-subtracted matrix of the target image; w and h represent the width and height of the template image and the target image respectively in their respective formulas; x and y represent the coordinates of the upper left corner element of the current search box in the target image matrix; x' and y' represent the relative coordinates of the elements framed by the search box in the target image matrix, and represent the coordinates of the template elements in the template image matrix; x'' and y'' represent the element coordinates of the template image matrix and the target image matrix respectively in their respective formulas.

[0051] Through similarity calculation, local feature similar regions can be effectively captured, especially suitable for scenarios where there is partial occlusion between the template and the target, further improving the generalization ability of template matching.

[0052] Among them, in the Kalman filter prediction in step two, a prior prediction is made on the tracking value at time t - 1, and the Kalman filter formula used is:

[0053]

[0054] is the prior prediction; x t-1 is the state information matrix of the tracking value at time t - 1, which is an 8-dimensional long vector [cx, cy, w, h, vx, vy, vw, vh], representing the position information and the corresponding speed information; F is the state transition matrix from time t - 1 to time t; dt is the time interval between consecutive frames; P t-1 and are respectively the 8 * 8 covariance matrix at time t - 1 and its prior prediction at time t. When the tracking value is initialized, its covariance matrix and mean matrix are generated from the height of the target box and the coordinate length and width information of the target box; Q is the motion estimation error of the Kalman filter, representing the degree of uncertainty.

[0055] By introducing the motion estimation error Q, the degree of uncertainty in the prediction process can be further quantified, making the prediction result more reliable, providing a solid foundation for subsequent data association and tracking.

[0056] Among them, for the cost matrix in step three, first calculate the cosine distance according to the features corresponding to each tracking value and the observation value. Cosine distance = 1 - cosine similarity, where each element in the cost matrix is the cosine distance; after obtaining the first cost matrix, adjust it through the Mahalanobis distance. If the Mahalanobis distance of an element in the cost matrix is greater than the threshold, modify the value; after the modification is completed, it is the final cost matrix for cascade matching.

[0057] By constructing a cost matrix by combining the cosine distance and the Mahalanobis distance, the accuracy of data association is improved, making the obtained cascaded matching cost matrix more reliable, which can optimize the association between tracking and observation and improve the performance of the overall tracking system.

[0058] Among them, the cosine similarity formula is:

[0059]

[0060] Among them, A and B are the features corresponding to the observation value and the tracking value respectively, which is a vector with a length of 128. After normalization, ||A||2 and ||B||2 are 1; i is the index value of the feature vector; n is the maximum value of the feature vector index value, which is 128.

[0061] Through the cosine similarity, the algorithm can focus on the distribution pattern of the target key features rather than the absolute intensity, providing a reliable feature comparison basis for long-term tracking.

[0062] Among them, the Kalman update formula in step four is:

[0063]

[0064] Among them, is the covariance matrix at the predicted time t; C is the measurement matrix; C T is the transpose of the measurement matrix; R is the noise matrix, which is a 4×4 diagonal matrix; is the state information matrix of the tracking value at the predicted time t; y k is the coordinate information (d_cx, d_cy, d_r, d_h) of the detected target observation value in the current frame.

[0065] Through the Kalman update formula, the sensor noise and the motion model error can be effectively suppressed, and the accuracy of target position correction can be significantly improved.

[0066] Among them, when performing personnel tracking and detection in step four, if the Ret i nanet template matching DeepSort is used for personnel target tracking and detection, and the target is detected in the first two frames and the personnel position remains unchanged in the third frame, it can be determined as interference exclusion.

[0067] By analyzing the temporal continuity to identify stationary or false targets, the problem of false detection caused by equipment vibration, light and shadow flickering or background interference is effectively solved. At the same time, the omission of real dangerous events is avoided, and the signal-to-noise ratio and practicality of the alarm system are further improved.

[0068] Among them, when updating the feature set in step four, if a certain tracking target fails to be successfully matched for multiple consecutive frames, its historical features will be automatically reduced in priority in the feature set and gradually replaced by the latest successfully matched observation features.

[0069] Avoid the problem of feature invalidation caused by long-term non-update, thereby improving the stability of continuous tracking of target personnel in complex scenarios.

[0070] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0071] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A personnel security detection method, characterized in that, The specific steps are as follows: Step 1: Personnel detection Object detection based on RetinaNet: Using ResNet as the backbone network, the feature maps of three layers, c3, c4, and c5, are extracted. The sizes of the feature maps are 1 / 8, 1 / 16, and 1 / 32 of the original image size respectively. After c3, c4, and c5 are fused with multi-scale features through the FPN feature pyramid network structure, the feature maps p3, p4, p5, p6, and p7 are obtained; multiple anchor boxes are preset on the feature maps from p3 to p7, and two sub-networks are used for the feature maps from p3 to p7 respectively, the classification network and the regression of the position offset of the detection box. Each layer generates the prediction of the target category and the position offset accordingly. Combining with the preset anchor boxes, the coordinate prediction on the multi-scale feature layers is obtained. After post-processing with NMS (Non-Maximum Suppression), the final detection category and coordinate results are obtained; Object detection based on opencv: First, a template image is established. On the target image, the template is moved from left to right and from top to bottom. At each position, the similarity between the template image and the original image area is calculated. The similarity matrix is calculated by the correlation coefficient matching method, and the position with the maximum similarity and the maximum similarity value are found. By comparing the maximum similarity value with the threshold, it is judged whether the matching is successful. If the maximum similarity value exceeds the threshold and the detection area meets the requirements, the matching result is marked with a rectangular box. By comparing the overlap degree of the template matching and the object detection results, and taking the object detection result as the main one, the comprehensive detection result is obtained; Step 2: Kalman filter prediction Use the RetinaNet and opencv correlation coefficient template matching algorithms in Step 1 to complete object detection. Personnel detection has been completed before tracking, and the personnel coordinate bounding box and feature set are obtained; After inputting it into the DeepSort algorithm, the Kalman filter first judges whether there is a tracking value. If there is, the prior probability prediction is made on its position information to obtain the prior prediction; Step 3: Matching of tracking value and observation value After the prior prediction is obtained, the matching of the tracking value and the observation value needs to be carried out. DeepSort uses a strategy of cascaded matching plus intersection over union (IoU) matching in series. The tracking value is divided into types that match the tracking prediction and the observation value, pending, and deleted according to the state; Cascaded matching is only executed for the tracking values that match the observation value successfully. They are batch-matched according to the time distance from the last match. The cost matrix used for matching is constructed by the cosine similarity distance and the Mahalanobis distance of the features. The Hungarian algorithm is used to calculate the cost matrix to obtain the matched and unmatched ones; after the tracking values that match the observation value successfully are matched using the Hungarian algorithm, the unmatched and pending tracking values are combined into a new set for IoU matching; IoU matching directly constructs an IoU cost matrix with all tracking values and observation values as elements, and the Hungarian algorithm is used for matching. The matching method is the same as that of cascaded matching; Step 4: Kalman filter update After matching, the matched, unmatched tracking values and unmatched observation values are obtained. Subsequently, the successfully matched tracking values are corrected, the unmatched tracking values are updated in status, the unmatched observation values are converted into tracking values, and the feature sets of the successfully matched tracking values are updated. After the Kalman update is completed, all the core operations of DeepSort are completed. The subsequent operations are to update the status of each tracking value, delete the dead tracking values, and update the feature sets of the successfully matched tracking values. After all the updates are completed, the detection and tracking of the next frame are entered.

2. The personnel security detection method according to claim 1, characterized in that: In step 1, if the maximum similarity value exceeds the threshold and the detection area meets the requirements, the matching result is marked with a rectangular box.

3. The personnel security detection method according to claim 1, characterized in that: In step 1, the correlation coefficient matching method is used to calculate the similarity matrix, and the formula is: where R is the similarity result matrix; R(x, y) represents the similarity between the area at x, y and the template; T is the template image matrix; I is the target image matrix; T' is the mean-subtracted matrix of the template image; I' is the mean-subtracted matrix of the target image; w and h represent the width and height of the template image and the target image in their respective formulas; x and y represent the coordinates of the upper-left corner element of the current search box in the target image matrix; x' and y' represent the relative coordinates of the elements within the search box in the target image matrix; represent the coordinates of the template elements in the template image matrix; x'' and y'' represent the element coordinates of the template image matrix and the target image matrix in their respective formulas.

4. A personnel security detection method according to claim 1, characterized in that: In step 2, when performing Kalman filter prediction, a prior prediction is made on the tracking value at time t - 1, and the Kalman filter formula is used: is the prior prediction; x t-1 is the state information matrix of the tracking value at time t-1, which is an 8-dimensional long vector [cx, cy, w, h, vx, vy, vw, vh], representing the position information and the corresponding speed information; F is the state transition matrix from time t-1 to time t; dt is the time interval between consecutive frames; P t-1 and are the 8*8 covariance matrix at time t-1 and its prior prediction at time t respectively. When the tracking value is initialized, its covariance matrix and mean matrix are generated from the height of the target box and the coordinate length and width information of the target box; Q is the motion estimation error of the Kalman filter, representing the degree of uncertainty.

5. A personnel security detection method according to claim 1, characterized in that: In step 3, for the cost matrix, first, the cosine distance is calculated based on the features corresponding to each tracking value and observation value. The cosine distance = 1 - cosine similarity, and each element in the cost matrix is the cosine distance. After obtaining the first cost matrix, it is adjusted by the Mahalanobis distance. If the Mahalanobis distance of an element in the cost matrix is greater than the threshold, the value is modified. After the correction is completed, it is the final cost matrix for cascade matching.

6. The personnel security detection method according to claim 5, characterized in that: The cosine similarity formula is: where A and B are the features corresponding to the observation value and the tracking value respectively, which is a vector with a length of 128. After normalization, ||A||2 and ||B||2 are 1; i is the index value of the feature vector; n is the maximum value of the feature vector index value, which is 128.

7. A personnel security detection method according to claim 1, characterized in that: The Kalman update formula in step 4 is: Among them, is the covariance matrix at the predicted time t; C is the measurement matrix; C T is the transpose of the measurement matrix; R is the noise matrix, which is a 4×4 diagonal matrix; is the tracking value state information matrix at the predicted time t; y k is the coordinate information (d_cx, d_cy, d_r, d_h) of the detected target observation value in the current frame.

8. A personnel security detection method according to claim 1, characterized in that: In step 4, when performing personnel tracking detection, if Retinanet template matching and DeepSort are used for personnel target tracking detection, and the target is detected in the first two frames and the personnel position remains unchanged in the third frame, it can be determined that the interference is excluded.

9. The personnel security detection method according to claim 1, wherein: In step 4, when updating the feature set, if a certain tracking target is not successfully matched for multiple consecutive frames, the priority of its historical features in the feature set is automatically reduced, and they are gradually replaced with the latest successfully matched observation features.

Citation Information

Patent Citations

  • Transparent wine bottle quality detection method

    CN115791817A

  • System and method for surveillance of goods

    US20230245460A1