A method for detecting the staying behavior of people in public places

Through the YOLOv5 model and Deep-SORT algorithm combined with ReID technology, accurate detection and real-time early warning of personnel staying behavior in public places is achieved, the problems of accuracy, cost and privacy protection in the existing technology are solved, and efficient monitoring and early warning functions are provided.

CN118334743BActive Publication Date: 2025-07-11SUZHOU LUOPAN NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410478030.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2025-07-11
Estimated Expiration
2044-04-19

AI Technical Summary

Technical Problem

The prior art personnel stay detection methods in public places have problems such as low accuracy, high cost, privacy protection problems and insufficient real-time performance, especially in terms of lighting, occlusions, sensor coverage and computing resources.

Method used

The YOLOv5 model is used for object detection and Deep-SORT algorithm combined with Kalman filtering and Hungarian algorithm for multi-object tracking, combined with ReID technology for pedestrian re-identification, and video streams are processed through encrypted storage and access control to achieve accurate stay behavior detection and real-time early warning.

Benefits of technology

It improves the accuracy and real-time detection, ensures data privacy protection, can automatically extract and analyze personnel's behavioral characteristic information, realize real-time monitoring and early warning, and is suitable for monitoring needs in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118334743B_ABST
    Figure CN118334743B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting the staying behavior of people in public places, which relates to the technical field of public behavior detection, and includes the following steps: Raw data: Install cameras in public places to continuously capture video streams, obtain raw video data, and transmit it to the preprocessing center through the network. Encrypt and store the raw video data and perform access control to obtain the preprocessed target data; Based on deep learning technology, the present invention uses the YOLOv5 model to detect, identify, screen, mark, and track the target pedestrians captured by the cameras installed in public places, accurately identify and track the abnormal staying and wandering behaviors of suspicious people, further monitor and process the target data through ReID technology, and combine the Deep-SORT algorithm to determine the trajectory of the associated target pedestrians in the multi-target tracking information, realizing the functions of automatically extracting and analyzing the behavior feature information and movement trajectories of people.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of public behavior detection, and particularly to a method for detecting the staying behavior of people in public places. Background Art

[0002] The person staying algorithm is mainly used to judge and analyze the staying time of people in a specific area. At present, the conventional technologies in the domestic and international industries include two types: video analysis-based and sensor-based.

[0003] The person staying algorithm based on video analysis usually uses computer vision technology and deep learning algorithms to detect and track people. By analyzing video frames, the entry and exit times of people can be detected, and thus the staying time can be calculated. This method does not require additional hardware devices and can directly utilize existing surveillance cameras. However, video analysis technology is sensitive to factors such as lighting conditions, occlusions, and the number of people, and requires a large amount of computing resources and storage space, with relatively high costs. In addition, since video data may involve privacy protection issues, it is necessary to strengthen the management and supervision of data.

[0004] The person staying algorithm based on sensors calculates the staying time by deploying various sensors such as infrared sensors and WiFi fingerprint sensors to detect the time when people enter and leave the area. The cost of sensors is relatively low and they are not very sensitive to environmental changes. Moreover, sensor technology needs to be deployed in specific areas and may be affected by factors such as the moving speed of people and the coverage range of sensors. In addition, since sensor data may involve privacy protection issues, it is necessary to strengthen the management and supervision of data.

[0005] The existing technologies have the following deficiencies:

[0006] 1. Accuracy problem: The video analysis-based method may be affected by factors such as lighting and occlusions, while the sensor-based method may be affected by factors such as the moving speed of people and the coverage range of sensors, resulting in insufficient accuracy.

[0007] 2. Cost problem: The video analysis-based method requires a large amount of computing resources and storage space, with relatively high costs. Although the sensor-based method has relatively low costs, it requires the deployment of a large number of sensors, and there are also cost problems.

[0008] 3. Privacy protection problem: The video analysis-based method may involve privacy protection issues, and it is necessary to strengthen the protection and management of video data. The sensor-based method may also involve privacy protection issues, and it is necessary to strengthen the management and supervision of sensor data.

[0009] 4. Real-time issue: For scenarios with high real-time requirements, conventional technologies may not be able to meet the needs, and more efficient algorithms and technologies need to be adopted to improve real-time performance.

[0010] In summary, although the conventional technologies of the personnel staying algorithm at the current stage have certain application values, there are still many disadvantages and deficiencies, and further improvement and perfection are needed.

[0011] The above information disclosed in the background art section is only used to strengthen the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0012] The object of the present invention is to provide a method for detecting the staying behavior of personnel in public places. The present invention performs detection, recognition, screening, marking and tracking processing on the target pedestrians captured by the cameras installed in public places by using the YOLOv5 model, further monitors the target data by using the ReID technology, and combines the Deep-SORT algorithm to determine the trajectories of the associated target pedestrians in the multi-target tracking information, and performs encrypted storage and access control processing on the video stream continuously captured by the cameras, so as to solve the problems in the above background art.

[0013] In order to achieve the above object, the present invention provides the following technical solution: A method for detecting the staying behavior of personnel in public places, including the following steps:

[0014] S1. Original data: Install cameras in public places to continuously capture video streams, obtain the original video data, and transmit it to the preprocessing center through the network, and perform encrypted storage and access control on the original video data to obtain the preprocessed target data;

[0015] S2. Target detection: Use the YOLOv5 model to detect, recognize, screen and mark the pedestrians in the target data, and generate a pedestrian target data set;

[0016] S3. Target tracking: Use the Deep-SORT algorithm combined with the Kalman filter and the Hungarian algorithm to perform cascade matching and trajectory prediction on the target pedestrians in the pedestrian target data set, realize multi-target tracking, and obtain the motion trajectory information of multiple targets;

[0017] S4. Pedestrian re-identification: Use the ReID technology to monitor and process the target data captured by different cameras in the public scene, perform re-identification of the target pedestrians with the motion trajectory information of multiple targets, and combine the Deep-SORT algorithm to determine the trajectories of the associated target pedestrians in the multi-target tracking information, and obtain trajectory association data to avoid ID jumps;

[0018] S5. Loitering Detection: After obtaining the motion trajectory information and trajectory association data of multiple targets, a determination method based on the distance of the motion trajectory is adopted to determine whether a target pedestrian is loitering according to displacement and distance, so as to determine whether the target person has a lingering behavior and generate a loitering detection result;

[0019] S6. Detection Result: When the loitering detection result determines that the target pedestrian is in a loitering state, an alarm signal is immediately triggered, and the abnormal behaviors and trajectories of the detected target pedestrians are communicated and transmitted in the form of images, video clips, and alarm information and reported back to the security personnel in public places, and response measures such as dispersing, sorting, and necessary early warning are taken in a timely manner.

[0020] Optionally, the preprocessing steps of the original video data are as follows:

[0021] By connecting several cameras in public places, continuously capture video streams to obtain the original video data, calibrated as Ovd, and where x and y respectively represent random variables in the video stream pixels, representing the pixel variable set in the continuous video stream, and are transmitted to the preprocessing center through the network;

[0022] Extract individual frames from the video stream and use a Gaussian filter to perform noise reduction and resolution adjustment on each frame. The noise reduction calculation formula is where G(Ovd) represents the Gaussian function of the original video data Ovd, and σ G represents the standard deviation of the Gaussian distribution of the random variables x and y. The resolution adjustment calculation formula is where P(x′, y′) represents the coordinates of the new pixel point after adjusting the resolution of P(x, y), and P(x, y) represents the coordinates of the source pixel point in the original video data Ovd. P(x + i, y + j) represents the neighbor pixel at P(x, y), and i and j respectively represent the neighborhood ranges around P(x, y), and w i,j represents the weight calculated by weighted averaging according to the distance between the point P(x, y) and the point P(x + i, y + j);

[0023] Before transmission and storage, the original video data Ovd is encrypted using the AES algorithm to obtain the encrypted data signal, calibrated as Eds. The encryption calculation formula is Eds = AES(Ovd + K0), and where AES represents the encryption function, Ovd represents the original video data, and K0 represents the key, respectively represent the randomly read pixel variables after encryption, represents the set of encrypted pixel variables;

[0024] Meanwhile, access control policies are implemented for personnel who can process data signals to ensure that only authorized personnel can access and process data signals;

[0025] For the encrypted data signal Eds, in order to facilitate transmission or storage, the encrypted data format is standardized, that is, hexadecimal encoding or Base64 encoding is used to process and change the representation of the data for storage and transmission, so as to obtain the encrypted data to be processed, designated as Tds, and where X and Y respectively represent the data formats obtained by re-standardizing and encoding the encrypted data signal Eds, and N1 represents a number of target data sets;

[0026] And the application of background subtraction technology to highlight moving targets, that is, the background difference algorithm is used to subtract the background model from the current frame to detect moving targets in the original video data. At the time t when the moving target in the current frame is located, at the position (x t , y t ) the pixel value is P′(x t , y t ), then the background difference calculation formula is D′(x t , y t ) = [P′(x t , y t ) - B′(x, y)]. In the formula, D′(x t , y t ) represents the absolute difference between the pixel at the position (x t , y t ) in the current frame at time t and the pixel of the background frame, B′(x, y) represents the pixel value at any time at the position in the background model, and x t , y t respectively represent the pixel values at the position of the moving target in the target data Tds at time t read from the target data Tds.

[0027] Optionally, the acquisition logic of the target data is as follows:

[0028] Read the encrypted data to be processed Tds, and perform the inverse algorithm decryption using the AES key. Then the AES decryption calculation formula is where Ovd represents the original video data, AES -1 represents the inverse algorithm of the encryption function, Tds represents the ciphertext data, represents the symmetric key;

[0029] Identify the original video data Ovd, and perform step-by-step processing calculations of formatting, scaling, and normalization on the original video data Ovd to obtain target data, calibrated as Md3. Among them, the formatting process automatically converts the color space of the operable frame sequence in the original video data Ovd using a video processing library, that is, from BGR to RGB format, to obtain formatted data, calibrated as Fd1, and Fd1 = {x1, y1} n , where x1 and y1 respectively represent the formatting results of variables x and y;

[0030] The scaling process uses the bilinear interpolation algorithm to scale each frame of the formatted data Fd1 to the resolution, to obtain scaled data, calibrated as Sd2. Then the bilinear interpolation calculation formula is where Sd2(x1′, y1′) represents the pixel value at the positions of x1′ and y1′, a, b, c, and d respectively represent the proximity of the relative positions between the target point of four adjacent pixel values in the image and the original pixel point, and x1′, y1′ represent the positions of points x1 and y1 in the original coordinate system mapped to the new coordinate system;

[0031] The normalization process normalizes each pixel value range [0, 255] in the image scaled data Sd2 to the range [0, 1] to obtain normalized data, that is, the final target data Md3. The normalization calculation formula is where Md3(x 0 , y 0 ) represents the value after normalization, x 0 , y 0 respectively represent the normalization results of the positions x′1 and y′1 in the new coordinate system, and Sd2(x′1, y′1) represents the pixel value at the position (x1′, y1′) after the original image is scaled.

[0032] Optionally, the target detection steps of the YOLOv5 model are as follows:

[0033] Input the target data Md3(x 0 , y 0 ) into the YOLOv5 network where the YOLOv5 model is located;

[0034] The convolutional layer, residual layer, downsampling layer, upsampling layer, and routing layer included in the YOLOv5 network perform feature extraction on the target data Md3 to obtain a target feature map, and the target feature map includes the prediction of the bounding box and the classification confidence;

[0035] The YOLOv5 model divides the image of the target feature map into grid cells. Based on the training and forward inference of the YOLOv5 model, bounding boxes can be detected, and multiple bounding boxes and corresponding confidence scores are predicted on each grid cell. Then, by using non-maximum suppression (NMS) to compare the intersection over union (IoU) threshold, overlapping bounding boxes are removed to ensure that each real target corresponds to only one bounding box;

[0036] The detection results with confidence higher than the set threshold are screened out, and the classification of moving pedestrian targets is performed;

[0037] The finally retained bounding boxes and classification results are parsed, and the detected pedestrian positions and class labels are output and stored.

[0038] Optionally, the acquisition logic of the pedestrian target dataset is as follows:

[0039] The YOLOv5 model receives the input target data Md3(x 0 , y 0 ), and performs feature extraction through convolutional layers, residual layers, downsampling layers, upsampling layers, and routing layers. For each convolutional layer, a convolutional kernel is used for feature extraction, and the output target feature map, calibrated as TS(x 0 , y 0 ), and the convolutional calculation formula is In the formula, TS(x 0 , y 0 ) represents the output feature map of the convolutional layer at the position (x 0 , y 0 ), K represents the convolutional kernel, Md3(x 0 +i0, y 0 +j0, k0) represents, K(i0, j0, k0) represents the value of the convolutional kernel at the position (x 0 , y 0 ), i0, j0, k0 respectively represent three different points x 0 +i0, y 0 +j0, k0 in the convolutional kernel K, x 0 +i0, y 0 +j0, k0 represent three different weight parameters in the convolutional kernel K, and b0 represents the bias term of the convolutional layer;

[0040] The target feature map TS(x 0 , y 0 ) after convolutional calculation is batch-normalized, and the batch-normalization calculation formula is And In the formula, represents the input target feature map TS(x 0 , y 0)The result of performing batch normalization for k batches, k 0 is represented as k batches, is represented as the target feature map TS(x 0 , y 0 ) input in k batches, B(k 0 ) is represented as the batch mean, ε is represented as the learnable parameter, is represented as the batch variance, c′ is represented as a positive number for dividing by zero to control the scaling ratio, Y(k 0 ) is represented as the predicted output value in the kth batch, λ is represented as the hyperparameter of the scaling factor, b′ is represented as the offset translation term for adjusting the predicted value;

[0041] The YOLOv5 model will use the feature map output by batch normalization for prediction training, generating the predicted output value Y(k 0 ) and the corresponding confidence P(Y), and P(Y) ∈ [0, 1], and generating prediction box information according to the predicted output value Y(k 0 ), and the prediction box information includes class probability, box position, and prediction box width and height, respectively calibrated as PC, BP, and WH, that is, Y(k 0 ) = {PC, BP, WH};

[0042] Then the class probability PC is to predict the probability distribution of each anchor box belonging to each class by using the normalized exponential function. The calculation formula for setting the class probability is and 0 ≤ m, n2 ≤ k. In the formula, PC is represented as the score of the target feature predicted by the YOLOv5 model, Y(k 0 ) m is represented as the feature vector of the predicted output value Y(k 0 ) in the current mth class, is represented as the feature vectors of k classes summation;

[0043] The box position BP adopts the anchor box mechanism, outputs the offset related to the anchor box through the Sigmoid function, and determines the final boundary box position according to the offset, the center coordinates of the original boundary box, and the information of the predicted box width and height WH. Then, according to the predicted output value Y(k 0 ), the center coordinate values of the predicted boundary box are set as (Yx, Yy), the width and height of the predicted boundary box are Yw and Yh respectively, and it is known that the upper left corner coordinates of the grid cell for image segmentation are (Gx, Gy), and the width and height are Gw and Gh respectively. Then the offset formula calculated by the Sigmoid function is dx = σ(Yx) + Gx, dy = σ(Yy) + Gy, and dw = Gw × e Yw , dh = Gh × e Yh, where dx and dy respectively represent the magnitudes of the offsets in the x-axis and y-axis directions, dw and dh represent the width and height of the final bounding box position, σ represents the hyperparameter of the predicted bounding box center coordinate values Yx and Yy, e represents the exponential function with the natural constant e as the base, that is, the box position BP = {dx, dy, dw, dh}, and the predicted box width and height WH = {dw, dh};

[0044] During the training of the YOLOv5 model, the network parameters are adjusted through the forward propagation algorithm, and the loss function is used to minimize the gap between the detection results and the actual annotations. The non-maximum suppression NMS is not adopted to remove redundant overlapping boxes, and several detection results of each pedestrian are retained for position fusion processing. After training, pedestrian detection is performed on the given new target data to generate a dataset containing only pedestrian targets, that is, the pedestrian target dataset, calibrated as Ptd. Among them, the calculation of the loss function and the method of position fusion processing are crucial. The loss function includes coordinate loss, object confidence loss, and class loss. The calculation formula of the coordinate loss is where L(CIoU) represents the loss value calculated by the CIoU loss function, and IoU represents the intersection over union, represents the Euclidean distance between the predicted box position BP and the center point of the grid cell true box (Gx, Gy), represents the diagonal length of the minimum enclosing region of the predicted box and the true box, and α represents the weight coefficient, represents the consistency of measuring the aspect ratio of the predicted box and the true box;

[0045] The calculation formula of the object confidence loss is L(BCE) = -P0ln(P(Y)) - (1 - P0)×(1 - ln(P(Y))). In the formula, L(BCE) represents the object confidence loss value calculated by using the binary cross-entropy loss algorithm BCE, P0 represents the true label of the detected person, and when the detected person exists, P0 = 1, when the detected person does not exist, P0 = 0, and P(Y) represents the detection object confidence output by the YOLOv5 model prediction result, and P(Y) ∈ [0, 1];

[0046] The calculation formula of the class loss is where L(PC) represents the difference loss value between the predicted class distribution and the true class distribution calculated by using the cross-entropy loss algorithm, k represents the total number of classes, p(Y k ) represents the true probability that the detected object belongs to class k, and p(PC) k represents the predicted probability that the detected object belongs to class k;

[0047] The position fusion process uses the average fusion algorithm to obtain the final bounding box coordinates of the detected target pedestrian and constructs the pedestrian target dataset Ptd. The calculation formula of the average fusion algorithm is and In the formula, and respectively represent the average values of the x-axis coordinate, y-axis coordinate, width, and height of the center point after the detection box fusion. represents the confidence of the fusion box, k represents the total number of categories, N2 represents the number of detected bounding boxes, and Yx i , Yy i , Yw i , Yh i respectively represent the coordinate positions and width and height of each bounding box.

[0048] Optionally, the multi-object tracking steps of the Deep-SORT algorithm for the pedestrian target dataset in combination with the Kalman filter and the Hungarian algorithm are as follows:

[0049] The Deep-SORT algorithm realizes efficient and accurate tracking of multiple targets in the pedestrian target dataset by comprehensively using Kalman filter prediction and data association of the Hungarian algorithm;

[0050] According to the pedestrian target dataset Ptd, the behavior targets of the pedestrians in the detected image are tracked. The Kalman filter is used to predict the position and speed state of the target at time t0 in the current frame. The prediction calculation formula of the Kalman filter is and In the formula, H(Ptd|t0) represents the state prediction value of the pedestrian target dataset Ptd at the current time t0. represents the state transition matrix of the pedestrian target dataset Ptd, and H(Ptd|t0 - 1) represents the state estimation value of the pedestrian target dataset Ptd at the previous time t0 - 1. represents the control input matrix of the pedestrian target dataset Ptd. represents the control input at time t0, P(t0) represents the predicted state covariance at time t0, P(t0 - 1) represents the state covariance matrix at the previous time t0 - 1, T represents the transpose. represents the process noise covariance matrix at time t0;

[0051] According to the target state H(Ptd|t0) predicted by the Kalman filter and the detection observed at the current time t0 and use the Hungarian algorithm to assign the trajectory where the detection target is located. The trajectory calculation formula of the Hungarian algorithm is In the formula, represents the target state and detection The total distance between, where λ represents a parameter for the equilibrium distance is expressed as a matrix product is expressed as the Mahalanobis distance based on Kalman filter prediction is expressed as the predicted value based on the cosine similarity between the target state H(Ptd|t0) and the detection ;

[0052] Again for each matched target and detection pair, use the update step of the Kalman filter to update the new state of the target position and velocity. Then the calculation formula for updating the new state of the Kalman filter is K(t0) = P(t0 - 1) × H(t0) T × (H(t0) × P(t0 - 1) × H(t0) T + R(t0)) -1 , and , P(t0 + 1) = (1 - K(t0) × H(t0)) × P(t0). In the formula, K(t0) represents the Kalman gain, H(t0) represents the observation model, R(t0) represents the observation noise covariance, P(t0 - 1) represents the state covariance matrix at the previous moment t0 - 1, T represents the transpose, H(Ptd|t0 + 1) represents, and H(Ptd|t0) represents the state prediction value of the pedestrian target dataset Ptd at the current moment t0 represents the control input matrix of the pedestrian target dataset Ptd, that is, the observed value of the pedestrian target dataset Ptd at the current moment t0, P(t0 + 1) represents the updated state probability at the next moment t0 + 1, and P(t0) represents the predicted state covariance at the moment t0

[0053] The change of the position of each tracked target over time is recorded to form a motion trajectory, and the motion trajectory of the target is optimized by predicting the current state and updating the state based on new observation data to achieve trajectory smoothing

[0054] Optionally, the acquisition logic of the motion trajectory information of the multi-target is as follows

[0055] Use the YOLOv5 model to detect the position information of the target pedestrians in the target data of the current frame to obtain the predicted target bounding boxes

[0056] Extract features from the detected predicted target bounding boxes, determine the confidence of the detection objects, perform target tracking from the pedestrian target dataset Ptd, and match the trajectories where the targets detected in the current frame are located

[0057] Use the Hungarian algorithm to match the trajectories where the detected targets are located, and use Kalman filtering to predict and update the positions of each trajectory of multi-target pedestrians in the next frame. Also, for each matched target-detection pair, use the update step of Kalman filtering again to update the new state of the target to improve the accuracy of data association;

[0058] According to the Kalman filtering smooth trajectory technique, by predicting the current state and updating the state based on new observation data, the motion trajectory is obtained, and then the motion trajectory information H(Ptd|t0+1) of multi-targets is obtained n , where H(Ptd|t0+1) represents the motion trajectory information of the detected target, and n represents the number of detected targets.

[0059] Optionally, the pedestrian re-identification steps of the ReID technology are as follows:

[0060] A1. Feature extraction: Adopt a convolutional neural network model to learn and extract the discriminative behavioral feature information of multi-target pedestrians from the pedestrian target dataset;

[0061] A2. Feature matching: By calculating the distance or similarity between features, match the pedestrian features extracted from the current target data with the pedestrian feature library extracted from the known pedestrian target dataset;

[0062] A3. Trajectory association: According to the results of feature matching, use the Hungarian algorithm to associate the target pedestrians of the target data with the motion trajectory information of the corresponding multi-targets in the known pedestrian target dataset, and find the best associated trajectory information.

[0063] Optionally, the acquisition logic of the trajectory association data is as follows:

[0064] Use the convolutional neural network CNN model to learn and extract the discriminative behavioral feature information from the pedestrian target dataset Ptd, denoted as f(Ptd). Then, the discriminative behavioral feature information is calculated using the convolutional layer formula in the CNN model as f(Ptd) = ReLU(W * Ptd + β0), where ReLU represents the activation function, W represents the convolutional kernel weight, β0 represents the bias term, and * represents the convolutional operation;

[0065] Use the Euclidean distance to calculate the distance between two feature vectors for feature matching, generating an associated distance matrix, denoted as Df i′j′ , then set two feature vectors according to the behavioral feature information f(Ptd) where i′ and j′ represent the feature of the detected target and the trajectory feature of the target pedestrian respectively, and are positive integers where 1 ≤ i′, j′ ≤ n. The formula for the Euclidean distance is

[0066] According to the distance matrix Df that generates the association i′j′ , use a matching algorithm such as the Hungarian algorithm to associate the behavioral feature information f(Ptd) of the target pedestrian and the motion trajectory information H(Ptd|t0+1) of multiple targets n , and associate the cost matrix and find the optimal match to obtain the trajectory association data, labeled as rg n , then the association degree calculation formula is and rg n ∈[0,1]. In the formula, rg n represents the association between the behavioral feature information f(Ptd) of the target pedestrian and the motion trajectory information H(Ptd|t0+1) of multiple targets n . represents the minimum distance cost between the detected target feature and the trajectory feature vector of the target pedestrian. n represents the number of detected targets. Then, calculate the association between the detected targets and the corresponding trajectories to obtain the optimal trajectory matching result;

[0067] To avoid ID jumps, Deep-SORT does not immediately delete the trajectory when the trajectory matching fails, but uses a timing counter to track the number of frames of each trajectory since the last successful match

[0068] Optionally, the acquisition logic of the loitering detection result is as follows:

[0069] Through the motion trajectory information H(Ptd|t0+1) of multiple targets n and the trajectory association data rg n , adopt the distance and displacement calculation method, and determine whether the target pedestrian stays or loiters in an area according to the ratio of the compared distance and displacement;

[0070] Among them, the best motion trajectory information for detecting and matching multiple target pedestrians in the common area is In the formula, represents the best motion trajectory information of multiple target pedestrians. x″ and y″ respectively represent the coordinates of the target pedestrian on the x-axis and y-axis in the plane position within the best motion trajectory information. t represents the time period for which the best motion trajectory information lasts. n3 represents the number of detected targets;

[0071] Arbitrarily select the best motion trajectory information of a target pedestrian and find the starting point to the ending point from the trajectory where the target pedestrian is located, labeled as Bti(x s ″,y s ″) t 、Bti(x e ″,y e ″) t, and counting each trajectory coordinate of M consecutive moving points of the target pedestrian on the optimal motion trajectory, calibrated as Bti(x″ M ,y″ M ) t ;

[0072] Then, calculate the distance and displacement of the target pedestrian from the starting point to the ending point of the trajectory. The displacement calculation formula is

[0073] The distance calculation formula is and M≥2;

[0074] The calculation result based on the ratio of the distance and displacement is Dt≠0, Pp∈[0,1]. In the formula, Pp represents the ratio of the distance and displacement. And for determining whether the target pedestrian stays or wanders in an area, a threshold is set, calibrated as Wt∈[0,0.75]. In the formula, Wt represents the behavior of staying or wandering in an area. Therefore, compare the magnitudes of the ratio and the threshold. When Pp>Wt, Pp∈(0.75,1], it is determined that the target pedestrian does not stay or wander in a public area; when Pp≤Wt, Pp∈[0,0.75], it is determined that the target pedestrian stays or wanders in a public area.

[0075] In the above technical solution, the technical effects and advantages provided by the present invention are as follows:

[0076] Based on deep learning technology, the present invention uses the YOLOv5 model to detect, identify, screen, mark and track the target pedestrians captured by the cameras installed in public places, accurately identify and track the abnormal staying and wandering behaviors of suspicious persons, further monitor and process the target data through the ReID technology, and combine the Deep-SORT algorithm to determine the trajectory of the associated target pedestrian in the multi-target tracking information, realizing the functions of automatically extracting and analyzing the behavior feature information and motion trajectory of personnel, achieving an accurate judgment of whether the target pedestrian belongs to the staying or wandering behavior, and improving the accuracy and reliability of monitoring;

[0077] By encrypting the storage and access control processing of the continuously captured video stream of the camera, the privacy protection of the collected data is further realized, and the security and confidentiality of the data are improved;

[0078] Based on the determination method of the distance of the movement trajectory, calculate the wandering detection result, trigger real-time early warning and information feedback to the security personnel in public places, realize the functions of real-time monitoring and early warning, notify the management personnel to take response measures such as dispersing, sorting out and necessary early warning, improve the efficiency and level of emergency response, provide strong support for public security work, and this application has wider applicability, real-time performance and flexibility, and can be customized and adjusted according to the special needs of different regions and different scenarios in public places to meet the monitoring needs and monitoring accuracy in different scenarios. Brief Description of the Drawings

[0079] Figure 1 It is a flowchart of the method for detecting the staying behavior of people in public places according to the present invention.

[0080] Figure 2 It is a flowchart of the pedestrian re-identification step according to the present invention. Detailed Embodiments

[0081] Now, the exemplary embodiments will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these exemplary embodiments are provided so that the present disclosure will be more complete and comprehensive, and will fully convey the concept of the exemplary embodiments to those skilled in the art.

[0082] The present invention provides a method for detecting the staying behavior of people in public places as shown in Figure 1 and includes the following steps:

[0083] S1. Original data: Install cameras in public places to continuously capture video streams, obtain original video data, and transmit it to the preprocessing center through the network. Encrypt and store the original video data and perform access control to obtain the preprocessed target data;

[0084] Specifically, the preprocessing steps of the original video data are as follows:

[0085] Connect several cameras to the public place to continuously capture video streams to obtain original video data, calibrated as Ovd, and where x and y respectively represent random variables in the video stream pixels, n1 represents the pixel variable set in the continuous video stream, and it is transmitted to the preprocessing center through the network;

[0086] Extract individual frames from the video stream and use a Gaussian filter to perform noise reduction and resolution adjustment on each frame to improve the accuracy of the original video data. The noise reduction calculation formula is where G(Ovd) represents the Gaussian function of the original video data Ovd, σ GDenoted as the standard deviation of the Gaussian distribution of random variables x and y, the resolution adjustment calculation formula is In the formula, P(x′, y′) represents the new pixel point coordinates after adjusting the resolution of P(x, y), and P(x, y) represents the source pixel point coordinates in the original video data Ovd. P(x + i, y + j) represents the neighbor pixels at P(x, y). i and j respectively represent the neighborhood ranges around P(x, y), w i,j Denoted as the weight calculated by weighted average according to the distance between point P(x, y) and point P(x + i, y + j);

[0087] Before transmission and storage, the original video data Ovd is encrypted using the AES algorithm to obtain the encrypted data signal, denoted as Eds, to ensure data security. The encryption calculation formula is Eds = AES(Ovd + K0), and In the formula, AES represents the encryption function, Ovd represents the plaintext of the original video data, and K0 represents the key, Denoted as the random pixel variables read after encryption respectively, Denoted as the set of encrypted pixel variables;

[0088] At the same time, an access control policy is implemented for the personnel who can process the data signal to ensure that only authorized personnel can access and process the data signal;

[0089] For the encrypted data signal Eds, in order to facilitate transmission or storage, the encrypted data format is standardized, that is, hexadecimal encoding or Base64 encoding is used to process and change the representation of the data for storage and transmission to obtain the data to be encrypted and processed, denoted as Tds, and In the formula, X and Y respectively represent the data formats obtained by re-standardizing and encoding the encrypted data signal Eds. N1 represents a number of target data sets to reduce storage space and improve processing efficiency;

[0090] And the background subtraction technology is applied to highlight the moving target, that is, the background difference algorithm is used to subtract the background model from the current frame to detect the moving target in the original video data. At the moment t when the moving target in the current frame is located at position (x t , y t ), the pixel value is P′(x t , y t ). Then the background difference calculation formula is D′(x t , y t ) = [P′(x t , y t ) - B′(x, y)]. In the formula, D′(x t , y t) is expressed as the absolute difference between the pixel at position (x t , y t ) at the current frame t and the pixel of the background frame. B′(x, y) is expressed as the pixel value at any time at the position in the background model. x t , y t respectively represent the pixel values at the position of the moving target at time t read from the target data Tds.

[0091] Specifically, the acquisition logic of the target data is as follows:

[0092] Read the encrypted data to be processed Tds, and perform inverse algorithm decryption using the AES key. The AES decryption calculation formula is In the formula, Ovd represents the original video data, AES -1 represents the inverse algorithm of the encryption function, Tds represents the ciphertext data, represents the symmetric key;

[0093] Identify the original video data Ovd, and perform step-by-step processing calculations of formatting, scaling, and normalization on the original video data Ovd to obtain the target data, calibrated as Md3. Among them, the formatting process is to automatically convert the color space of the operable frame sequence in the original video data Ovd using a video processing library, that is, convert from BGR to RGB format to obtain the formatted data, calibrated as Fd1, and Fd1 = {x1, y1} n , where x1 and y1 respectively represent the formatting results of variables x and y;

[0094] The scaling process uses the bilinear interpolation algorithm to scale each frame of the formatted data Fd1 to the resolution, and obtain the scaled data, calibrated as Sd2. The bilinear interpolation calculation formula is In the formula, Sd2(x1′, y1′) represents the pixel value at the positions of x1′ and y1′, a, b, c, and d respectively represent the proximity of the relative positions between the adjacent four pixel value target points and the original pixel points in the image, and x1′, y1′ represent the positions of points x1, y1 in the original coordinate system mapped to the new coordinate system;

[0095] The normalization process normalizes the range [0, 255] of each pixel value in the image scaled data Sd2 to the range [0, 1] to obtain the normalized data, that is, the final target data Md3. The normalization calculation formula is In the formula, Md3(x 0 , y 0 ) represents the value after normalization, x 0 , y 0They are respectively expressed as the normalization results of the positions x1' and y1' in the new coordinate system, and Sd2(x1', y1') represents the pixel value at the position (x1', y1') after the original image is scaled.

[0096] S2. Object detection: Use the YOLOv5 model to detect, identify, screen, and mark the pedestrians in the target data, and generate a pedestrian target data set.

[0097] Specifically, the object detection steps of the YOLOv5 model are as follows:

[0098] Input the target data Md3(x 0 , y 0 ) into the YOLOv5 network where the YOLOv5 model is located.

[0099] The convolutional layer, residual layer, downsampling layer, upsampling layer, and routing layer included in the YOLOv5 network perform feature extraction on the target data Md3 to obtain a target feature map, and the target feature map includes the prediction of the bounding box and the classification confidence.

[0100] The YOLOv5 model divides the image of the target feature map into grid cells one by one. Based on the training and forward inference of the YOLOv5 model, the bounding box can be detected, and multiple bounding boxes and corresponding confidence scores are predicted on each grid cell. Among them, the confidence score is expressed as the product of the probability that the predicted box contains the target and the accuracy of the predicted box. Then, by using non-maximum suppression (NMS) to compare the intersection-over-union (IoU) threshold, the overlapping bounding boxes are removed to ensure that each real target has only one corresponding bounding box.

[0101] Screen out the detection results with a confidence higher than the set threshold, and classify the moving pedestrian targets.

[0102] Analyze the finally retained bounding boxes and classification results, and output and store the detected pedestrian positions and class labels.

[0103] Specifically, the acquisition logic of the pedestrian target data set is as follows:

[0104] The YOLOv5 model receives the input target data Md3(x 0 , y 0 ) and performs feature extraction through the convolutional layer, residual layer, downsampling layer, upsampling layer, and routing layer. Then, for each convolutional layer, a convolutional kernel is used for feature extraction, and the output target feature map is calibrated as TS(x 0 , y 0 ). The convolutional calculation formula is In the formula, TS(x 0 , y 0 ) represents the convolutional layer at the position (x 0 , y0 ) The output feature map, K represents the convolutional kernel, Md3(x 0 + i0, y 0 + j0, k0) is expressed as, K(i0, j0, k0) is expressed as the value of the convolutional kernel at the position (x 0 , y 0 ), i0, j0, k0 respectively represent three different points x in the convolutional kernel K 0 + i0, y 0 + j0, k0's parameters, x 0 + i0, y 0 + j0, k0 represent three different weight parameters in the convolutional kernel K, b0 represents the bias term of the convolutional layer;

[0105] The target feature map TS(x 0 , y 0 ) after convolutional calculation is batch-normalized, and the formula for batch normalization is and In the formula, represents the result of normalizing the input target feature map TS(x 0 , y 0 ) for k batches, k 0 represents k batches, represents the input target feature map TS(x 0 , y 0 ) in k batches, B(k 0 ) represents the batch mean, ε represents the learnable parameter, represents the batch variance, c′ represents a non-zero positive number for controlling the scaling ratio, Y(k 0 ) represents the predicted output value in the k-th batch, λ represents the hyperparameter of the scaling factor, b′ represents the offset translation term for adjusting the predicted value;

[0106] The YOLOv5 model uses the feature map output by batch normalization for prediction training, generating the predicted output value Y(k 0 ) and the corresponding confidence P(Y), and P(Y) ∈ [0, 1], and generating prediction box information according to the predicted output value Y(k 0 ), and the prediction box information includes class probability, box position, and prediction box width and height, which are respectively calibrated as PC, BP, and WH, that is, Y(k 0 ) = {PC, BP, WH};

[0107] Then the class probability PC is obtained by using the normalized exponential function to predict the probability distribution of each anchor box belonging to each class. The formula for setting the class probability is and 0 ≤ m, n2 ≤ k, where PC represents the score of the target feature predicted by the YOLOv5 model, and Y(k 0 ) m represents the feature vector of the predicted output value Y(k 0 ) in the current m categories, represents the feature vectors of k categories summation, and the softmax function ensures that the sum of all category probabilities equals one;

[0108] The bounding box position BP adopts the anchor box mechanism, outputs the offsets related to the anchor box through the Sigmoid function, and determines the final bounding box position according to the offsets, the center coordinates of the original bounding box, and the width and height WH of the predicted box. Then, according to the predicted output value Y(k 0 ), the center coordinate values of the predicted bounding box are set as (Yx, Yy), the width and height of the predicted bounding box are Yw and Yh respectively, and it is known that the upper left coordinates of the grid cell for image segmentation are (Gx, Gy), and the width and height are Gw and Gh respectively. Then, the offset formula calculated by the Sigmoid function is dx = σ(Yx) + Gx, dy = σ(Yy) + Gy, and dw = Gw × e Yw , dh = Gh × e Yh , where dx and dy respectively represent the magnitudes of the offsets in the x-axis and y-axis directions, dw and dh represent the width and height of the final bounding box position, σ represents the hyperparameter of the center coordinate values Yx and Yy of the predicted bounding box, and e represents the exponential function with the natural constant e as the base. That is, the bounding box position BP = {dx, dy, dw, dh}, and the predicted box width and height WH = {dw, dh};

[0109] When the YOLOv5 model is training, it adjusts the network parameters through the forward propagation algorithm, uses the loss function to minimize the gap between the detection result and the actual annotation, does not use non-maximum suppression NMS to remove redundant overlapping boxes, retains several detection results of each pedestrian for position fusion processing, and after training, conducts pedestrian detection on the given new target data to generate a dataset containing only pedestrian targets, that is, the pedestrian target dataset, calibrated as Ptd, for facilitating subsequent behavior analysis and pedestrian flow statistics. Among them, the calculation of the loss function and the method of position fusion processing are crucial. The loss function includes coordinate loss, object confidence loss, and class loss. The calculation formula of the coordinate loss is where L(CIoU) represents the loss value calculated by the CIoU loss function, and IoU represents the intersection over union, represents the Euclidean distance between the predicted bounding box position BP and the center point of the ground truth box (Gx, Gy) of the grid cell, represents the diagonal length of the minimum enclosing region of the predicted box and the ground truth box, and α represents the weight coefficient, It is expressed as the consistency of measuring the aspect ratio of the predicted bounding box and the ground truth bounding box;

[0110] The calculation formula for the object confidence loss is L(BCE) = -P0ln(P(Y)) - (1 - P0)×(1 - ln(P(Y))), where L(BCE) represents the object confidence loss value calculated using the binary cross-entropy loss algorithm BCE, P0 represents the true label of the detected person, and when the detector exists, P0 = 1, when the detector does not exist, P0 = 0, and P(Y) represents the confidence of the detected object output by the YOLOv5 model prediction result, and P(Y) ∈ [0, 1];

[0111] The calculation formula for the class loss is where L(PC) represents the difference loss value between the predicted class distribution and the true class distribution calculated using the cross-entropy loss algorithm, k represents the total number of classes, p(Y k ) represents the true probability that the detected object belongs to class k, and p(PC) k represents the predicted probability that the detected object belongs to class k;

[0112] The position fusion process is to use the average fusion algorithm to obtain the final bounding box coordinates of the detected target pedestrians to construct the pedestrian target dataset Ptd. Then, the calculation formula for the average fusion algorithm is

[0113] where and respectively represent the average values of the x-axis coordinate, y-axis coordinate, width, and height of the center point after the detection box fusion. represents the confidence of the fusion box, k represents the total number of classes, represents the number of detected bounding boxes, and Yx i 、Yy i 、Yw i 、Yh i respectively represent the coordinate positions and width and height of each bounding box.

[0114] S3. Target tracking: The Deep-SORT algorithm is used in combination with the Kalman filter and the Hungarian algorithm to perform cascade matching and trajectory prediction on the target pedestrians in the pedestrian target dataset, achieve multi-target tracking, and obtain the motion trajectory information of multiple targets;

[0115] Specifically, the steps of the Deep-SORT algorithm combined with the Kalman filter and the Hungarian algorithm for multi-target tracking of the pedestrian target dataset are as follows:

[0116] The Deep-SORT algorithm realizes the efficient and accurate tracking of multiple targets in the pedestrian target dataset by comprehensively utilizing Kalman filter prediction and data association of the Hungarian algorithm. It can maintain high tracking stability and accuracy in a dynamically changing environment, providing reliable target motion trajectory information for subsequent analysis and applications;

[0117] According to the pedestrian target dataset Ptd, conduct behavioral target tracking on the pedestrians in the detected image, and use the Kalman filter to predict the position and velocity state of the target at time t0 in the current frame. Then the prediction calculation formula of the Kalman filter is and In the formula, H(Ptd|t0) represents the state prediction value of the pedestrian target dataset Ptd at the current time t0, represents the state transition matrix of the pedestrian target dataset Ptd, and H(Ptd|t0 - 1) represents the state estimation value of the pedestrian target dataset Ptd at the previous time t0 - 1, represents the control input matrix of the pedestrian target dataset Ptd, represents the control input at time t0, P(t0) represents the prediction state covariance at time t0, P(t0 - 1) represents the state covariance matrix at the previous time t0 - 1, T represents the transpose, represents the process noise covariance matrix at time t0;

[0118] Based on the target state H(Ptd|t0) predicted by the Kalman filter and the detection observed at the current time t0 and use the Hungarian algorithm to assign the trajectory where the detected target is located. Then the trajectory calculation formula of the Hungarian algorithm is In the formula, represents the total distance between the target state H(Ptd|t0) and the detection λ represents the parameter for balancing the distance, represents the matrix product, represents the Mahalanobis distance based on the Kalman filter prediction, represents the prediction value based on the cosine similarity between the target state H(Ptd|t0) and the detection ;

[0119] Once again, for each matched target and detection pair, use the update step of the Kalman filter to update the new state of the target position and velocity. Then the calculation formula for updating the new state of the Kalman filter is K(t0) = P(t0 - 1) × H(t0) T × (H(t0) × P(t0 - 1) × H(t0) T + R(t0)) -1 , and , P(t0 + 1) = (1 - K(t0)×H(t0))×P(t0), where K(t0) represents the Kalman gain, H(t0) represents the observation model, R(t0) represents the observation noise covariance, P(t0 - 1) represents the state covariance matrix at the previous time t0 - 1, T represents the transpose, H(Ptd|t0 + 1) represents, and H(Ptd|t0) represents the state prediction value of the pedestrian target dataset Ptd at the current time t0. It is represented as the control input matrix of the pedestrian target dataset Ptd, that is, the observed value of the pedestrian target dataset Ptd at the current time t0. P(t0 + 1) represents the updated state probability at the next time t0 + 1, and P(t0) represents the predicted state covariance at time t0.

[0120] The change in the position of each tracking target over time is recorded to form a motion trajectory, and the motion trajectory of the target is optimized by predicting the current state and updating the state based on new observation data to smooth the trajectory, reducing the influence of jumps and noise, thereby improving the overall quality of tracking.

[0121] Specifically, the acquisition logic of the motion trajectory information of multiple targets is as follows:

[0122] Use the YOLOv5 model to detect the position information of target pedestrians in the target data of the current frame to obtain the predicted target bounding box.

[0123] Extract features from the detected predicted target bounding box, determine the confidence of the detection object, perform target tracking from the pedestrian target dataset Ptd, and match the trajectory where the target detected in the current frame is located.

[0124] Use the Hungarian algorithm to match the trajectory where the detected target is located, and use Kalman filtering to predict and update the position of each trajectory of multiple target pedestrians in the next frame. For each matched target and detection pair, use the update step of Kalman filtering again to update the new state of the target to improve the accuracy of data association.

[0125] According to the Kalman filter smoothing trajectory technology, by predicting the current state and updating the state based on new observation data, the motion trajectory is obtained, and then the motion trajectory information H(Ptd|t0 + 1) of multiple targets is obtained. n , where H(Ptd|t0 + 1) represents the motion trajectory information of the detected target, and n represents the number of detected targets.

[0126] S4, Pedestrian Re-identification: The ReID technology is used to monitor and process the target data captured by different cameras in public scenarios, and then re-identify the target pedestrians with the multi-object motion trajectory information. The Deep-SORT algorithm is combined to determine the trajectory of the associated target pedestrians in the multi-object tracking information, and trajectory association data is obtained to avoid ID jumps;

[0127] Specifically, the pedestrian re-identification steps of the ReID technology are as follows:

[0128] A1, Feature Extraction: A convolutional neural network model is used to learn and extract the discriminative behavioral feature information of multi-object pedestrians from the pedestrian target dataset;

[0129] A2, Feature Matching: By calculating the distance or similarity between features, the pedestrian features extracted from the current target data are matched with the pedestrian feature library extracted from the known pedestrian target dataset;

[0130] A3, Trajectory Association: According to the results of feature matching, the Hungarian algorithm is used to associate the target pedestrians of the target data with the motion trajectory information of the corresponding multi-objects in the known pedestrian target dataset, and the best associated trajectory information is found.

[0131] Specifically, the acquisition logic of the trajectory association data is as follows:

[0132] Using the convolutional neural network CNN model to learn and extract the discriminative behavioral feature information from the pedestrian target dataset Ptd, calibrated as f(Ptd), then the discriminative behavioral feature information is calculated by the convolutional layer formula in the CNN model as f(Ptd) = ReLU(W * Ptd + β0), where ReLU represents the activation function, W represents the convolutional kernel weight, β0 represents the bias term, and * represents the convolutional operation;

[0133] Using the Euclidean distance to calculate the distance between two feature vectors for feature matching, generating an associated distance matrix, calibrated as Df i′j′ , then two feature vectors are set according to the behavioral feature information f(Ptd) In the formula, i′ and j′ respectively represent the detected target feature and the trajectory feature of the target pedestrian, and are positive integers where 1 ≤ i′, j′ ≤ n. The calculation formula of the Euclidean distance is And using the cosine similarity to calculate two feature vectors The formula for the similarity is In the formula, Sf i′j′ represents the cosine similarity, · represents the dot product of vectors, and ||×|| represents the Euclidean norm of vectors;

[0134] According to the generated associated distance matrix Df i′j′, a matching algorithm such as the Hungarian algorithm is used to associate the behavioral feature information f(Ptd) of the target pedestrian and the motion trajectory information H(Ptd|t0+1) of multiple targets n , and associate the cost matrix and find the optimal matching to obtain trajectory association data, calibrated as rg n , then the formula for the association degree is and rg n ∈[0,1]. In the formula, rg n represents the association between the behavioral feature information f(Ptd) of the target pedestrian and the motion trajectory information H(Ptd|t0+1) of multiple targets n of. represents the minimum distance cost between the detected target feature and the trajectory feature vector of the target pedestrian. n represents the number of detected targets. Then, the association between each other is calculated according to the detected target and the corresponding trajectory, and the optimal trajectory matching result is obtained;

[0135] To avoid ID jumps, Deep-SORT does not immediately delete the trajectory when the trajectory matching fails. Instead, a timing counter is used to track the number of frames of each trajectory since the last successful match. If the trajectory of the target pedestrian to be detected still cannot find a match within the predetermined timing trajectory threshold, the trajectory of the target pedestrian to be detected will be deleted. At the same time, when the matching cost of the newly detected target pedestrian trajectory is too high, it is initialized as a new trajectory.

[0136] S5. Loitering detection: After obtaining the motion trajectory information of multiple targets and the trajectory association data, a determination method based on the distance of the motion trajectory is adopted, and the loitering of the target pedestrian is determined according to the displacement and the distance to determine whether the target person has a lingering behavior, and a loitering detection result is generated;

[0137] Specifically, the acquisition logic of the loitering detection result is as follows:

[0138] Through the motion trajectory information H(Ptd|t0+1) of multiple targets n and the trajectory association data rg n , a distance and displacement calculation method is adopted, and whether the target pedestrian stays or loiters in an area is determined according to the ratio of the compared distance and displacement;

[0139] Among them, the best motion trajectory information of multiple target pedestrians detected and matched in the public area is In the formula, represents the best motion trajectory information of multiple target pedestrians. x″ and y″ respectively represent the coordinates of the target pedestrian on the x-axis and y-axis in the plane position of the best motion trajectory information. t represents the time period during which the best motion trajectory information lasts. n3 represents the number of detected targets;

[0140] Arbitrarily select the best motion trajectory information of a target pedestrian and find the starting point to the ending point from the trajectory of the target pedestrian, which are respectively marked as Bti(x s ″, y s ″) t , Bti(x e ″, y e ″) t , and count the trajectory coordinates of each of the M points where the target pedestrian continuously moves on the optimal motion trajectory, which are marked as Bti(x″ M , y″ M ) t ;

[0141] Then calculate the distance and displacement of the target pedestrian respectively according to the starting point to the ending point of the target pedestrian's trajectory. The displacement calculation formula is

[0142] The distance calculation formula is and M ≥ 2;

[0143] According to the calculation result of the ratio of the distance and the displacement is Dt ≠ 0, Pp ∈ [0, 1]. In the formula, Pp represents the ratio of the distance and the displacement, and set a threshold for determining whether the target pedestrian stays or wanders in an area, which is marked as Wt ∈ [0, 0.75]. In the formula, Wt represents the behavior of staying or wandering in an area. Therefore, compare the magnitudes of the ratio and the threshold. When Pp > Wt, Pp ∈ (0.75, 1], it is determined that the target pedestrian does not stay or wander in a public area; when Pp ≤ Wt, Pp ∈ [0, 0.75], it is determined that the target pedestrian stays or wanders in a public area.

[0144] S6. Detection result: When the wandering detection result determines that the target pedestrian is in a wandering state, immediately trigger an alarm signal, and communicate and transmit the detected abnormal behavior and trajectory of the target pedestrian in the form of images, video clips, and alarm information, and feedback and report to the security personnel in public places, and take response measures such as dispersing, sorting, and necessary early warnings in a timely manner.

[0145] Specifically, the ways of the alarm signal include audible and visual alarms, text message and email notification alarms, and intelligent monitoring program push alarms, and the content of the alarm signal includes the occurrence time, location, duration of the target pedestrian's stay or wander, and the trajectory image of the target pedestrian.

[0146] Specifically, the content of communication transmission and feedback report includes using the existing monitoring camera system to directly transmit the real-time video stream of the captured loitering behavior to the monitoring center or the mobile devices of security personnel; intercepting and saving the images or video clips of the detected loitering behavior through automated software and sending them to security personnel via the network; and generating an alarm report of the details of the loitering event and sending it to relevant security personnel via email or the internal communication network. The security personnel are required to verify the content of the alarm report, conduct on-site inspections, and take on-site responses and contingency plan processing.

[0147] For the method for detecting the staying behavior of people in a public place provided by an embodiment of the present invention, the specific method and process are as detailed in the embodiment of the method for detecting the staying behavior of people in a public place described above, and will not be elaborated here.

[0148] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0149] As described above, the above are only specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

Claims

1. A method for detecting the staying behavior of people in public places, characterized in that, The steps are as follows: S1. Original data: Cameras are installed in public places to continuously capture video streams, obtaining original video data, which is then transmitted over the network to a preprocessing center. The original video data is encrypted and stored and access-controlled to obtain the preprocessed target data; The preprocessing steps for the original video data are as follows: By connecting to several cameras in public places, continuously capturing video streams, obtaining original video data, calibrated as Ovd, and where x and y respectively represent random variables in the video stream pixels, n1 represents the set of pixel variables in the continuous video stream, and is transmitted to the preprocessing center through the network; Extract individual frames from the video stream and use a Gaussian filter to denoise and adjust the resolution of each frame. The denoising calculation formula is In the formula, G(Ovd) represents the Gaussian function of the original video data Ovd, and σ G represents the standard deviation of the Gaussian distribution of the random variables x and y. The resolution adjustment calculation formula is In the formula, P(x′, y′) represents the coordinates of the new pixel point after adjusting the resolution of P(x, y), and P(x, y) represents the coordinates of the source pixel point in the original video data Ovd. P(x + i, y + j) represents the neighboring pixel at P(x, y). i and j respectively represent the neighborhood ranges around P(x, y), and w i,j represents the weight calculated by weighted averaging based on the distance between the point P(x, y) and the point P(x + i, y + j); Before transmission and storage, the original video data Ovd is encrypted using the AES algorithm to obtain an encrypted data signal, designated as Eds. The encryption calculation formula is Eds = AES(Ovd + K0), and In the formula, AES represents the encryption function, Ovd represents the original video data, and K0 represents the key. They respectively represent the randomly read pixel variables after encryption. It represents the set of pixel variables after encryption; Meanwhile, access control policies are implemented for personnel who can process data signals, ensuring that only authorized personnel can access and process data signals; For the encrypted data signal Eds, in order to facilitate transmission or storage, the encrypted data format is standardized, that is, hexadecimal encoding or Base64 encoding is used to process and change the representation of the data for storage and transmission, so as to obtain the data to be encrypted and processed, designated as Tds, and wherein, X and Y respectively represent the data formats obtained by re-standardizing and encoding the encrypted data signal Eds, and N1 represents a number of target data sets; and applying background subtraction technology to highlight moving objects, that is, using the background difference algorithm to subtract the background model from the current frame to detect moving objects in the original video data. At the moment t when the moving object in the current frame is located, the pixel value at the position (x t ,y t ) is P′(x t ,y t ), then the background difference calculation formula is D′(x t ,y t ) = [P′(x t ,y t ) - B′(x, y)], where D′(x t ,y t ) represents the absolute difference between the pixel at the position (x t ,y t ) in the current frame at time t and the pixel of the background frame, B′(x, y) represents the pixel value at any position in the background model, and x t 、y t respectively represent the pixel values of the position where the moving object is located at time t read from the target data Tds; S2. Target detection: The YOLOv5 model is used to detect, identify, screen, and label pedestrians in the target data, generating a pedestrian target dataset; The acquisition logic for the pedestrian target dataset is as follows: The YOLOv5 model receives the input target data Md3(x 0 , y 0 ), and performs feature extraction through convolutional layers, residual layers, downsampling layers, upsampling layers, and routing layers. For each convolutional layer, a convolutional kernel is used for feature extraction, and the output target feature map is calibrated as TS(x 0 , y 0 ). The convolutional calculation formula is In the formula, TS(x 0 , y 0 ) represents the output feature map of the convolutional layer at the position (x 0 , y 0 ), K represents the convolutional kernel, Md3(x 0 + i0, y 0 + j0, k0) represents the operation input value of the target data Md3(x 0 , y 0 ) on the convolutional kernel K, K(i0, j0, k0) represents the value of the convolutional kernel K at the position (x 0 , y 0 ), i0, j0, and k0 respectively represent three different points x 0 + i0, y 0 + j0, k0 in the convolutional kernel K, x 0 + i0, y 0 + j0, k0 represent three different weight parameters in the convolutional kernel K, and b0 represents the bias term of the convolutional layer; After performing convolution calculation on the target feature map TS(x 0 , y 0 ), batch normalization is carried out, and the calculation formula for batch normalization is And In the formula,[[]] represents the result of performing k-batch normalization on the input target feature map TS(x 0 , y 0 ), k 0 represents k batches, represents the input target feature map TS(x 0 , y 0 ) in k batches, B(k 0 ) represents the batch mean, ε represents a learnable parameter, represents the batch variance, c′ represents a non-zero positive number for controlling the scaling ratio, Y(k 0 ) represents the predicted output value in the k-th batch, λ represents a hyperparameter of the scaling factor, and b′ represents an offset translation term for adjusting the predicted value; The feature map output by batch normalization in the YOLOv5 model will be used for prediction training to generate the predicted output value Y(k 0 ), the corresponding confidence P(Y), where P(Y) ∈ [0, 1], and based on the predicted output value Y(k 0 ) of the feature map, the prediction box information is generated, and the prediction box information includes the class probability, box position, and the width and height of the prediction box, which are calibrated as PC, BP, and WH respectively, that is, Y(k 0 ) = {PC, BP, WH}; The class probability PC is obtained by using the softmax function to predict the probability distribution of each anchor box belonging to each class. The calculation formula for the class probability is set as and 0 ≤ m, n2 ≤ k. In the formula, PC represents the score of the target feature predicted by the YOLOv5 model, and Y(k 0 ) m represents the feature vector of the predicted output value Y(k 0 ) in the current m-th class, represents the feature vector of k classes summation, and n2 represents the n2-th class among a total of k classes; The box position BP adopts the anchor box mechanism, outputs the offsets related to the anchor box through the Sigmoid function, and determines the final boundary box position based on the offsets, the center coordinates of the original bounding box, and the width and height WH of the predicted box. Then, according to the predicted output value Y(k 0 ), the center coordinate values of the predicted bounding box are set as (Yx, Yy), the width and height of the predicted bounding box are Yw and Yh respectively, and the upper left coordinates of the grid cell for image segmentation are known as (Gx, Gy), and the width and height are Gw and Gh respectively. Then, the offset formula calculated using the Sigmoid function is dx = σ(Yx) + Gx, dy = σ(Yy) + Gy, and dw = Gw × e Yw , dh = Gh × e Yh . In the formula, dx and dy respectively represent the magnitudes of the offsets in the x-axis and y-axis directions, dw and dh represent the width and height of the final bounding box position, σ represents the hyperparameter of the center coordinate values Yx and Yy of the predicted bounding box, and e represents the exponential function with the natural constant e as the base. That is, the box position BP = {dx, dy, dw, dh}, and the predicted box width and height WH = {dw, dh}; When the YOLOv5 model is training, it adjusts the network parameters through the forward propagation algorithm, uses the loss function to minimize the gap between the detection results and the actual annotations, does not adopt non-maximum suppression (NMS) to remove redundant overlapping boxes, retains several detection results of each pedestrian for position fusion processing, and after training, performs pedestrian detection on the given new target data to generate a dataset that only contains pedestrian targets, namely the pedestrian target dataset, calibrated as Ptd. Among them, the calculation of the loss function and the method of position fusion processing are crucial. The loss function includes coordinate loss, object confidence loss, and class loss. The calculation formula of the coordinate loss is In the formula, L(CIoU) represents the loss value calculated by the CIoU loss function, and IoU represents the intersection over union represents the Euclidean distance between the predicted box position BP and the center point of the true box (Gx, Gy) of the grid cell represents the diagonal length of the minimum closed area of the predicted box and the true box, and α represents the weight coefficient represents the consistency of the aspect ratio of the predicted box and the true box The calculation formula for object confidence loss is L(BCE) = -P0ln(P(Y)) - (1 - P0)×(1 - ln(P(Y))), where L(BCE) represents the object confidence loss value calculated using the binary cross-entropy loss algorithm BCE, P0 represents the true label of the detected person, and when the detected person exists, P0 = 1, when the detected person does not exist, P0 = 0, and P(Y) represents the detection object confidence output by the YOLOv5 model prediction result, and P(Y) ∈ [0, 1]; The calculation formula for the class loss is In the formula, L(PC) represents the difference loss value between the predicted class distribution and the true class distribution calculated using the cross-entropy loss algorithm, k represents the total number of classes, p(Y k ) represents the true probability that the detection object belongs to class k, and p(PC) k represents the predicted probability that the detection object belongs to class k; The position fusion process uses the average fusion algorithm to obtain the final bounding box coordinates of the detected target pedestrian and constructs the pedestrian target dataset Ptd. The calculation formula of the average fusion algorithm is And In the formula,[[]]END]] And respectively represent the average values of the x-axis coordinate, y-axis coordinate, width, and height of the center point after the detection box fusion. represents the confidence level of the fusion box, k represents the total number of categories, N2 represents the number of detected bounding boxes, Yx i , Yy i , Yw i , Yh i respectively represent the coordinate positions and width and height of each bounding box. S3. Target tracking: The Deep-SORT algorithm combines Kalman filtering and the Hungarian algorithm to perform cascade matching and trajectory prediction on the target pedestrians in the pedestrian target dataset, achieving multi-target tracking and obtaining the motion trajectory information of multiple targets; S4. Pedestrian re-identification: After using ReID technology to monitor and process the target data captured by different cameras in a public scene, pedestrian re-identification of the target pedestrians is performed with the motion trajectory information of multiple targets, and the Deep-SORT algorithm is combined to determine the trajectory of the associated target pedestrians in the multi-target tracking information, obtaining trajectory association data to avoid ID jumps; The acquisition logic for the trajectory association data is as follows: The convolutional neural network CNN model is used to learn and extract discriminative behavioral feature information from the pedestrian target dataset Ptd, denoted as f(Ptd). The discriminative behavioral feature information is calculated using the convolutional layer in the CNN model as f(Ptd) = ReLU(W*Ptd + β0), where ReLU represents the activation function, W represents the convolutional kernel weight, β0 represents the bias term, and * represents the convolution operation; Feature matching is performed by calculating the distance between two feature vectors using the Euclidean distance, generating an associated distance matrix, labeled as Df i′j′ , then two feature vectors are set according to the behavioral feature information f(Ptd) In the formula, i' and j' respectively represent the trajectory features of the detected target feature and the target pedestrian, and are positive integers where 1 ≤ i', j' ≤ n. The calculation formula for the Euclidean distance is According to the distance matrix Df that generates the association i′j′ , use a matching algorithm such as the Hungarian algorithm to associate the behavioral feature information f(Ptd) of the target pedestrian and the motion trajectory information H(Ptd|t0+1) of multiple targets n , and associate the cost matrix and find the optimal match to obtain the trajectory association data, calibrated as rg n , then the association degree calculation formula is and rg n ∈[0,1]. In the formula, rg n represents the association between the behavioral feature information f(Ptd) of the target pedestrian and the motion trajectory information H(Ptd|t0+1) of multiple targets n of represents the minimum distance cost between the detected target feature and the trajectory feature vector of the target pedestrian. n represents the number of detected targets. Then, calculate the association between each other according to the detected target and the corresponding trajectory, and obtain the optimal trajectory matching result; To avoid ID jumps, Deep-SORT does not immediately delete a trajectory when the trajectory matching fails, but instead uses a time counter to track the number of frames since the last successful match for each trajectory; S5. Loitering detection: After obtaining the motion trajectory information of multiple targets and the trajectory association data, a determination method based on the motion trajectory distance is used to determine whether a target pedestrian is loitering according to displacement and distance, generating a loitering detection result; The acquisition logic for the loitering detection result is as follows: Through the multi-target motion trajectory information H(Ptd|t0+1) n and the trajectory correlation data rg n , the distance and displacement calculation method is adopted, and whether the target pedestrian stays or lingers in a region is determined according to the ratio of the compared distance and displacement; Among them, the best motion trajectory information for detecting and matching multiple target pedestrians in the public area is In the formula,[[]]END]] represents the best motion trajectory information of multiple target pedestrians. x″ and y″ respectively represent the coordinates of the target pedestrians on the x-axis and y-axis in the plane position within the best motion trajectory information. t represents the time period during which the best motion trajectory information persists. n3 represents the number of detected targets; Arbitrarily select the best motion trajectory information of a multi-target pedestrian and find the start point to the end point from the trajectory where the target pedestrian is located, and respectively calibrate them as Bti(x″ s ,y″ s ) t , Bti(x″ e ,y″ e ) t , and count the trajectory coordinates of each of the M points where the target pedestrian continuously moves on the best motion trajectory, and calibrate them as Bti(x″″,y″ M ) t ; Then, calculate the distance and displacement of the target pedestrian from the starting point to the ending point of the target pedestrian's trajectory. The displacement calculation formula is The distance calculation formula is and M ≥ 2; The calculation result according to the ratio of the distance and the displacement is In the formula, Pp represents the ratio of the distance and the displacement, and a threshold is set for determining whether the target pedestrian stays or lingers in an area, calibrated as Wt ∈ [0, 0.75]. In the formula, Wt represents the behavior of staying or lingering in an area. Therefore, by comparing the magnitudes of the ratio and the threshold, when Pp > Wt and Pp ∈ (0.75, 1], it is determined that the target pedestrian does not stay or linger in a public area; when Pp ≤ Wt and Pp ∈ [0, 0.75], it is determined that the target pedestrian stays or lingers in a public area; S6. Detection Results: When the wandering detection result determines that the target pedestrian is in a wandering state, an alarm signal is immediately triggered, and the abnormal behaviors and trajectories of the detected target pedestrian are communicated and transmitted in the form of images, video clips, and alarm information, and reported back to the security personnel in public places, and response measures such as dispersing, sorting, and necessary early warnings are taken in a timely manner.

2. The method for detecting the staying behavior of people in a public place according to claim 1, characterized in that, The acquisition logic of the target data is as follows: Read the encrypted data to be processed Tds, and perform inverse algorithm decryption using the AES key. The AES decryption calculation formula is In the formula, Ovd represents the original video data, and AES -1 represents the inverse algorithm of the encryption function, Tds represents the ciphertext data, represents the symmetric key; Identify the original video data Ovd, and perform step-by-step processing calculations of formatting, scaling, and normalization on the original video data Ovd to obtain target data, calibrated as Md3. Among them, the formatting process automatically converts the color space of the operable frame sequence in the original video data Ovd by using a video processing library, that is, converts from BGR to RGB format, to obtain formatted data, calibrated as Fd1, and Fd1 = {x1, y1} n , where x1 and y1 respectively represent the formatting results of variables x and y; The scaling process uses the bilinear interpolation algorithm to scale each frame of the formatted data Fd1 to the resolution, obtaining the scaled data, calibrated as Sd2. The bilinear interpolation calculation formula is as follows In the formula, Sd2(x1′, y1′) represents the pixel value at the positions of x1′ and y1′. a, b, c, and d respectively represent the proximity of the relative positions between the target point of four adjacent pixel values in the image and the original pixel point. x1′ and y1′ represent the positions where the points x1 and y1 in the original coordinate system are mapped to the new coordinate system; Normalization is to normalize each pixel value range [0, 255] in the image scaling data Sd2 to the range [0, 1] to obtain the normalized data, that is, the final target data Md3. The calculation formula for normalization is In the formula, Md3(x 0 , y 0 ) represents the value after normalization. x 0 , y 0 respectively represent the normalization results of the positions x1′, y1′ in the new coordinate system. Sd2(x1′, y1′) n represents the pixel value at the position (x1′, y1′) after the original image is scaled.

3. The method for detecting the staying behavior of people in a public place according to claim 2, wherein, The target detection steps of the YOLOv5 model are as follows: Input the target data Md3(x 0 , y 0 ) into the YOLOv5 network where the YOLOv5 model is located; The convolutional layer, residual layer, downsampling layer, upsampling layer, and routing layer included in the YOLOv5 network extract features from the target data Md3 to obtain a target feature map, and the target feature map includes the prediction of the bounding box and the classification confidence. The YOLOv5 model divides the image of the target feature map into grid cells, and based on the training and forward inference of the YOLOv5 model, the bounding boxes can be detected, and multiple bounding boxes and corresponding confidence scores are predicted on each grid cell. Then, by using non-maximum suppression (NMS) to compare the intersection over union (IoU) threshold, the overlapping bounding boxes are removed to ensure that each real target has only one corresponding bounding box. Filter out the detection results with a confidence higher than the set threshold and classify the moving pedestrian targets. Parse the finally retained bounding boxes and classification results, and output and store the detected pedestrian positions and class labels.

4. The method for detecting the staying behavior of people in a public place according to claim 3, characterized in that, The multi-object tracking steps of the Deep-SORT algorithm for the pedestrian target dataset in combination with the Kalman filter and the Hungarian algorithm are as follows: The Deep-SORT algorithm realizes the efficient and accurate tracking of multiple targets in the pedestrian target dataset by comprehensively using the Kalman filter prediction and the data association of the Hungarian algorithm. According to the pedestrian target dataset Ptd, the pedestrians in the detection image are tracked for behavior targets. The Kalman filter is used to predict the position and velocity state of the target at time t0 within the current frame. The prediction calculation formula of the Kalman filter is And In the formula, H(Ptd|t0) represents the state prediction value of the pedestrian target dataset Ptd at the current time t0. Represents the state transition matrix of the pedestrian target dataset Ptd, and H(Ptd|t0 - 1) represents the state estimation value of the pedestrian target dataset Ptd at the previous time t0 - 1. Represents the control input matrix of the pedestrian target dataset Ptd. Represents the control input at time t0, P(t0) represents the predicted state covariance at time t0, P(t0 - 1) represents the state covariance matrix at the previous time t0 - 1, T represents the transpose. Represents the process noise covariance matrix at time t0. The target state H(Ptd|t0) predicted by the Kalman filter and the detection Ptd observed at the current time t0 t0 , and the Hungarian algorithm is used to assign the trajectory where the detection target is located. The trajectory calculation formula of the Hungarian algorithm is In the formula,[[]] represents the total distance between the target state H(Ptd|t0) and the detection Ptd t0 , λ represents the parameter for balancing the distance,[[]] represents the matrix product,[[]] represents the Mahalanobis distance predicted based on the Kalman filter,[[]] represents the predicted value based on the cosine similarity between the target state H(Ptd|t0) and the detection ; Again, for each matched target and detection pair, use the update step of the Kalman filter to update the new state of the target position and velocity. The calculation formula for updating the new state of the Kalman filter is K(t0) = P(t0 - 1) × H(t0) T × (H(t0) × P(t0 - 1) × H(t0) T + R(t0)) -1 , and H(Ptd|t0 + 1) = H(Ptd|t0) + K(t0) × (B(Ptd) t0 - H(t0) × H(Ptd|t0)), P(t0 + 1) = (1 - K(t0) × H(t0)) × P(t0). In the formula, K(t0) represents the Kalman gain, H(t0) represents the observation model, R(t0) represents the observation noise covariance, P(t0 - 1) represents the state covariance matrix at the previous moment t0 - 1, T represents the transpose, H(Ptd|t0 + 1) represents the state prediction value of the pedestrian target dataset Ptd at the next moment t0 + 1, H(Ptd|t0) represents the state prediction value of the pedestrian target dataset Ptd at the current moment t0, B(Ptd) t0 represents the control input matrix of the pedestrian target dataset Ptd, that is, the observed value of the pedestrian target dataset Ptd at the current moment t0, P(t0 + 1) represents the updated state probability at the next moment t0 + 1, and P(t0) represents the predicted state covariance at the moment t0; The position change of each tracking target over time is recorded to form a motion trajectory, and the motion trajectory of the target is optimized by predicting the current state and updating the state based on the new observation data to achieve trajectory smoothing.

5. A method for detecting the staying behavior of people in a public place according to claim 4, characterized in that, The acquisition logic of the motion trajectory information of the multiple targets is as follows: Use the YOLOv5 model to detect the position information of the target pedestrian in the target data of the current frame to obtain the predicted target bounding box. Extract features from the detected predicted target bounding box, determine the confidence of the detection object, perform target tracking from the pedestrian target dataset Ptd, and match the trajectory where the target detected in the current frame is located. Use the Hungarian algorithm to match the trajectory where the detected target is located, use the Kalman filter to predict and update the position of each trajectory of the multiple target pedestrians in the next frame, and for each matched target and detection pair, use the update step of the Kalman filter again to update the new state of the target to improve the accuracy of data association. According to the Kalman filter smoothing trajectory technology, the motion trajectory is obtained by predicting the current state and updating the state based on new observation data, and then the motion trajectory information H(Ptd|t0+1) of multiple targets is obtained. n , where H(Ptd|t0+1) represents the motion trajectory information of the detected target, and n represents the number of detected targets.

6. The method for detecting the staying behavior of personnel in a public place according to claim 5, characterized in that, The pedestrian re-identification steps of the ReID technology are as follows: A1. Feature Extraction: Use a convolutional neural network model to learn and extract the discriminative behavioral feature information of multiple target pedestrians from the pedestrian target dataset. A2. Feature Matching: By calculating the distance or similarity between features, match the pedestrian features extracted from the current target data with the pedestrian feature library extracted from the known pedestrian target dataset. A3. Trajectory Association: Based on the result of feature matching, the Hungarian algorithm is used to associate the target pedestrians in the target data with the motion trajectory information of multiple targets corresponding to the known pedestrian target dataset, and the optimal associated trajectory information is found.

Citation Information

Patent Citations

  • Construction site pedestrian wandering detection method, device and equipment and storage medium

    CN111860318A

  • Lightweight pedestrian tracking method in complex scene

    CN115984969A