Abnormality detection method for hidden social security behaviors
By introducing regional concealment and observability and using Transformer autoencoder to learn related sequences, the problem of difficulty in detecting hidden social security behavior in the existing technology is solved, and more accurate identification of latency behaviors and improved the effectiveness of social security monitoring.
Patent Information
- Application Number
- CN202510144716.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The prior art is difficult to detect hidden social security behaviors because these behaviors are not abnormal in terms of speed, posture, etc., and are difficult to detect through existing actions and timing information.
Regional concealment and observability are introduced, and the region number sequence, concealment sequence and observability sequence are learned by constructing a Transformer autoencoder, and abnormal scores are calculated to determine hidden social security behavior.
Through the integration of social attribute analysis, social security behaviors during incubation can be detected more accurately, the ability to identify hidden behaviors is improved, and the effectiveness of social security monitoring is enhanced.
Smart Images

Figure CN120071249A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of social security monitoring and relates to an abnormal detection method for concealed social security behaviors. Background Art
[0002] In important places in modern cities, especially banks, the security risks are severe. According to data from the Federal Bureau of Investigation of the United States, there were 1,740 bank robberies in the United States in 2022, causing huge economic losses and threats to personal safety. In order to effectively prevent crimes and reduce social security hazards, a large number of intelligent surveillance cameras and monitoring systems have been widely deployed in urban areas, and key places are the top priorities. At the same time, driven by video analysis and artificial intelligence, video abnormal detection technology has developed rapidly.
[0003] Those social security dangerous persons, such as bank robbers, will repeatedly scout before committing a crime, observing the positions of cameras, the situation of guard patrols, etc., and carefully planning the crime. The behaviors during this period are concealed social security behaviors. If they can be accurately detected, early warnings can be given to ensure social security. Currently, there are already a variety of related technical solutions. Morais et al. proposed an abnormal behavior detection method based on skeleton feature prediction. First, the pedestrian skeleton trajectory is extracted and decomposed into two parts: the global trajectory and the local body posture. Then, MPED-RNN is used for modeling, and the abnormality is judged by measuring the difference between the prediction and the actual situation. However, there are many skeleton trajectory points, and the algorithm has a large computational amount. The abnormal behavior detection model based on multiple time scales by Rodrigues et al. is a multi-layer structure. Each layer learns the distribution law of normal posture trajectories at a specific time scale. When detecting, the prediction results of each layer are combined to judge the abnormality, but it is difficult to accurately balance the prediction results of each layer. The abnormal trajectory detection network designed by Sun et al. uses the YOLOv5 and Deep-Sort algorithms to extract the target trajectory, obtains the trajectory abnormal score through a series of processes, and also designs a speed calculation module to obtain the speed abnormal score, and comprehensively judges whether the target is abnormal based on the two. Georgescu et al. proposed a video abnormal detection framework based on adversarial training, which extracts the forward movement, backward movement and appearance information of the target, and judges the abnormality by reconstructing through an encoder and fusing the reconstruction error.
[0004] However, the detection inputs of the existing technologies are all pedestrian postures and trajectory appearances. These motion information and timing information can only detect abnormalities in terms of actions and speeds, that is, significant abnormal behaviors. However, concealed social security behaviors have no abnormalities in terms of speed, posture, etc., so it is difficult to be detected. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides an anomaly detection method for hidden social security behaviors, aiming to introduce regional concealment to measure the concealment of the activity space and introduce regional observability to measure the observability of the activity space, and accordingly establish a detection and early warning model for casing behaviors. Compared with the existing spatio-temporal analysis methods, the model introduces social attribute analysis, and fuses and analyzes social features and spatio-temporal features through regional concealment and regional observability, exploring a new way for the early warning research of latent social security behaviors.
[0006] The object of the present invention can be achieved by the following technical solutions:
[0007] The present application provides an anomaly detection method for hidden social security behaviors, including the following steps:
[0008] S1. Obtain the external surveillance video stream data of the target building for a period of time as the training set, and use the trajectory monitoring algorithm to extract the pedestrian trajectories in the surveillance video stream;
[0009] S2. Divide each frame of the video stream image into grids, assign a separate ID to each grid area, and generate a sequence composed of ID numbers in the order of the grid areas passed by the pedestrian trajectory, denoted as the area number sequence;
[0010] S3. Define the concealment of each grid area according to the ratio of the number of pedestrians passing through each grid area to the total number of pedestrians in all grid areas;
[0011] S4. Generate a sequence composed of concealments in the order of the grid areas passed by the pedestrian trajectory, denoted as the concealment sequence;
[0012] S5. Take the center of each grid area as the observation point, extend the line of sight from the observation point to the target building, and obtain the visible length of the target building;
[0013] S6. Calculate the ratio of the visible length of the target building to the total length to obtain the observability of the observation point to the target monitoring;
[0014] S7. Generate a sequence composed of observabilities in the order of the grid areas passed by the pedestrian trajectory, denoted as the observability sequence;
[0015] S8. Construct a Transformer autoencoder to learn the area number sequence, concealment sequence and observability sequence;
[0016] S9. Obtain the surveillance video stream to be detected of the target building, and use the area number sequence, concealment sequence and observability sequence of the pedestrian trajectory of the video to be detected as the input sequence to input into the Transformer autoencoder to obtain the reconstructed sequence;
[0017] S10. Calculate the anomaly score based on the error between the reconstructed sequence and the input sequence. When the anomaly score exceeds a preset threshold, it is determined that there is a hidden social security behavior.
[0018] Further, the step of extracting the pedestrian trajectory in the surveillance video stream by using the trajectory monitoring algorithm includes the following steps:
[0019] S11. Video stream data reading and preprocessing: Read the surveillance video stream data outside the target building and load it into the processing system. Subsequently, perform preprocessing on each frame of the image, enhance the clarity through image enhancement, and reduce noise interference through denoising processing;
[0020] S12. Pedestrian target detection: Load the pre-trained pedestrian detection model, scan each frame of the video stream image, identify the pedestrian area and output its position information and confidence score, and determine the pedestrian target accordingly;
[0021] S13. Initialization of pedestrian tracking: Set a threshold according to the confidence score output by the pedestrian detection to screen effective pedestrian targets, and then initialize the tracker for the effective pedestrian targets. Use the tracker to predict the current frame position based on the previous frame state of the pedestrian, and start the subsequent tracking process;
[0022] S14. Pedestrian tracking and trajectory update: In the subsequent frames, the tracker continuously tracks the initialized pedestrian targets, updates the tracking state according to the matching situation between the actual position and the predicted position of the pedestrian, and records the position of the pedestrian in each frame at the same time to form and continuously improve the pedestrian trajectory information.
[0023] Further, the concealment degree is calculated by the formula:
[0024]
[0025] In the formula, H i is the concealment degree of the i-th grid area; n i is the number of pedestrians passing through the i-th grid area; N is the total number of pedestrians in all grid areas.
[0026] Further, the observability is calculated by the formula:
[0027]
[0028] In the formula, V i is the observability of the target building; Visiblelength is the visible length of the target building; Totallength is the total length of the target building.
[0029] Further, the Transformer autoencoder includes the following construction steps:
[0030] S81. Determine input, output and model architecture: Use the region number sequence, concealment sequence and observability sequence as model input, and the original input sequence as the expected output. Select the Transformer architecture to build an autoencoder.
[0031] S82. Build the encoder: first create an input embedding layer to convert various input sequences into vector representations; build a multi-head attention layer and set multiple attention heads to capture the complex relationships between sequence elements from multiple angles; add a feedforward neural network layer to perform nonlinear transformation to extract high-level features; finally, perform layer normalization and residual connection operations after the multi-head attention layer and feedforward neural network layer respectively;
[0032] S83, build decoder: set up input embedding layer to process initial input of decoder; set up masked multi-head attention layer, use masking technology to decode in sequence; add multi-head attention layer to interact with encoder to use encoder features; set up feedforward neural network layer to complete sequence reconstruction task, and perform layer normalization and residual connection;
[0033] S84, set the loss function and optimizer: select the loss function to measure the difference between the decoder restored sequence and the original input sequence, and minimize the difference as the training goal; select the optimizer to adjust the model parameters according to the gradient update strategy to improve the training efficiency and convergence speed;
[0034] S85, Model training and evaluation: Divide the data set containing related sequences into training set, validation set and test set, use the training set to train the built model, calculate the difference based on the loss function and update the parameters until the predetermined goal is achieved. After the training is completed, use the test set to evaluate the model performance. If it does not meet the standard, adjust the architecture, parameters or hyperparameters and retrain.
[0035] Furthermore, in step S84, the loss function is configured as a mean square error loss function; and the optimizer is configured as stochastic gradient descent.
[0036] Furthermore, in step S10, the calculation formula of the abnormal score is:
[0037]
[0038] Where AS represents the anomaly score; l i is the input concealment sequence, l′ i is the reconstructed concealment sequence; u i is the observable sequence of the input, u′ i is the reconstructed observable sequence; n is the number of elements in the sequence; p y-gramDenotes the probability that all consecutive y IDs in the reconstructed region number sequence are the same as those in the input region number sequence. k is the number of consecutive identical IDs in the reconstructed and input region number sequences. Indicates that the larger the number y of consecutive identical IDs, the larger this value; ω l 、ω u and ω y Are the weight coefficients of the concealment degree sequence, the observability degree sequence, and the region number sequence, respectively.
[0039] Furthermore, the weight coefficients are determined by calculating the correlation coefficients between each sequence in the historical period and whether the known concealed social security behavior occurs.
[0040] Furthermore, in step S10, when it is determined that there is a concealed social security behavior, identify the pedestrians with abnormal behaviors in the input sequence.
[0041] Advantages of the present invention:
[0042] By dividing the grid for the external surveillance video stream image of the target building; generating the region number sequence, the concealment degree sequence, and the observability degree sequence according to the order of the grid areas passed by the pedestrian trajectory; constructing a Transformer autoencoder to learn the region number sequence, the concealment degree sequence, and the observability degree sequence; obtaining the surveillance video stream to be detected of the target building, and taking the region number sequence, the concealment degree sequence, and the observability degree sequence of the pedestrian trajectory of the video to be detected as the input sequence and inputting them into the Transformer autoencoder to obtain the reconstructed sequence; calculating the anomaly score according to the error between the reconstructed sequence and the input sequence, and when the anomaly score exceeds the preset threshold, it is determined that there is a concealed social security behavior. The present invention solves the problem that it is difficult to detect concealed social security behaviors in the prior art. Description of the Drawings
[0043] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the drawings.
[0044] Figure 1 Is the flowchart of an anomaly detection method for concealed social security behaviors in the present invention.
[0045] Figure 2 Is the flowchart of extracting the pedestrian trajectory in the surveillance video stream using the trajectory monitoring algorithm in an embodiment of the present invention.
[0046] Figure 3 Is the flowchart of constructing a Transformer autoencoder in an embodiment of the present invention. Detailed Embodiments
[0047] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features and their effects according to the present invention as follows.
[0048] Please refer to Figures 1 - 3 , this application provides an anomaly detection method for hidden social security behaviors, including the following steps:
[0049] S1. Obtain the external monitoring video stream data of the target building for a period of time as the training set, and use the trajectory monitoring algorithm to extract the pedestrian trajectories in the monitoring video stream;
[0050] In this embodiment, obtain the external monitoring video stream data of the building for a period of time as the training set. First, determine a specific time period to be collected, such as the past week or month, etc., to ensure that this period can cover the personnel activities in different time periods and different weathers. Then, collect all the video stream data recorded by each monitoring camera set around the target building (such as a bank, etc.) during this period. These unprocessed original videos constitute rich basic materials.
[0051] Next, use the trajectory monitoring algorithm to extract the pedestrian trajectories from it. First, apply advanced object detection technology to enable the algorithm to accurately lock the pedestrian individuals in each frame of the video image. This relies on a detection model trained with a large amount of data and having excellent pedestrian feature recognition capabilities, such as the commonly used YOLO series models, etc., which can effectively distinguish pedestrians from other objects even in complex backgrounds.
[0052] After the detection is completed, the pedestrian tracking link is immediately carried out. By assigning a unique identifier to each detected pedestrian, based on the association of features such as position and appearance between consecutive frames, with the help of a tracking algorithm similar to the Kalman filter, predict the current frame position based on the state information of the pedestrian in the previous frame and match and correct it to ensure the continuity of tracking, even if the pedestrian is briefly blocked or the motion state changes, it is not affected.
[0053] Finally, arrange the position information of each tracked pedestrian in different frames in chronological order to generate a complete motion trajectory, which is presented in the form of coordinate points and attached timestamps, speed estimation values, etc. In this way, the pedestrian trajectories in the monitoring video stream are successfully extracted, providing an important basis for subsequent analysis.
[0054] Furthermore, the step of using the trajectory monitoring algorithm to extract the pedestrian trajectories in the monitoring video stream includes the following steps:
[0055] S11. Reading and preprocessing of video stream data: Read the external monitoring video stream data of the target building and load it into the processing system, and then perform preprocessing on each frame of the image, enhance the clarity through image enhancement and reduce the noise interference through denoising processing;
[0056] S12. Pedestrian target detection: Load a pre-trained pedestrian detection model (such as YOLO, Faster R-CNN, etc.), scan each frame of the video stream image, identify the pedestrian area and output its position information and confidence score, and determine the pedestrian target accordingly;
[0057] S13. Initialization of pedestrian tracking: Set a threshold according to the confidence score output by pedestrian detection to screen effective pedestrian targets, then initialize a tracker (such as a Kalman filter tracker, etc.) for the effective pedestrian targets, use the tracker to predict the current frame position based on the previous frame state of the pedestrian, and start the subsequent tracking process;
[0058] S14. Pedestrian tracking and trajectory update: In subsequent frames, the tracker continuously tracks the initialized pedestrian targets, updates the tracking state according to the matching situation between the actual position and the predicted position of the pedestrian, and at the same time records the position of the pedestrian in each frame to form and continuously improve the pedestrian trajectory information.
[0059] S2. Divide each frame of the video stream image into grids, assign a separate ID to each grid area, and generate a sequence composed of ID numbers in the order of the grid areas passed by the pedestrian trajectory, denoted as the area number sequence;
[0060] In this embodiment, in the process of deeply analyzing the external surveillance video stream data of the target building to extract more valuable information, a method of dividing each frame of the video stream image into grids is adopted. Specifically, for each frame of the image in the video stream, it is evenly divided into several grid areas of the same size according to the pre-set rules. These grids are like a dense "net", completely covering the entire image frame, ensuring that each pixel point in the image can belong to a specific grid area. In order to facilitate the accurate tracking and recording of the flow of the pedestrian trajectory between different grid areas in the future, a separate and unique ID number is assigned to each grid area. This ID is like the exclusive "ID card" of each grid area, enabling it to be clearly and accurately identified and distinguished in the entire analysis system.
[0061] When the pedestrian trajectory in the surveillance video stream is successfully extracted by using the trajectory monitoring algorithm, a unique sequence can be generated according to the path the pedestrian walks in the image, in the order of the grid areas passed by its trajectory. This sequence is completely composed of the ID numbers corresponding to the grid areas passed by the pedestrian in sequence, and is specifically denoted as the area number sequence. The area number sequence generated in this way contains rich information. It can show the movement route of the pedestrian between different areas of the video picture in a simple and intuitive way, providing a very important data basis for further analyzing the pedestrian's behavior pattern, activity range, and interaction relationship with the surrounding environment, etc.
[0062] S3. Define the concealment degree of each grid area according to the ratio of the number of pedestrians passing through each grid area to the total number of pedestrians in all grid areas.
[0063] In this embodiment, after making a detailed grid division of the external surveillance video stream image of the target building and obtaining the sequence of area numbers corresponding to the pedestrian trajectories, a specific method is used to define the concealment degree of each grid area. The specific approach is to count the number of pedestrians passing through each grid area, calculate the total number of pedestrians in all grid areas at the same time, and then use the ratio of the number of pedestrians passing through each grid area to the total number of pedestrians in all grid areas to determine the concealment degree of this grid area.
[0064] This concealment degree index is of great significance for identifying hidden social security behaviors. For those who intend to carry out hidden social security behaviors, such as criminals when scouting or preparing to carry out criminal activities, they often tend to choose relatively concealed areas to observe, wait for opportunities or carry out preliminary actions. And the concealment degree of the grid area calculated through this can accurately identify which areas are relatively less concerned in pedestrian activities and have a lower frequency of pedestrians passing by, that is, relatively concealed areas.
[0065] When analyzing the surveillance video data, if it is found that the activity trajectories of certain personnel frequently appear in areas with a higher concealment degree and their behaviors are abnormal, such as staying for a long time, abnormal behaviors, etc., this may imply the risk of hidden social security behaviors. The quantitative analysis of the concealment degree provides an effective means from the perspective of environmental characteristics, assisting in more sensitively detecting those abnormal behaviors that may hide bad intentions among numerous normal pedestrian activities, thereby helping to give early warnings and take corresponding measures to prevent the occurrence of social security incidents.
[0066] Furthermore, the formula for the concealment degree is:
[0067]
[0068] In the formula, H i is the concealment degree of the i-th grid area; n i is the number of pedestrians passing through the i-th grid area; N is the total number of pedestrians in all grid areas.
[0069] It should be noted that, as can be seen from the above formula, the larger the concealment degree of a grid area, the fewer the number of pedestrians passing through.
[0070] S4. Generate a sequence composed of concealment degrees in the order of the grid areas passed by the pedestrian trajectory, denoted as the concealment degree sequence.
[0071] In this embodiment, when pedestrians are active in the external area of the target building, their movement trajectories will pass through different grid areas in sequence. By sorting out the previously recorded pedestrian trajectory information, the concealment value corresponding to each grid area is extracted according to the order in which they pass through each grid area. Then, these concealment values obtained in sequence are arranged in sequence to form a complete concealment sequence. The concealment sequence generated in this way contains rich and valuable information. It can present the changes in the concealment level of the areas experienced by pedestrians during the entire activity process in a continuous and intuitive way. It is of great significance for analyzing pedestrian behavior patterns and identifying possible hidden social security behaviors. For example, if it is found that the concealment sequence of a pedestrian shows abnormal situations such as frequent stays in high-concealment areas or long-term stays in areas with gradually increasing concealment, combined with other behavioral characteristics of the pedestrian, it is possible to infer whether he is suspected of concealing social security behaviors, thereby providing strong data support and decision-making basis for related security monitoring and risk prevention work.
[0072] S5. The center of each grid area is taken as the observation point, and the sight line is extended from the observation point to the target building to obtain the visible length of the target building;
[0073] In this embodiment, when conducting an in-depth analysis of the external environment of the target building, an observation method based on each grid area is adopted to obtain visual information about the target building. Specifically, the center of each grid area is set as an observation point. This observation point is like a specific observation position, from which we can carry out relevant observation operations on the target building.
[0074] Once the observation point is determined, the next step is to extend the line of sight from this observation point to the target building. This process simulates the line of sight when a person is at the center of the grid area looking at the target building. When extending the line of sight, various factors in the actual environment need to be considered, such as whether there are obstructions (such as trees, other buildings, etc.), which may affect the extension of the line of sight and make it impossible to directly see the full picture of the target building.
[0075] As the line of sight extends continuously from the observation point to the target building, we measure and record the length of the part that can clearly see the target building without being blocked. We define this length as the visible length of the target building. This visible length is a relative metric that reflects the effective range size of the target building that can be directly observed starting from the observation point at the center of a specific grid area. By obtaining the visible length of the target building corresponding to each grid area, we can further analyze the visible relationship between different grid areas and the target building, providing an important data basis for subsequent related research such as analyzing the attention of personnel at different positions to the target building and evaluating potential security monitoring fields of view.
[0076] S6. Calculate the ratio of the visible length of the target building to the total length to obtain the observability of the target monitoring at the observation point;
[0077] In this embodiment, by setting the center of each grid area as the observation point and extending the line of sight from this observation point to the target building, the visible length of the target building starting from each observation point is determined. Next, in order to more accurately measure the degree of monitoring of the target building from each observation point, the index of "observability" is introduced. The specific calculation method is that first, the total length of the target building needs to be determined. This total length can be obtained through on-site measurement, building design drawings, or other reliable measurement methods, which represents the overall actual length scale of the target building. Then, the visible length of the target building starting from each observation point obtained previously is used for a ratio operation with the total length of the target building.
[0078] Through such a calculation, the obtained result is the observability of the target building monitoring at the observation point. This ratio can intuitively reflect the degree to which the target building can be observed from a specific observation point. For example, when the value of the observability is 1, it means that the entire picture of the target building can be completely observed from this observation point; and the closer the value of the observability is to 0, it means that the less part of the target building can be observed from this observation point, the greater the influence of factors such as occlusion, and the more limited the monitoring effect on the target building.
[0079] Furthermore, the formula for the observability is:
[0080]
[0081] In the formula, V i is the observability of the target building; Visiblelength is the visible length of the target building; Totallength is the total length of the target building.
[0082] S7. Generate a sequence composed of observabilities in the order of the grid areas passed by the pedestrian trajectory, denoted as the observability sequence;
[0083] In this embodiment, according to the order in which pedestrians pass through these grid regions, the corresponding observability values are extracted in sequence and arranged in an orderly manner, thus forming a sequence composed of observability, which is specifically recorded as the observability sequence. This observability sequence carries rich information and can intuitively and continuously show the dynamic change of the observability of the areas passed by pedestrians during the entire activity process with respect to the target building. It is of great significance in identifying hidden social security behaviors. For example, if it is found that the observability sequence of a certain pedestrian shows extremely low observability and long stay in a specific area, or extremely abrupt changes in observability, combined with other behavioral characteristics of the pedestrian, such as sneaky actions and frequent appearances in hidden areas, it can provide strong clues for inferring whether the pedestrian has hidden social security behaviors, and thus provide important data support and decision-making basis for relevant security monitoring and risk prevention work.
[0084] S8. Construct a Transformer autoencoder to learn the region number sequence, concealment sequence, and observability sequence;
[0085] In this embodiment, the Transformer autoencoder is an advanced deep learning model constructed based on the Transformer architecture, mainly consisting of an encoder and a decoder working together. The core of its principle lies in the attention mechanism. The multi-head attention mechanism in the encoder is the key. It can simultaneously examine each element of the input sequence from multiple perspectives, accurately calculate the correlation weights between elements, and thus deeply capture the complex relationships between elements at different positions in the sequence. For example, for the region number sequence, concealment sequence, and observability sequence we are involved in, it can insight into the internal connections between different grid regions, concealment, and changes in observability. Then, the feed-forward neural network layer will perform a non-linear transformation on the data processed by the attention mechanism to further extract high-level features.
[0086] The work of the decoder is equally important. Its masked multi-head attention mechanism restricts the calculation of attention weights to only seeing the sequence elements before the current position through masking, so as to ensure the sequential reconstruction of the original input sequence step by step and avoid unreasonable reconstruction caused by obtaining future information in advance. And the decoder is also equipped with a feed-forward neural network layer to assist in completing the final sequence reconstruction. In the current context of analyzing hidden social security behaviors, the Transformer autoencoder plays a key role in many aspects.
[0087] In terms of data feature mining, it can deeply analyze sequence data such as area number, concealment, and observability, and dig out the hidden patterns and laws. For example, it finds specific sequence patterns that frequently appear in areas with high concealment and low observability, as well as the correlation between different grid areas and changes in concealment and observability in pedestrian trajectories. In terms of understanding pedestrian behavior patterns, it comprehensively grasps the behavioral dynamics of pedestrians around target buildings by analyzing the mutual relationship between each sequence and the internal change law, and infers whether there are abnormalities in their behavioral intentions, such as whether they deliberately avoid areas with high observability. In terms of abnormal behavior recognition, after it fully learns various sequence patterns under normal circumstances, once it monitors a sequence combination that is significantly different from the normal pattern, it can detect it in time, providing a strong basis for identifying hidden social security behaviors, thereby helping to ensure social security.
[0088] Furthermore, the Transformer autoencoder includes the following construction steps:
[0089] Step 1: Determine input, output and model architecture
[0090] Clarify input and output: First, determine the input and output of the model according to the task to be processed. In the current situation, the input is the area number sequence, the concealment sequence, and the observability sequence, which record the information related to the activities of pedestrians around the target building from different angles. The output is to restore these input sequences as accurately as possible, so as to learn the patterns and rules in the sequence by comparing the differences between the input and output.
[0091] Determine the model architecture: Choose to build the autoencoder based on the Transformer architecture. The Transformer architecture is well-known for its powerful ability to process sequence data. Its core components include multi-head attention mechanism, feedforward neural network layer, etc. These components will be configured and built in detail in subsequent steps.
[0092] Step 2: Build the encoder
[0093] Input embedding layer: Create an input embedding layer to embed different types of input sequences (region number sequence, concealment sequence, observability sequence) and convert them into vector representations suitable for model processing. This step can use common embedding methods, such as linear embedding or embedding based on pre-trained models, to ensure that each element has a suitable representation in the vector space for subsequent calculations.
[0094] Multi-Head Attention Layer: Next, build a multi-head attention layer and set multiple attention heads (usually 2 to 8 or determined according to specific requirements). Each attention head independently calculates the correlation weights between each position in the input sequence and other positions. By calculating multiple attention heads in parallel, the model can capture the complex relationships between sequence elements from multiple perspectives. For example, for the sequence of area numbers, the degree of mutual association between different grid areas in the pedestrian trajectory can be analyzed; for the concealment sequence, the mutual influence of concealment changes at different positions can be understood, etc.
[0095] Feed-Forward Neural Network Layer: After the multi-head attention layer, add a feed-forward neural network layer. This layer consists of multiple fully connected neurons, which perform further non-linear transformations on the features processed by the multi-head attention layer to extract more abstract and higher-level feature representations. It can enhance the model's ability to capture the features of the input sequence and transform the sequence data into a form more suitable for subsequent processing and analysis.
[0096] Layer Normalization and Residual Connection: To make the model training more stable and efficient, perform layer normalization operations after the multi-head attention layer and the feed-forward neural network layer respectively, normalizing the input of each layer to reduce problems such as gradient vanishing or gradient explosion. At the same time, adopt the residual connection method, adding the input of each layer to the output after passing through this layer, ensuring that the gradient can be transmitted more directly between layers, which helps to alleviate the degradation problem that occurs during the training of deep neural networks.
[0097] Step 3: Build the decoder part
[0098] Input Embedding Layer: According to the specific design of the decoder, it may be necessary to create an input embedding layer to process the initial input of the decoder. If the input of the decoder is the feature vector output by the encoder, then this embedding layer may perform further embedding processing on it to make it more suitable for the internal calculations of the decoder.
[0099] Masked Multi-Head Attention Layer: Build a masked multi-head attention layer, similar to the multi-head attention layer in the encoder, but here the masking technology is adopted. The mask restricts the model to only see the sequence elements before the current position when calculating the attention weights, which can prevent the model from seeing future information in advance during the decoding process and ensure the sequentiality and rationality of the decoding process. In this way, the decoder can gradually reconstruct the original input sequence based on the existing information.
[0100] Interaction between the Multi-Head Attention Layer and the Encoder: Add a multi-head attention layer to interact the result output by the masked multi-head attention layer with the feature vector output by the encoder. By applying the multi-head attention mechanism again, the decoder can better utilize the features extracted by the encoder to further accurately reconstruct the original input sequence.
[0101] Feedforward neural network layer: A feedforward neural network layer is also set up to perform further nonlinear transformation on the previously processed features to complete the task of reconstructing the original input sequence.
[0102] Layer normalization and residual connection: Like the encoder, layer normalization and residual connection operations are also performed between the layers of the decoder to ensure stable and efficient training and alleviate the degradation problem.
[0103] Step 4: Set up the loss function and optimizer
[0104] Select the loss function: In order to measure the difference between the sequence restored by the decoder and the original input sequence, select an appropriate loss function. A common loss function is the mean square error (MSE) loss function, which quantifies the difference by calculating the average of the sum of squares of the differences between the corresponding elements of the restored sequence and the original sequence. The training goal of the model is to minimize the value of this loss function so that the decoder can reconstruct the original input sequence as accurately as possible.
[0105] Select optimizer: At the same time, select a suitable optimizer to adjust the parameters of the model so that the model can continue to improve in the direction of minimizing the loss function. Commonly used optimizers include stochastic gradient descent (SGD) and its variants (such as Adagrad, Adadelta, Adam, etc.). These optimizers update the parameters of the model according to different gradient update strategies to improve the training efficiency and convergence speed of the model.
[0106] Step 5: Model training and evaluation
[0107] Training data preparation: The collected data sets containing region number sequences, concealment sequences, and observability sequences are divided into training sets, validation sets, and test sets according to a certain ratio. Usually, the training set is used to train the model, the validation set is used to adjust the parameters and hyperparameters of the model during the training process, and the test set is used to finally evaluate the performance of the model.
[0108] Model training: Use the training set to train the Transformer autoencoder. During the training process, the training set data is sent to the encoder part of the model in batches. After the encoding and decoding process, the restored sequence is obtained. Then the difference between the restored sequence and the original sequence is calculated according to the set loss function, and the parameters of the model are updated according to this difference through the optimizer. This training process is repeated until the model reaches the predetermined training goal, such as the value of the loss function is reduced to an acceptable level or a certain number of training rounds have passed.
[0109] Model evaluation: After training is completed, use the test set to evaluate the model. Evaluate the performance of the model by calculating metrics such as the loss function value, accuracy, recall rate, etc. on the test set (determined according to specific task requirements). If the performance of the model does not meet expectations, it may be necessary to adjust the model architecture, parameters, or hyperparameters, and then retrain until satisfactory performance is obtained.
[0110] S9. Obtain the surveillance video stream to be detected for the target building, and use the sequence of area numbers, concealment degrees, and observability degrees of the pedestrian trajectories in the video to be detected as the input sequence and input it into the Transformer autoencoder to obtain the reconstructed sequence;
[0111] In this embodiment, in the process of deeply analyzing and monitoring the pedestrian activities and related environmental characteristics around the target building, we first need to obtain the surveillance video stream data to be detected for the target building. These surveillance video streams are like a dynamic information treasure trove, completely recording the personnel flow and scene changes around the target building within a specific time period.
[0112] Next, for the obtained surveillance video stream to be detected, we use a series of analysis methods adopted before to extract the key information of the pedestrian trajectories. Specifically, according to the order of the grid areas passed by the pedestrian trajectories, generate the corresponding sequence of area numbers, concealment degrees, and observability degrees respectively.
[0113] The sequence of area numbers can clearly show the order of the grid areas passed by pedestrians outside the target building, reflecting the distribution characteristics of the pedestrian activity paths in different areas; the concealment degree sequence quantifies the concealment degree of each area from the perspective of pedestrian activities based on the ratio of the number of pedestrians passing through each grid area to the total number of pedestrians in all grid areas; the observability degree sequence is obtained by calculating the ratio of the visible length of the target building to the total length for each grid area as an observation point, which reflects the change in the observability degree of observing the target building from different areas.
[0114] After generating the above three sequences, use them as the input sequence and input it into the well - constructed and fully trained Transformer autoencoder. The Transformer autoencoder, with its powerful ability to extract and learn features from sequence data, will deeply process these input sequences.
[0115] Through components such as the multi - head attention mechanism and the feed - forward neural network layer in the encoder part, perform feature encoding on the input sequence to capture the internal relationships and hidden patterns between sequence elements. Then, the decoder part attempts to restore a sequence as similar as possible to the original input based on the feature vectors output by the encoder, and this finally obtained sequence is the reconstructed sequence.
[0116] The obtained reconstructed sequence is of great significance for further analyzing pedestrian behavior patterns and identifying potential hidden social security behaviors. By comparing the differences between the input sequence and the reconstructed sequence, we can discover anomalies in pedestrian activities. For example, the concealment or observability in certain areas shows changes inconsistent with the normal pattern in the reconstructed sequence, or the sequence of area numbers of the pedestrian trajectory deviates abnormally, etc., thus providing strong data support and decision-making basis for relevant security monitoring and risk prevention work.
[0117] S10. Calculate the anomaly score based on the error between the reconstructed sequence and the input sequence. When the anomaly score exceeds a preset threshold, it is determined that there is a hidden social security behavior.
[0118] In this embodiment, the anomaly score is calculated based on the error between the reconstructed sequence and the input sequence. This error calculation process is an important means to measure the deviation degree of the model's reconstruction effect of the input sequence from the original input. Specifically, for each element in the input sequence (corresponding to the values of the area number, concealment, and observability at each time step or position), we calculate the difference between it and the corresponding element in the reconstructed sequence. Then, according to certain calculation rules, these differences are combined to form a value that can quantify the overall error degree, and this value is the anomaly score.
[0119] After calculating the anomaly score, we compare it with the preset threshold. This preset threshold is determined based on a large number of experiments, data analyses, and evaluations of possible normal and abnormal situations in combination with the actual application scenario. It represents a boundary for distinguishing normal behavior patterns and potential hidden social security behaviors. When the anomaly score exceeds the preset threshold, it is determined that there is a hidden social security behavior. This is because under normal circumstances, a well-trained Transformer autoencoder should be able to reconstruct the input sequence well, keeping the error between the reconstructed sequence and the input sequence within a relatively small range. Once the anomaly score exceeds the threshold, it indicates that the characteristics presented by the current input sequence related to the pedestrian trajectory deviate significantly from the normal pattern learned by the model. Such deviations are likely due to abnormal behaviors of pedestrians, such as deliberately choosing hidden areas to stay for a long time, avoiding areas with high observability for suspicious activities, etc. These behaviors are often associated with hidden social security behaviors. In this way, by calculating the anomaly score based on the error between the reconstructed sequence and the input sequence and making a determination according to the preset threshold, an effective quantitative analysis method is provided for timely discovering and identifying hidden social security behaviors, which helps to strengthen the security monitoring and risk prevention of the surrounding environment of the target building.
[0120] Furthermore, the formula for the anomaly score is:
[0121]
[0122] In the formula, AS represents the anomaly score; l i is the input concealment degree sequence, and l′ i is the reconstructed concealment degree sequence; u i is the input observability sequence, and u′ i is the reconstructed observability sequence; n is the number of elements in the sequence; p y-gram represents the probability that all consecutive y IDs in the reconstructed area number sequence are the same as those in the input area number sequence, and k is the number of consecutive IDs that are the same in the reconstructed and input area number sequences. It means that the larger the number y of consecutive identical IDs, the larger this value; ω l 、ω u and ω y are the weight coefficients of the concealment degree sequence, the observability sequence, and the area number sequence respectively.
[0123] The weight coefficients are determined by analyzing and statistically processing a large amount of historical data (including pedestrian-related sequence data with confirmed existence or non-existence of concealed social security behaviors) to determine the weights of each sequence.
[0124] Correlation analysis: Calculate the correlation coefficients between each sequence and the known occurrence or non-occurrence of concealed social security behaviors. The higher the correlation of a sequence, the higher the weight should be given in the judgment. For example, after analysis, it is found that the concealment degree sequence has the strongest correlation with the occurrence of concealed social security behaviors, followed by the observability sequence, and then the area number sequence. Then, the weight allocation can be adjusted accordingly. For example, the weight of the concealment degree sequence is set to 50%, the weight of the observability sequence is set to 30%, and the weight of the area number sequence is set to 20%.
[0125] In addition, some feature importance evaluation tools in machine learning algorithms can also be used, such as the feature importance index in the random forest algorithm. Taking the area number sequence, the concealment degree sequence, and the observability sequence as features and inputting them, through model training and analysis, the importance scores of each feature (i.e., each sequence) are obtained, and then the weights are allocated according to the score ratio.
[0126] Furthermore, in step S10, when it is determined that there is a concealed social security behavior, identify the pedestrians with abnormal behaviors in the input sequence.
[0127] In this embodiment, it is necessary to further analyze which pedestrians in the input sequence have abnormal behaviors. Only by identifying the specific pedestrians with abnormal behaviors can subsequent monitoring, investigation, or preventive measures be taken more pertinently. To identify the pedestrians with abnormal behaviors, a more detailed analysis of the input sequence is carried out. First, review the sequence of area numbers and check whether there are abnormal patterns in the order of grid areas passed by each pedestrian's trajectory. For example, whether some pedestrians frequently enter and exit grid areas with high concealment and low observability, or whether their trajectories show detours or long stays that do not conform to the normal pedestrian flow direction in specific areas. By analyzing the characteristics of these area number sequences, some pedestrians with suspicious behaviors can be initially screened out.
[0128] Secondly, combine the concealment sequence for further confirmation. Observe the concealment of the grid areas passed by the pedestrians initially screened out. If it is found that a certain pedestrian often passes through areas with extremely high concealment and stays in these areas for a long time, this further increases the possibility of their abnormal behavior. In addition, the observability sequence can also provide important clues. If a pedestrian's trajectory always appears in areas with extremely low observability, or the observability suddenly drops significantly when passing through certain areas and is accompanied by abnormal staying behaviors, this also implies that the pedestrian may have abnormal behaviors.
[0129] By comprehensively considering the characteristics of the area number sequence, concealment sequence, and observability sequence, pedestrians with abnormal behaviors in the input sequence can be more accurately identified, thus laying a solid foundation for subsequent further monitoring, investigation, and taking corresponding safety preventive measures for these specific pedestrians, and effectively ensuring the safety and order of the surrounding environment of the target building.
[0130] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed as above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to form equivalent embodiments with equivalent changes, but as long as they do not depart from the technical content of the present invention, any brief modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for detecting anomalies of concealed social security behaviors, characterized by: The following steps are involved: S1. Obtain external surveillance video stream data of the target building for a period of time as a training set, and use the trajectory monitoring algorithm to extract pedestrian trajectories in the surveillance video stream; S2, gridding each frame of the video stream image, assigning a separate ID to each grid area, and generating a sequence of ID numbers according to the order of the grid areas passed by the pedestrian trajectory, recorded as the area number sequence; S3, defining the concealment degree of each grid area according to the ratio of the number of pedestrians passing through each grid area to the total number of pedestrians passing through all grid areas; S4, generating a sequence of concealment degrees according to the order of the grid areas passed by the pedestrian trajectory, recorded as a concealment degree sequence; S5. The center of each grid area is taken as the observation point, and the sight line is extended from the observation point to the target building to obtain the visible length of the target building; S6. Calculate the ratio of the visible length and the total length of the target building to obtain the observability of the observation point for target monitoring; S7, generating a sequence of observable degrees according to the order of the grid areas passed by the pedestrian trajectory, recorded as the observable degree sequence; S8. Construct a Transformer autoencoder to learn the region number sequence, concealment sequence, and observability sequence; S9, obtain the surveillance video stream of the target building to be detected, and input the area number sequence, concealment degree sequence and observability degree sequence of the pedestrian trajectory in the video to be detected as the input sequence into the Transformer autoencoder to obtain the reconstructed sequence; S10. Calculate an anomaly score based on the error between the reconstructed sequence and the input sequence. When the anomaly score exceeds a preset threshold, it is determined that there is a behavior of concealing social security.
2. The method for detecting anomalies of concealed social security behaviors according to claim 1, characterized in that: The method of extracting pedestrian trajectories from a surveillance video stream using a trajectory monitoring algorithm comprises the following steps: S11, video stream data reading and preprocessing: read the external surveillance video stream data of the target building and load it into the processing system, then preprocess each frame of the image, improve the clarity through image enhancement and reduce noise interference through denoising; S12, pedestrian target detection: load the pre-trained pedestrian detection model, scan each frame of the video stream, identify the pedestrian area and output its location information and confidence score, and determine the pedestrian target accordingly; S13, pedestrian tracking initialization: according to the confidence score output by pedestrian detection, a threshold is set to screen valid pedestrian targets, and then a tracker is initialized for the valid pedestrian targets. The tracker is used to predict the current frame position based on the state of the pedestrian in the previous frame, and the subsequent tracking process is started; S14. Pedestrian tracking and trajectory update: In subsequent frames, the tracker continues to track the initialized pedestrian target, updates the tracking status based on the match between the actual position of the pedestrian and the predicted position, and records the position of the pedestrian in each frame to form and continuously improve the pedestrian trajectory information.
3. The method for detecting anomalies of concealed social security behaviors according to claim 1, characterized in that: The concealment degree is calculated as follows: In the formula, H i is the concealment degree of the i-th grid area; n i is the number of pedestrians passing through the i-th grid area; N is the total number of pedestrians in all grid areas.
4. The method for detecting anomalies of concealed social security behaviors according to claim 1, characterized in that: The observability is calculated as follows: Where V i is the observability of the target building; Visible length is the visible length of the target building; Totallength is the total length of the target building.
5. The method for detecting anomalies of concealed social security behaviors according to claim 1, characterized in that: The Transformer autoencoder includes the following construction steps: S81. Determine input, output and model architecture: Use the region number sequence, concealment sequence and observability sequence as model input, and the original input sequence as the expected output. Select the Transformer architecture to build an autoencoder. S82. Build the encoder: first create an input embedding layer to convert various input sequences into vector representations; build a multi-head attention layer, set multiple attention heads to capture the complex relationship between sequence elements from multiple angles; add a feedforward neural network layer to perform nonlinear transformation to extract high-level features; Finally, layer normalization and residual connection operations are performed after the multi-head attention layer and the feedforward neural network layer respectively; S83, build a decoder: set up an input embedding layer to process the initial input of the decoder; Build a masked multi-head attention layer and use masking technology to decode in sequence; Add a multi-head attention layer that interacts with the encoder to exploit encoder features; set up a feed-forward neural network layer to complete the sequence reconstruction task, and perform layer normalization and residual connections; S84, set the loss function and optimizer: select the loss function to measure the difference between the decoder restored sequence and the original input sequence, and minimize the difference as the training goal; select the optimizer to adjust the model parameters according to the gradient update strategy to improve the training efficiency and convergence speed; S85, model training and evaluation: Divide the data set containing related sequences into training set, validation set and test set, use the training set to train the built model, calculate the difference based on the loss function and update the parameters until the predetermined goal is achieved.
6. The method for detecting anomalies of concealed social security behaviors according to claim 5, characterized in that: In step S84, the loss function is configured as a mean square error loss function; and the optimizer is configured as stochastic gradient descent.
7. The method for detecting anomalies of concealed social security behaviors according to claim 1, characterized in that: In step S10, the anomaly score is calculated as follows: Where AS represents the anomaly score; l i is the input concealment sequence, is the reconstructed concealment sequence; u i is the observable sequence of the input, is the reconstructed observable sequence; n is the number of elements in the sequence; p y-gram represents the probability that all y consecutive IDs in the reconstructed area number sequence are the same as the input area number sequence, k is the number of consecutive IDs that are the same in the reconstructed and input area number sequences, It means that the more consecutive identical IDs y are, the larger the value will be; ω l ,ω u and ω y are the weight coefficients of concealment sequence, observability sequence and area number sequence respectively.
8. The method for detecting anomalies of concealed social security behaviors according to claim 7, characterized in that: The weight coefficient is determined by calculating the correlation coefficient between each sequence in the historical period and whether the known concealed social security behavior occurs or not.
9. The method for detecting anomalies of concealed social security behaviors according to claim 1, characterized in that: In step S10, when it is determined that there is a concealed social security behavior, pedestrians with abnormal behavior in the input sequence are identified.
Citation Information
Patent Citations
Personnel group behavior monitoring method based on video analysis
CN114863352A
Airport security check behavior pattern analysis and abnormal state detection method and system
CN114898257A
Anomalous pattern discovery
US20120237081A1
A system and method to determine anomalous behavior
WO2023150642A1
Object maskinformation for supplementalenhancement information message
WO2024213070A1