Airport crossing safety management method and system based on video capture
Through the airport crossing safety management method based on video capture, multiple targets are identified and tracked in real time, and their motion trajectories are generated. Through dynamic detention warning and abnormal behavior alarm signals, the problem of the inability to effectively identify multi-objective dynamic behavior and safety hazards in the existing technology is solved, and efficient and accurate safety management is achieved.
Patent Information
- Application Number
- CN202510144361.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The prior art cannot effectively identify multi-target dynamic behavior, trailing violations, and target retention behavior, resulting in safety hazards, monitoring blind spots and limited human resources.
The airport crossing safety management method based on video capture is adopted to obtain and preprocess video data in real time, identify and track multiple targets, generate their motion trajectories, and solve the problems of monitoring blind spots and limited personnel through dynamic detention warnings and abnormal behavioral alarm signals.
Real-time monitoring and behavioral trajectory analysis of multiple targets are achieved, abnormal static and trailing violations are effectively identified, and real-time, accuracy and response efficiency of road crossing safety management are improved.
Smart Images

Figure CN120071218A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of crossing safety management, and particularly to an airport crossing safety management method and system based on video capture. Background Art
[0002] In modern airport operations, crossing safety management is a key link to ensure smooth airport operations and passenger safety. As the main passage for vehicles and personnel to enter and exit, once a safety accident occurs at the airport crossing, it will have a serious impact on the normal operation of the airport and even threaten the lives of passengers and staff. In order to prevent illegal elements from tailing vehicles into the area, unauthorized personnel from breaking into restricted areas, and other safety hazards, it is necessary to strengthen the monitoring and management of the crossing area. Therefore, establishing an effective airport crossing safety management system that can monitor, identify, and warn of potential safety threats in real time is crucial for improving the overall safety protection ability of the airport.
[0003] In the prior art, the safety management of airport crossings mainly relies on a combination of video surveillance systems and manual inspections, with manual inspections being the main method. By setting up surveillance cameras, staff can monitor the crossing area to promptly detect abnormal behaviors or potential safety hazards. However, the surveillance footage usually needs to be viewed manually in real time. Given the large monitoring area of the crossing, frequent vehicle and personnel flow, and limited human resources, it is difficult to achieve full coverage and real-time response. In addition, due to the limited number of staff and the gradually aging age structure of the inspection personnel, the energy and reaction speed of manual inspections are affected. Moreover, there are line-of-sight blind spots and monitoring dead zones in some areas of the airport crossing, making safety inspections vulnerable to human factors and difficult to achieve real-time analysis and accurate judgment of multi-target behaviors. This management mode mainly relying on human resources is prone to lagging warnings or missed reports when facing complex interaction behaviors, affecting the overall safety management level of the crossing.
[0004] In summary, there are still many problems in the airport crossing safety management in the prior art, especially in the multi-target dynamic monitoring of the crossing area and the identification of abnormal behaviors. Although the existing monitoring systems can cover the crossing to a certain extent, due to the lack of accurate identification and interaction analysis of multiple targets, the monitoring systems cannot handle the complex behavior patterns between multiple targets simultaneously. Especially when there are behaviors such as multi-target interaction and target retention, it is often difficult to identify and issue warning signals in a timely manner. In addition, during manual inspections, due to limited monitoring personnel and the inability to have someone constantly monitor the surveillance footage 24 / 7, it is impossible to promptly detect tailing and illegal behaviors in dead zones of large vehicles or objects (such as long lanes), increasing safety hazards. Therefore, the prior art cannot provide comprehensive, accurate, and efficient safety management for the airport crossing area, and there is an urgent need for a solution to solve the problems of multi-target dynamic monitoring and abnormal behavior identification. Summary of the Invention
[0005] In view of the above actual situation, the present application proposes an airport crossing safety management method and system based on video capture to solve the problems existing in the prior art, such as the inability to simultaneously track multiple target interactions, complex behavior pattern analysis, and monitoring dead angles, especially the potential safety hazards of being unable to effectively identify multi-target dynamic behaviors, tailgating violations, and target retention behaviors.
[0006] An airport crossing safety management method based on video capture, the method comprising the following steps:
[0007] S1, obtaining real-time data to be processed, where the real-time data to be processed is video crossing picture data captured in real time by a video monitoring device;
[0008] S2, preprocessing the real-time data to be processed to obtain real-time preprocessed data;
[0009] S3, performing identification processing on a first object and a second object for the real-time preprocessed data, generating dynamic data of the first object and the second object, and generating respective motion trajectories based on the dynamic data, and assigning an independent ID to each target;
[0010] S4, performing trajectory filtering processing on the moving trajectory of the first object under the unique ID, identifying an abnormal stationary target and outputting a dynamic retention warning signal;
[0011] S5, when outputting the dynamic retention warning signal, performing screening and matching analysis on the moving trajectory of the second object, identifying abnormal trajectories and triggering an abnormal behavior alarm signal and an associated response action.
[0012] Further, the S2 step includes the following sub-steps:
[0013] S201, performing a frame extraction operation on the real-time data to be processed, decomposing continuous video pictures into a discrete frame sequence and performing normalization processing to obtain discrete normalized data, and performing Gaussian filtering on the discrete normalized data for denoising processing to generate denoised frame data;
[0014] S202, performing background modeling and foreground segmentation processing on the denoised frame data, extracting a moving foreground region by using a segmentation algorithm of a convolutional neural network to generate foreground segmentation data, and adding a timestamp mark to each frame of the foreground segmentation data to generate real-time preprocessed data.
[0015] Further, the identification processing of the first object in the S3 step includes the following sub-steps:
[0016] S301. Separate the foreground segmentation data in the real-time preprocessed data in the spatio-temporal domain, filter the motion features of the foreground region through an adaptive spatio-temporal filtering algorithm to generate motion pattern data, and the spatio-temporal domain separation is performed by combining a convolutional neural network and LSTM;
[0017] S302. Apply a partition processing algorithm based on tensor decomposition to the motion pattern data to dynamically construct a target motion attribute model within the region, thereby generating first object dynamic data, and the partition processing algorithm is implemented through local maximum clustering and an adaptive threshold algorithm;
[0018] S303. Use an optimized Kalman filtering algorithm based on joint time-frequency analysis to model the motion trajectory of the first object dynamic data, generate the moving trajectory of the first object, and assign an independent ID to it. The trajectory modeling uses the fusion of weighted mean filtering and a multi-dimensional Gaussian model.
[0019] Furthermore, the processing of the second object recognition in step S3 includes the following sub-steps:
[0020] S311. Generate a local saliency map for the foreground segmentation data in the real-time preprocessed data, and expand and contract the object region through a boundary tracking method based on a spatio-temporal composite optimization algorithm to extract second object candidate region data;
[0021] S312. For the second object candidate region data, perform hierarchical feature extraction in combination with a deep autoencoder neural network, and generate second object dynamic data through a dynamic flow field model and a motion clustering algorithm. The feature extraction is jointly processed by a graph convolutional network and a motion prediction algorithm;
[0022] S313. Apply a multi-scale convolutional neural network based on spatio-temporal correlation analysis to the second object dynamic data for temporal sequence modeling, generate the moving trajectory of the second object, and assign an independent ID to it.
[0023] Furthermore, step S4 includes the following sub-steps:
[0024] S401. Perform time series analysis on the moving trajectory of the first object under the unique ID, calculate the stationary duration within each time period, and compare it with a preset benchmark motion model to obtain stationary behavior deviation data;
[0025] S402. Cumulatively process the stationary behavior deviation data under the unique ID, calculate the cumulative deviation value for multiple time periods, and generate a deviation change rate according to the cumulative deviation change situation of adjacent time periods. The deviation change rate is generated through differential analysis and a weighted moving average algorithm;
[0026] S403. Construct a spatio-temporal change trend model of the target stationary behavior based on the cumulative deviation value and deviation change rate under the unique ID. The model is modeled through a spatio-temporal convolutional network to obtain the change trend of the target stationary behavior.
[0027] S404. Conduct causal inference analysis on the spatio-temporal change trend model under the unique ID to generate target abnormal behavior determination data, and perform comprehensive analysis in combination with a multi-objective decision-making network to output a dynamic stay warning signal.
[0028] Further, the S5 step includes the following sub-steps:
[0029] S501. When a dynamic stay warning signal is output, screen the movement trajectory of the second object based on the independent ID of the second object and the spatio-temporal constraint model. The spatio-temporal constraint model generates a difference data set between the trajectory abnormal points and the normal behavior trajectory through dynamic spatio-temporal window and spatial neighborhood clustering analysis, so as to generate an effective target data set.
[0030] S502. Use an adaptive trajectory similarity matching algorithm to perform similarity comparison on the effective target data set, compare the effective target data set with a preset normal action trajectory template, and generate a matching analysis result.
[0031] S503. Process the matching result based on the spatio-temporal deviation fusion analysis model, identify the deviation from the preset trajectory exceeding the threshold, output an abnormal behavior alarm signal, and trigger an associated response action.
[0032] Further, the timing modeling process in the S313 includes cross-frame motion feature extraction, spatio-temporal correlation modeling, and trajectory generation and ID assignment.
[0033] Further, the dynamic time warping algorithm is used for the time series analysis in the S401 step.
[0034] Further, the similarity comparison in the S502 step includes feature extraction processing and similarity calculation processing. The similarity calculation processing is performed using an adaptive trajectory similarity matching algorithm. The adaptive trajectory similarity matching is calculated through the dynamic time warping algorithm. The spatio-temporal deviation fusion analysis model in the S503 step includes deviation value extraction processing and spatio-temporal fusion calculation processing.
[0035] In addition, the present application also proposes an airport crossing safety management system based on video capture, which is characterized in that the system includes:
[0036] An acquisition unit for acquiring real-time data to be processed, where the real-time data to be processed is video crossing picture data captured by a video monitoring device in real time.
[0037] A preprocessing unit for preprocessing real-time data to be processed to obtain real-time preprocessed data;
[0038] An identification unit for identifying a first object and a second object from the real-time preprocessed data, generating dynamic data of the first object and the second object, modeling their respective movement trajectories based on the dynamic data, and assigning independent IDs to each target;
[0039] A retention detection unit for filtering the movement trajectory of the first object under the unique ID, identifying abnormal stationary targets, and outputting a dynamic retention warning signal;
[0040] An anomaly detection unit for screening and matching analysis of the movement trajectory of the second object when the dynamic retention warning signal is output, identifying trajectory anomalies, and triggering an abnormal behavior alarm signal and associated response actions.
[0041] An airport crossing safety management method and system based on video capture proposed by this application realizes real-time monitoring and behavior trajectory analysis of multiple targets, can effectively identify abnormal stationary and trailing violation behaviors, and solves the problems of monitoring blind spots and limited personnel through the triggering of dynamic retention warnings and abnormal behavior alarm signals, improving the real-time performance, accuracy, and response efficiency of crossing safety management. Description of the Drawings
[0042] Figure 1 It is a schematic flowchart of a method for airport crossing safety management based on video capture proposed by this application;
[0043] Figure 2 It is a schematic flowchart of the preprocessing data in a method for airport crossing safety management based on video capture proposed by this application;
[0044] Figure 3 It is a schematic flowchart of the first object identification process in a method for airport crossing safety management based on video capture proposed by this application;
[0045] Figure 4 It is a schematic flowchart of the second object identification process in a method for airport crossing safety management based on video capture proposed by this application;
[0046] Figure 5 It is a schematic flowchart of outputting a dynamic retention warning signal in a method for airport crossing safety management based on video capture proposed by this application;
[0047] Figure 6 It is a schematic flowchart of triggering an abnormal behavior alarm signal and associated response actions in a method for airport crossing safety management based on video capture proposed by this application;
[0048] Figure 7 Schematic structural diagram of a production parameter reverse prediction system based on tar content provided by an embodiment of the present application; Specific embodiments
[0049] The following will clearly and completely describe the simulation technical route in the embodiments of the present invention in conjunction with the drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0050] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments.
[0051] The following further describes the features and performance of the present invention in conjunction with embodiments. Please refer to the attached Figure 1 As shown, a road crossing safety management method based on video capture, the method includes the following steps:
[0052] S1. Obtain real-time data to be processed, where the real-time data to be processed is video crossing picture data captured by a video monitoring device in real time;
[0053] In this embodiment, this step involves obtaining real-time data to be processed, specifically by a video monitoring device capturing crossing picture data in real time. The data to be processed comes from a camera or video monitoring device installed near the crossing. The video monitoring device can continuously collect video images in the crossing area and convert the collected image data into electronic signals, and transmit them to the processing unit in real time. The video monitoring device forms a coverage area in the crossing area, which can comprehensively capture the dynamic situation in the crossing to ensure that scene data is obtained without dead angles.
[0054] In some embodiments, the video monitoring device is a high-definition digital video acquisition device, and the device has high resolution and real-time processing capabilities. The video device can capture the image information of the crossing in real time at a certain frame rate, and the frame rate can be adjusted according to requirements, generally reaching 30 frames per second or more. At this time, the video monitoring device processes the continuously collected image signals according to a specific video coding standard, such as the H.264 or H.265 standard, and then compresses and transmits them to the data processing platform.
[0055] The video crossing scene data not only includes the static background information within the crossing area, but also includes various dynamically changing information. This information is transmitted to the data stream to be processed through video monitoring devices, constituting the core data source of this step. The data to be processed includes each frame of video image and the timestamp information of each frame. The timestamp identifies the acquisition moment of each frame of data, ensuring the continuity and accuracy of the data in the time series.
[0056] In this embodiment, in order to ensure the accuracy and real-time performance of the data, the image acquisition principle adopted by the video monitoring device is based on optoelectronic sensing technology. The sensor senses the optical signal in the crossing area and converts it into an electrical signal. This electrical signal undergoes decoding and formatting processing by the image processing system and finally generates digital image data. In some implementation manners, the sensor of the video monitoring device adopts a CMOS or CCD image sensor. These sensors have high photosensitivity and can still maintain high image quality in low-light environments, ensuring clear images can be obtained under different lighting conditions.
[0057] The data to be processed is not limited to static image data, but also includes the motion information extracted from each frame of image. During the video acquisition process, the device can extract the moving target information within the crossing area through real-time difference analysis or motion detection algorithms, thereby providing more detailed dynamic data. This data provides a basis for subsequent target recognition, behavior analysis, and other processing.
[0058] In some implementation manners, the real-time data to be processed will undergo preliminary processing by the data buffer and storage unit to ensure that the data will not be lost during the processing. The video monitoring device has a certain memory capacity for storing the acquired video frame data and caches it according to the set duration. This process helps to ensure the continuity of the data stream and avoid data omission or loss during high-speed acquisition.
[0059] Through the real-time captured video crossing scene data, the data to be processed not only has real-time performance but also has complete spatio-temporal information, which can provide a reliable basis for subsequent image analysis, target recognition, and behavior judgment. The data processing flow is based on the real-time capture, transmission, and processing of video signals, ensuring the system's real-time monitoring and dynamic response capabilities for the crossing situation.
[0060] S2. Preprocess the real-time data to be processed to obtain real-time preprocessed data;
[0061] Specifically, please refer to the appendix Figure 2 As shown, this step includes the following sub-steps:
[0062] S201, perform a frame extraction operation on the real-time data to be processed, decompose the continuous video frames into a discrete frame sequence and perform normalization processing to obtain discrete normalized data, and perform Gaussian filtering on the discrete normalized data for denoising processing to generate denoised frame data;
[0063] In some embodiments, the frame extraction operation is implemented by a sampling algorithm, which can extract each frame of image from the continuous video stream according to the set frame rate. The frame rate is set according to the actual application scenario and processing requirements, and in this embodiment, it is 30 frames per second or higher. Each frame of image data after extraction is stored in the form of a two-dimensional matrix, where each matrix element corresponds to the pixel value of a certain point in the image.
[0064] Perform normalization processing on the extracted discrete frame data. The purpose is to convert the pixel values of each frame of image into a unified numerical range, thereby improving the stability and accuracy of subsequent processing. In the embodiment, through normalization, the pixel values of the image will be adjusted to the range of [0, 1], ensuring that the image avoids deviations caused by differences in pixel value ranges during the processing.
[0065] In this embodiment, further perform Gaussian filtering on the normalized discrete data to remove the noise in the image. Noise often originates from environmental interference or sensor limitations and will affect the accuracy of target detection and image analysis. Specifically, the following two-dimensional Gaussian kernel function is used to perform convolution processing on the image: where (x, y) is the position in the filter kernel, and σ is the standard deviation of the Gaussian kernel, which controls the smoothness of the filter. The Gaussian filter reduces the interference of noise on image details by performing weighted averaging on each pixel point and its neighborhood. During the implementation process, the filter size and the standard deviation value σ will be adjusted according to the type and intensity of the noise in the image.
[0066] In this embodiment, the Gaussian filtering processing is implemented through the following convolution formula: where I x,y is the pixel value at the position (x, y) in the original image, G(i, j) is the filtering weight at the position (i, j) in the Gaussian kernel, and k is the radius of the filter. Through the above convolution operation, the noise in the image is effectively suppressed, and the generated denoised frame data has a good smoothing effect, which helps to improve the accuracy of subsequent image analysis tasks.
[0067] Through the steps of frame extraction, normalization processing, and Gaussian filtering, the obtained denoised frame data is clearer and has less noise compared to the original image. The denoised image data provides high-quality input data for subsequent target recognition, behavior analysis, etc. In some embodiments, the denoised data can be further subjected to other image enhancement processing, such as edge detection or brightness adjustment, to further improve the image quality.
[0068] S202. Perform background modeling and foreground segmentation on the denoised frame data. Use the segmentation algorithm of the convolutional neural network to extract the moving foreground region to generate foreground segmentation data, and add a timestamp mark to each frame of the foreground segmentation data to generate real-time preprocessed data.
[0069] In this embodiment, this step performs background modeling and foreground segmentation on the denoised frame data. The purpose of background modeling is to distinguish the static background and dynamic foreground in the image, and then extract the moving objects or changing regions, which represent the moving targets existing in the scene. In some embodiments, background modeling is achieved by statistical modeling methods, comparing the value of each pixel in the image with the background model to determine whether it belongs to the background.
[0070] The background modeling uses a segmentation algorithm based on the convolutional neural network (CNN). In this embodiment, the convolutional neural network automatically learns features from the denoised frame data through deep learning and performs pixel-level classification to distinguish the foreground and background. The core structure of the convolutional neural network includes multiple convolutional layers, activation function layers, pooling layers, etc. Each layer learns different spatial hierarchical features in the image. Through multiple convolutional operations, the convolutional neural network can effectively extract the local features in the image and separate the foreground and background by mapping the features to the output space through the fully connected layer.
[0071] In some embodiments, the training process of the convolutional neural network is based on a labeled dataset and is optimized through the backpropagation algorithm. The training objective is to minimize the difference between the predicted label and the true label, and usually the cross-entropy loss function is used to measure the accuracy of the output. The specific form of the loss function is as follows: where N is the total number of pixels in the image, y i is the true label, indicating whether the pixel is the foreground (1 for foreground, 0 for background), is the foreground probability predicted by the network.
[0072] After the network training is completed, apply the obtained convolutional neural network to each frame of the denoised image to extract the moving foreground region. The convolutional neural network distinguishes the dynamic part (i.e., the foreground region) and the static part (i.e., the background region) in each frame of the image through pixel-by-pixel classification to obtain the foreground segmentation data. The foreground region usually represents the target objects in the scene, and the motion information of these target objects will be tracked and analyzed in subsequent processing steps.
[0073] In some embodiments, to further enhance the accuracy of foreground segmentation, spatio-temporal constraints are used to optimize the foreground segmentation results. By introducing temporal continuity constraints, the foreground segmentation algorithm can identify the consistency of moving objects in multiple frames of images, reducing missegmentation caused by light changes or temporary interferences. In some embodiments, the foreground segmentation data is processed through temporal smoothing to further remove missegmented regions caused by motion blur or occasional noise.
[0074] In addition, a timestamp is added to the foreground segmentation data of each frame to record the acquisition time of the image. The timestamp provides temporal information for each frame of image data, which is helpful for subsequent moving object tracking and behavior analysis. The timestamp is usually automatically generated by the video surveillance device and stored together with each frame of image data. In some embodiments, the accuracy of the timestamp matches the video frame rate to ensure precise temporal alignment of the data.
[0075] The generated real-time preprocessed data includes denoising, segmentation, and timestamp information, including the foreground regions of the image and the temporal characteristics of these regions. This data provides accurate and structured input data for subsequent object recognition, object tracking, and behavior analysis.
[0076] S3 performs recognition processing on the first object and the second object for the real-time preprocessed data, generates dynamic data of the first object and the second object, models their respective motion trajectories based on the dynamic data, and assigns an independent ID to each target;
[0077] Specifically, the data to be processed is the filtered dynamic part, that is, the regions with motion in the image. However, the motion patterns within these dynamic regions are diverse. To better identify and track the different characteristics of these objects in the complex airport intersection safety management environment, the present application designs the first object and the second object. In this embodiment, the first object represents a relatively regular and predictable motion pattern, such as ground transportation vehicles, specifically including but not limited to passenger service vehicles (shuttle buses, boarding vehicles), cargo and baggage handling vehicles (baggage trailers, cargo lift platform vehicles), aircraft service vehicles (food trucks, fuel trucks, water trucks, power supply vehicles), ground support vehicles (tugboats, guiding vehicles), and other special vehicles (airport bird repellent vehicles, tractors); the second object represents dynamic and irregular targets, such as people and animals. Therefore, differentiating the dynamic foreground part to ensure accurate identification and efficient tracking of different types of targets helps improve the safety and response efficiency of the system.
[0078] Specifically, please refer to the appendix Figure 3 As shown, the recognition processing of the first object includes the following sub-steps:
[0079] S301. Separate the foreground segmentation data in the real-time preprocessed data in the spatio-temporal domain, filter the motion features of the foreground region through an adaptive spatio-temporal filtering algorithm, and generate motion pattern data. The spatio-temporal domain separation is performed by combining a convolutional neural network and an LSTM.
[0080] In some embodiments, in this step, the foreground segmentation data in the real-time preprocessed data is separated in the spatio-temporal domain, and then the dynamic data that meets the target requirements is screened and extracted to generate motion pattern data. In this embodiment, the spatio-temporal domain separation is performed by combining a convolutional neural network and a long short-term memory network, aiming to effectively distinguish different motion patterns and filter out irrelevant dynamic features, thereby optimizing the effects of subsequent target recognition and trajectory modeling.
[0081] In some embodiments, the foreground segmentation data is pixel-by-pixel classification implemented by a convolutional neural network, and the dynamic part has been separated from the static background. At this time, the foreground data still contains various different dynamic patterns. Therefore, the spatio-temporal separation algorithm analyzes the spatio-temporal features of the foreground region to extract the motion pattern data. The basic idea of the spatio-temporal domain separation is to combine the spatial features extracted by the CNN and the time series features processed by the LSTM to precisely filter the motion pattern data from two dimensions of time and space.
[0082] In this embodiment, the convolutional neural network is responsible for extracting spatial features from the image, and these features represent the local information within the foreground region. The CNN decomposes the input foreground image into multiple feature maps through a series of convolutional operations, and each feature map represents the feature expression of the image in a certain dimension. Assuming the input image is I(x, y, t), where x and y are spatial coordinates and t is the time dimension. The convolutional neural network obtains the feature map F(x, y, t) through convolutional operations, and this feature map contains the motion pattern features in space: where K(m, n) is the convolutional kernel, and m, n are the spatial positions of the convolutional kernel.
[0083] The extracted spatial features will be used as the input of the LSTM for modeling time series information. The role of the LSTM network is to capture the law of change of the motion pattern within the foreground region over time and identify the motion behavior of the target in different time frames. The core principle of the LSTM is to control the information flow through a gating mechanism (input gate, forget gate, and output gate) to model the time dependence within the foreground region. Assuming the input of the LSTM is x t , and its output is h t , the state transition equation of the LSTM is as follows: f t = σ(W f · [h t-1 , x t + b f ), i t= σ(W i · [h t-1 , x t + b i ), o t = σ(W o · [h t-1 , x t + b o ), h t = o t · tanh(C t ), where f t is the forget gate, i t is the input gate, is the candidate memory cell, C t is the cell state, o t is the output gate, h t is the output of the LSTM. After being processed by the LSTM network, the motion pattern data at each time point can be obtained.
[0084] In some embodiments, the motion pattern data obtained from the LSTM is further refined by an adaptive spatio-temporal filtering algorithm. The role of this filtering algorithm is to denoise the motion pattern, remove the irrelevant features therein, and retain the dynamic features related to the target recognition task. The adaptive spatio-temporal filtering algorithm is based on the motion features of the foreground region and their temporal variation rules, and by dynamically adjusting the filter parameters, it can update the filtering effect in real time. Assuming the motion pattern data is M(x, y, t), the filtering algorithm uses an adaptive filter to smooth the motion pattern data, thereby improving the accuracy of target recognition. Its mathematical model is: M filtered (x, y, t) = α · M(x, y, t) + (1 - α) · M filtered (x, y, t - 1), where α is the adaptive weight parameter, representing the weighted ratio of the motion pattern data of the current frame to the previous frame. Through frame-by-frame weighted averaging, the filter can dynamically adjust its response to spatio-temporal changes, thereby extracting the motion pattern data that meets the requirements of target recognition.
[0085] After being processed by the above spatio-temporal domain separation and adaptive spatio-temporal filtering algorithm, the dynamic features of the foreground region are effectively extracted and optimized, generating the motion pattern data. This data reflects the behavioral characteristics of the moving objects in the foreground region and contains the motion states of the target at different time periods. At this time, the motion pattern data not only removes the noise but also better highlights the motion rules of the target, providing accurate input data for subsequent target recognition and trajectory modeling.
[0086] S302. Apply the partition processing algorithm based on tensor decomposition to the motion pattern data to dynamically construct a target motion attribute model within the region, thereby generating first object dynamic data. The partition processing algorithm is implemented through local maximum clustering and an adaptive threshold algorithm.
[0087] In some embodiments, this step is based on the motion pattern data. By applying the partition processing algorithm based on tensor decomposition, a target motion attribute model within the region is dynamically constructed, thereby generating first object dynamic data. That is, by analyzing and clustering the local features of the motion pattern data, the moving targets within the foreground region are identified and their dynamic behavior features are extracted. Through the partition processing algorithm, complex motion pattern data can be transformed into more structured information for subsequent target recognition and trajectory modeling.
[0088] In this embodiment, the tensor decomposition method is used to efficiently process multi-dimensional data structures, thereby providing an effective means for modeling the target motion characteristics. Set the motion pattern data as M(x, y, t), which represents the dynamic characteristics at position (x, y) and time t. Since the motion pattern data exhibits the characteristics of spatio-temporal interweaving, using tensors to represent this data can effectively capture the correlations in the spatial and temporal dimensions. The tensor is represented as: where M(x, y, t) is the motion pattern data represented by the tensor, a ijk is the decomposition coefficient of the tensor, u i (x), v j (y), w k (t) represent the basis vectors at spatial positions x, y, and time t respectively. By decomposing this tensor, the feature matrices in each dimension are obtained, thereby revealing the correlations between different dimensions.
[0089] In some embodiments, the feature matrices obtained after tensor decomposition need to be further processed by partitioning. To accurately extract the target dynamic characteristics, local maximum clustering and an adaptive threshold algorithm are used to cluster the spatial region, thereby identifying the moving targets within the region. In this embodiment, the local maximum clustering algorithm is based on the local extreme value characteristics of the motion pattern data, automatically selects the significant change regions of the motion pattern, and further optimizes the clustering accuracy of the target by setting an adaptive threshold.
[0090] The mathematical model of local maximum clustering finds the maximum change region of the motion pattern by searching for the local extreme points of the motion pattern data within a spatio-temporally continuous region. Set the local region of the motion pattern data as R(x, y, t), and its local maximum point P(x 0 , y 0 , t 0 ) satisfies the following conditions: M(x 0 , y 0 , t0 ) > M(x, y, t) where N(x 0 , y 0 , t 0 ) represents the local neighborhood. This condition indicates that the point P(x 0 , y 0 , t 0 ) is a local maximum point, that is, it has the maximum value within its neighborhood.
[0091] Through the local maximum clustering algorithm, it is possible to extract the target region with strong motion characteristics from the spatio-temporal region and generate candidate target region data. Then the adaptive threshold algorithm is used to filter out the regions that do not meet the target recognition requirements. The adaptive threshold algorithm is dynamically adjusted based on the motion intensity of the target and the background noise, so as to accurately identify the target region. Set the dynamic threshold as T(x, y, t), and this threshold is adaptively adjusted according to the local characteristics of the motion pattern data: T(x, y, t) = α·max(M(x, y, t)) + (1 - α)·min(M(x, y, t)), where α is the adjustment coefficient, controlling the weights of the maximum and minimum values, and dynamically adjusting the threshold to ensure that the moving target can be effectively identified under different background conditions.
[0092] Through the above-mentioned partitioning process and clustering algorithm, it is possible to extract the dynamic behavior characteristics of the target from the motion pattern data and generate the first object dynamic data. At this time, the result of the partitioning process is a motion attribute model of the target within a spatio-temporal region, and this model reflects the motion law of the target in space and time. The mathematical representation of this model is: where D 1 (x, y, t) represents the dynamic data of the first object, b ijk is the new tensor decomposition coefficient obtained through local maximum clustering and the adaptive threshold algorithm, u i ′(x), v′ j (y), w′ k (t) are the processed spatial and temporal feature matrices.
[0093] S303. For the first object dynamic data, use the optimized Kalman filtering algorithm based on time-frequency joint analysis to model the motion trajectory, generate the motion trajectory of the first object, and assign a unique ID to it. The trajectory modeling uses the fusion of weighted mean filtering and multi-dimensional Gaussian model.
[0094] In this embodiment, the Kalman filtering algorithm is optimized into a time-frequency joint analysis method to more accurately track the changes when processing the motion data of dynamic targets. Set the dynamic data of the first object as D1(x, y, t), which represents the dynamic changes of the target in the spatio-temporal dimension. The goal of motion trajectory modeling is to smooth the true motion trajectory of the target by filtering the dynamic data.
[0095] The Kalman filter consists of two main parts: state update and covariance prediction. In some embodiments, the state update is optimized through the joint analysis of time and frequency information. Set the state vector as x(t), which contains the position information and velocity information of the target at time t: where x(t) and y(t) are the position coordinates of the target, and are the velocity components of the target. The state update equation of the Kalman filter is: x(t) = A·x(t - 1)+B·u(t), where A is the state transition matrix, representing the evolution law of the target state, B is the control matrix, and u(t) is the control input. In this embodiment, the control input u(t) contains frequency domain information to capture the changes in the target motion frequency.
[0096] In some embodiments, to enhance the stability and accuracy of the Kalman filtering algorithm, weighted mean filtering and a multi-dimensional Gaussian model are fused. The weighted mean filtering method smooths the motion trajectory of the target according to the relative position and velocity information of the target in the spatio-temporal domain, reducing the error caused by noise. Set the weighting factor of the weighted mean filtering as w i (t), and its calculation formula is: where x i (t) is the target state vector at the i-th moment, and w i (t) is the weight coefficient related to the motion characteristics at the current moment t. This coefficient is obtained through spatio-temporal distance measurement to ensure that the filter can adapt to the motion state of the target.
[0097] The multi-dimensional Gaussian model is used to describe the probability distribution of the target motion. Assume that the motion trajectory of the target follows a Gaussian distribution, and the distribution of the target state vector x(t) is expressed as: N(x(t); μ(t), Σ(t)), where μ(t) is the expected state vector of the target at time t, and Σ(t) is the covariance matrix of the target state. The Gaussian model can model the target trajectory, thereby helping the filtering algorithm better predict the future motion state of the target.
[0098] Based on the above-optimized Kalman filter algorithm, the obtained target state vector x(t) represents the motion trajectory of the target at different time points. Through the filtered trajectory information, the motion trajectory of the target can be modeled, and a unique ID is assigned to each target. The ID of each target is assigned based on its unique features on the motion trajectory to ensure accurate differentiation of different targets in subsequent trajectory analysis.
[0099] In this embodiment, the trajectory modeling performs temporal modeling through the filtered target state vector x(t) to construct the continuous motion trajectory of the target. The ID of the target is assigned based on the uniqueness of its motion trajectory. By comparing the similarity of different target trajectories, a matching algorithm based on trajectory features is used to assign a unique identifier to each target. Specifically, assume that the trajectory of target Ti is x i (t), and the trajectory of target Tj is x j (t). If the trajectory similarity of the two targets meets a certain threshold, they are considered to belong to the same target; otherwise, different IDs are assigned.
[0100] Please continue to refer to the appendix Figure 4 As shown, the recognition process for the second object includes the following sub-steps:
[0101] S311. Generate a local saliency map for the foreground segmentation data in the real-time preprocessed data, and expand and contract the object region through a boundary tracking method based on a spatio-temporal composite optimization algorithm to extract the second object candidate region data. The boundary tracking is realized by the joint optimization of image pyramid and wavelet transform;
[0102] In some embodiments, after the foreground segmentation data is processed by a convolutional neural network, the dynamic region has been distinguished from the static background. To more accurately extract the second object from these dynamic regions, a local saliency map is generated for the foreground segmentation data. The generation of the saliency map is based on the combination of image gradient calculation in the spatial domain and temporal information. Specifically, the saliency map is constructed by the following formula: where S(x, y, t) represents the saliency value at the time and spatial position (x, y, t), I(x, y, t) is the pixel intensity of the image at the position (x, y) and time t, i, j, k are image translation and time delay parameters, and n is the size of the sliding window. By generating the saliency map, the parts with significant changes in the foreground region are extracted, providing a basis for subsequent object tracking.
[0103] In this embodiment, the expansion and contraction of the object region are further performed through a boundary tracking method based on a spatio-temporal composite optimization algorithm, by means of the spatio-temporal domain expansion and contraction strategy during the boundary tracking process. Specifically, the boundary tracking adopts a joint optimization strategy of image pyramid and wavelet transform. The image pyramid effectively reduces the resolution of the image and achieves a rough localization of the region by constructing a multi-scale image hierarchy. The wavelet transform extracts the detailed information of the image through multi-scale analysis, thereby determining the boundary of the target region.
[0104] In some embodiments, the image pyramid is generated through the following steps: The image is divided into multiple scale layers, and the resolution of each layer is 1 / 2 of the previous layer. That is: where I p (x, y) represents the pixel value of the p-th layer pyramid, G p is the Gaussian filter kernel, and I p-1 (i, j) is the image of the previous layer pyramid. The target is detected and located at different scales through the layer-by-layer processing of the image pyramid.
[0105] The wavelet transform is used to extract the details of the boundary region. The image is decomposed into different frequency sub-bands through the wavelet transform, and the changes in the image edges are analyzed. For the wavelet transform at each scale, its transform is defined as: where W(x, y) is the wavelet transform result of the image at the position (x, y), and h(m, n) is the wavelet filter kernel. The subtle changes in the boundary are captured by performing wavelet transform on the boundary of each layer, thereby tracking the boundary of the target in each layer of the image pyramid.
[0106] Combining the image pyramid and the wavelet transform, through the spatio-temporal composite optimization algorithm, boundary tracking is performed on the foreground segmentation data to further expand and contract the target region, and finally the second object candidate region data is extracted. This method can effectively avoid the over-expansion or over-contraction of the target region through the collaborative processing of spatio-temporal domain information, ensuring the accuracy of the candidate region.
[0107] S312. For the second object candidate region data, hierarchical feature extraction is performed in combination with a deep autoencoder neural network, and second object dynamic data is generated through a dynamic flow field model and a motion clustering algorithm. The feature extraction is jointly processed by a graph convolutional network and a motion prediction algorithm;
[0108] In some embodiments, for the second object candidate region data extracted from the previous step, hierarchical feature extraction is performed through a deep autoencoder neural network. The deep autoencoder network can automatically learn the feature representation of the candidate region and map it into a low-dimensional feature space. Specifically, the deep autoencoder network consists of an encoder and a decoder. The encoder maps the input data into a low-dimensional latent space representation, and the decoder restores it back to the original input data. In this embodiment, the input data is the image data of the second object candidate region. The encoder compresses the features of this data, and the decoder reconstructs it. The specific encoding process is represented by the following mathematical formula: h = f(Wx + b), where x is the input second object candidate region data, W and b are the weight matrix and bias of the encoder respectively, h is the low-dimensional feature representation obtained through the encoder, and f(·) is the activation function. The decoding process is completed through the following formula: where is the reconstructed candidate region data, W′ and b′ are the weight matrix and bias of the decoder, and g(·) is the decoding function. Through the hierarchical learning of the encoder and decoder, high-level features of the second object can be extracted, thus better supporting the subsequent generation of dynamic data.
[0109] In this embodiment, the output feature representation of the deep autoencoder neural network will be used as the input, and further combined with a graph convolutional network to extract local structure information. The graph convolutional network can process the local structure information in the image through the graph structure to capture the relationships between different regions. For each node v i , that is, a pixel or local region in the graph, the graph convolution operation is calculated through the following formula: where represents the feature representation of node v i at the k-th layer, N(i) is the set of neighbor nodes of node v i , c ij is the normalization coefficient, W(k) and b(k) are the weights and biases of the k-th layer respectively, and σ(·) is the activation function. Through the multi-layer propagation of the graph convolutional network, local structure information of the second object can be extracted in the spatial domain.
[0110] In one embodiment, the output features of the graph convolutional network are combined with a motion prediction algorithm to generate dynamic data of the second object. The motion prediction algorithm is based on the feature information of the second object candidate region to predict the motion trajectory and state changes of the target. Specifically, a dynamic flow field model is constructed to model the motion features, and the following formula is used to predict the motion trajectory of the object: where $\hat{y}(t + 1)$ is the predicted state at time $t + 1$, $y(t)$ is the target state at time $t$, $u(t)$ is the control input (e.g., the acceleration or direction of the target), and $A$ and $B$ are the state transition matrix and the control matrix. Through the dynamic flow field model, the motion law of the second object in the spatio-temporal domain can be considered, and more accurate dynamic data can be generated.
[0111] In some embodiments, the dynamic data of the second object is clustered based on a motion clustering algorithm. The motion clustering algorithm clusters similar dynamic behaviors according to the motion characteristics of the target, thereby identifying the motion pattern of the second object. The classification of the target behavior pattern is achieved by clustering the motion data of the second object. The clustering result is generated by the following formula: where $C$ is the clustering result, $x$ i is the $i$-th data point, and $c$ k is the $k$-th clustering center. In this way, the motion characteristics of the second object are effectively extracted, providing an important basis for subsequent target behavior analysis.
[0112] Through the above steps, the generated dynamic data of the second object includes its motion state, motion trajectory, and its classification information. That is to say, the dynamic data of the second object contains two parts: time-series motion state data and the clustered motion pattern. Specifically expressed as: where $y(t)$ is the state at the original time, $\hat{y}(t)$ is the predicted state, and $C$ is the clustering result of the motion pattern.
[0113] S313. For the dynamic data of the second object, apply a multi-scale convolutional neural network based on spatio-temporal correlation analysis for time-series modeling, generate the motion trajectory of the second object, and assign a unique ID to it;
[0114] The time-series modeling of the dynamic data of the second object is carried out in a way that combines multi-dimensional convolution and graph analysis. Through the processing of cross-frame motion features, accurate trajectory prediction and ID assignment of the second object are achieved. In some embodiments, the time-series modeling process includes cross-frame motion feature extraction, spatio-temporal correlation modeling, and trajectory generation and ID assignment.
[0115] In some embodiments, the cross-frame motion feature extraction is performed by a multi-dimensional convolutional neural network. The input dynamic data contains the state information of the second object at multiple time steps. Through the convolutional neural network, the motion features between different time steps are extracted, and these features can reflect the changes and motion patterns of the second object over time. This operation can effectively extract cross-frame information from the dynamic data and capture the dynamic behavior of the second object.
[0116] In some embodiments, spatio-temporal correlation modeling is achieved by combining a multi-scale convolutional neural network with graph analysis. The multi-scale convolutional neural network is used to extract motion features at different time scales. Through convolutional operations at different scales, the motion patterns of the second object at different moments are obtained. Graph analysis combines spatial and temporal information. By modeling the relationship between space and time and considering the interactions between objects, it can capture the complex dynamic behavior of the second object from the spatio-temporal dimension, providing complete spatio-temporal correlation information for subsequent trajectory modeling.
[0117] In some embodiments, the results of temporal modeling are integrated by the weighted average method to generate the movement trajectory of the second object. The movement trajectory reflects the actual motion pattern of the second object in different time periods. The trajectory of each target is fitted according to the weighted average value of spatio-temporal features to generate a motion trajectory that conforms to the actual motion law. Based on the motion trajectory, in some embodiments, a unique ID can be assigned to each target to ensure that the second object can be uniquely identified and tracked in subsequent monitoring and analysis processes.
[0118] S4. Perform trajectory filtering on the motion trajectory of the first object under the unique ID, identify abnormally stationary targets, and output a dynamic retention warning signal.
[0119] Specifically, please refer to the appendix Figure 5 As shown, the steps include the following sub-steps:
[0120] S401. Perform time series analysis on the motion trajectory of the first object under the unique ID, calculate the stationary duration within each time period, and compare it with a preset reference motion model to obtain stationary behavior deviation data.
[0121] In some embodiments, the dynamic time warping algorithm is used for time series analysis. This algorithm solves the time series matching problem caused by misaligned time axes by measuring the similarity between the target at different time points. Through this algorithm, the movement differences between the target at different time points can be calculated, and the stationary situation of the target during the monitoring period can be evaluated based on these differences. If the movement differences of the target within a certain time period are very small, the target is considered to be in a stationary state. The stationary duration is calculated by statistically summing up the time when the target is stationary during the monitoring period and is used for subsequent analysis of the target's behavior.
[0122] In some embodiments, the preset reference motion model represents the expected motion pattern of the target in the normal state. This model analyzes the motion trajectories of a large number of normal targets to obtain the normal motion characteristics of the target, such as motion speed and acceleration. Under this reference model, the target should maintain a certain motion state within a certain period of time and should not be stationary for too long. By comparing the reference model with the actual stationary duration of the target, stationary behavior deviation data can be obtained. If the stationary duration of the target exceeds the normal range in the reference model, it indicates that the target has abnormal stationary behavior.
[0123] In some embodiments, the stationary behavior deviation data reflects the difference between the stationary behavior of the target and the normal motion model. By calculating the difference between the actual stationary duration of the target and the stationary duration in the reference model, the stationary behavior deviation of the target is obtained. This deviation data can provide a basis for subsequent abnormal behavior judgment and help identify whether the target has potential safety hazards or abnormal behaviors.
[0124] S402, perform cumulative processing on the stationary behavior deviation data under the unique ID, calculate the cumulative deviation values for multiple time periods, and generate a deviation change rate according to the change of the cumulative deviation between adjacent time periods. The deviation change rate is generated through differential analysis and weighted moving average algorithm;
[0125] In this embodiment, the stationary behavior deviation data ΔT static (t) is cumulatively processed within multiple time periods to obtain the cumulative deviation value for each time period. Set the time period t i The corresponding stationary behavior deviation data is ΔT static (t i ) By cumulatively processing these data, the cumulative deviation C static (t) is obtained: Where C static (t) is the cumulative deviation value at the t-th moment, indicating the cumulative situation of the stationary behavior deviation of the target from the initial moment to the current moment. The cumulative deviation value reflects the total deviation of the target's stationary behavior and provides a measure of the difference between the target's motion pattern and normal behavior for subsequent analysis.
[0126] In some embodiments, the deviation change rate is used to quantify the deviation change between adjacent time periods. The deviation change rate R static (t) is generated through differential analysis and weighted moving average algorithm. Set the cumulative deviation values of two adjacent time periods t i-1 and t i to be C static (t i-1 ) and C static (t i ), then the deviation change rate R static(t) is: Rate of change of deviation R static (t) is obtained by calculating the difference of the cumulative deviations in adjacent time periods, and the change trend of the target stationary behavior is obtained. In this embodiment, the weighted moving average algorithm is used to smooth this rate of change to reduce the influence of noise interference on the calculation result of the rate of change. The calculation formula of the weighted moving average is: where R smooth (t) is the smoothed rate of change of deviation, w k is the weighting coefficient, representing the weight of the rate of change at the previous moment, and N is the length of the weighting window. The weighted moving average algorithm performs weighted averaging on the rate of change through the weighting window, avoiding the influence of outliers on the calculation result and obtaining the smoothed rate of change.
[0127] In this embodiment, the cumulative deviation and the rate of change of deviation are comprehensively used to detect the abnormality of the target stationary behavior. By performing cumulative processing on the stationary behavior deviation data and generating the rate of change of deviation, the evolution trend of the target stationary behavior is analyzed, and whether there are signs of long-term stationary or abnormal stationary behavior is identified.
[0128] S403. According to the cumulative deviation value and the rate of change of deviation under the unique ID, construct a spatio-temporal change trend model of the target stationary behavior. The model is modeled by a spatio-temporal convolutional network to obtain the change trend of the target stationary behavior;
[0129] In some embodiments, the spatio-temporal convolutional network combines the learning of spatial features and time series features, and can effectively capture the change patterns of the target behavior in the spatial and time dimensions. By inputting the cumulative deviation value and the rate of change of deviation data, the network can simultaneously analyze the stationary distribution of the target in space and the change law in time, so as to extract the spatio-temporal features of the target stationary behavior. Specifically, the structure of the spatio-temporal convolutional network consists of multiple convolutional layers, and each convolutional layer gradually extracts the spatial and time features of the target stationary behavior.
[0130] In this embodiment, the input data of the spatio-temporal convolutional network includes the stationary deviation (cumulative deviation value) of the target in different time periods and the change trend of the deviation (rate of change of deviation). Through the convolution operation, the network can perform joint analysis of these two kinds of information in the spatio-temporal domain, and gradually extract the spatio-temporal features of the target stationary behavior. Each convolution operation learns the spatial and time dependence relationships in the data, so as to better capture the evolution law of the target stationary behavior.
[0131] The output of the spatio-temporal convolutional network is the spatio-temporal change trend of the target's stationary behavior. This output represents the stationary behavior pattern of the target in the time and space dimensions, as well as the change trend of this behavior pattern. Through the multi-layer convolutional structure, the network gradually extracts and optimizes these features, thus accurately revealing the changes in the target behavior at different time periods and different positions.
[0132] In some embodiments, the output of the spatio-temporal change trend model provides a dynamic spatio-temporal change view of the target's stationary behavior. Based on these outputs, it is further determined whether the target is in an abnormal stationary state. For example, if the output of the model reflects a sudden change in the target behavior or an abnormal stationary pattern, it is considered that the target has abnormal stationary behavior. The results of this model will serve as the basis for subsequent anomaly detection, helping to identify potential safety hazards in a timely manner.
[0133] In this embodiment, the spatio-temporal change trend model obtained through the spatio-temporal convolutional network can capture the evolution law of the target's stationary behavior in both the space and time dimensions, and effectively determine whether it is an abnormal behavior.
[0134] S404, perform causal reasoning analysis on the spatio-temporal change trend model under the unique ID, generate target abnormal behavior determination data, and perform comprehensive analysis in combination with the multi-object decision-making network to output a dynamic stay warning signal.
[0135] In this embodiment, the causal reasoning analysis uses the stationary behavior pattern data extracted from the spatio-temporal change trend model T trend (t) to conduct causal relationship derivation. This analysis constructs a causal graph model G, where the nodes of the graph represent stationary behavior states and the edges represent the causal relationships between behaviors.
[0136] Set the stationary behavior deviation as the behavior deviation of the target within the time period t. In this embodiment, let the causal reasoning model of the target behavior be: where pa i is the parent node of node i, representing the causal factor of this node. The parameters in the causal graph are optimized through maximum likelihood estimation to obtain the causal reasoning result of the target behavior. Thus, according to the behavior pattern output by the causal reasoning, it is determined whether the target has abnormal stationary behavior.
[0137] In some embodiments, the multi-object decision-making network is combined to perform comprehensive analysis on the target abnormal behavior determination data. The multi-object decision-making network combines the behavior information of multiple targets and can make reasonable decisions in a complex environment. For each target, the network first comprehensively evaluates the behavior deviation of the target based on the output of the spatio-temporal change trend model and the causal reasoning result.
[0138] The core of the multi-objective decision-making network lies in optimizing the decision weights of each target behavior to ensure that the output results of the network can comprehensively consider the dynamic behavior information of multiple targets. Let W target be the decision weight of target i, be the final behavior deviation of target i, and the decision rule of the target is: where N is the number of targets, and W target (j) is the weight of target j. By performing weighted summation on the behavior data of multiple targets, a comprehensive abnormal behavior assessment for each target is finally generated.
[0139] In this embodiment, the comprehensive abnormal behavior assessment of the target will be compared with a preset abnormal behavior threshold . If , it is considered that the target has abnormal stationary behavior. This determination result will further generate a dynamic stay warning signal.
[0140] S5. When outputting the dynamic stay warning signal, perform screening and matching analysis on the movement trajectory of the second object, identify trajectory anomalies, and trigger an abnormal behavior alarm signal and associated response actions.
[0141] Specifically, in the safety management of the crossing, analyze the stationary behavior of the first object, calculate data such as its stationary duration, deviation value, and change rate to confirm whether there is an abnormal stay situation. When a stay warning signal is detected, the system uses it as a preliminary warning of potential risks, indicating that there may be problems in this area. However, this stay signal does not directly represent a safety threat and requires further observation. Next, the system performs screening and matching analysis on the behavior trajectory of the second object to determine whether there is tailing or other abnormal behaviors. When the behavior of the second object deviates from the preset normal mode or detects non-compliant behaviors, the system will trigger an alarm signal and associated response actions. This process improves the accuracy and timeliness of processing by first prompting potential risk signals and then responding when actual threats occur, avoiding false alarms and effectively improving the safety management level. Please refer to the appendix Figure 6 as shown, this step includes the following sub-steps:
[0142] S501. When outputting the dynamic stay warning signal, screen the movement trajectory of the second object based on the independent ID of the second object and the spatio-temporal constraint model. The spatio-temporal constraint model generates a difference data set between the trajectory anomaly points and the normal behavior trajectory through dynamic spatio-temporal window and spatial neighborhood clustering analysis, thereby generating an effective target data set;
[0143] In this embodiment, the spatio-temporal constraint model constructs a multi-dimensional space model that can capture the change characteristics of the target trajectory in combination with the dynamic behavior information of the second object. Let X i(t) = (x i (t), y i (t)) is the position data of the second object at time t, where x i (t) and y i (t) represent coordinates in a two-dimensional space. By introducing a spatio-temporal window mechanism, local feature modeling of the target trajectory is performed. The spatio-temporal window is defined as W(t) = [t - ξt, t + ξt], where ξt is the time window size. The movement data of the target within the time interval [t - ξt, t + ξt] is extracted through the window, and its spatio-temporal characteristics are analyzed. In some embodiments, the size of the spatio-temporal window is dynamically adjusted according to the speed and behavior characteristics of the target trajectory. In this embodiment, the application of the dynamic spatio-temporal window is based on the speed and direction information of the target within a given time period.
[0144] Within the spatio-temporal window, the target trajectory data X i (t) will be subjected to clustering analysis according to its position change characteristics to screen out possible abnormal points in the trajectory. The abnormal point judgment formula is set as: ΔX(t) = ||X i (t) - X norm (t)||, where X norm (t) is the estimated point of the normal behavior trajectory, and the criterion for abnormal points is that the distance between the trajectory point and the normal trajectory exceeds a certain set threshold, that is: ΔX(t) > ε. When the trajectory point X i (t) of the target and the normal trajectory X norm (t) exceeds the threshold ε, it is determined as an abnormal point.
[0145] In some embodiments, spatial neighborhood clustering analysis further determines the abnormal situation of the trajectory data by performing regional clustering on the target trajectory within the spatio-temporal window. The clustering method is based on the spatial distribution of the target trajectory, and the trajectory points are divided into multiple clusters through a clustering algorithm. In this embodiment, the density-based clustering algorithm is used for spatial neighborhood clustering analysis. The spatial distance metric of the target trajectory is set as: The DBSCAN algorithm calculates the distance d(X i (t), X j (t)) between trajectory points, groups the trajectory points with closer distances into the same cluster, and determines whether there are noise points or abnormal points. By setting appropriate neighborhood radius ε and minimum number of points MinPts, abnormal points in the trajectory can be effectively identified.
[0146] In this embodiment, the trajectory data screened by the dynamic spatio-temporal window and spatial neighborhood clustering analysis will generate an effective target data set. Let D val id be the effective target data set, which contains the screened target trajectory points X i(t), these trajectory points represent the difference dataset of the target between the normal behavior trajectory and the abnormal behavior trajectory. The abnormal behavior characteristics of the target are determined by comparing the deviation of the abnormal points from the normal trajectory.
[0147] In some embodiments, the generation of the effective target dataset is based on the following conditions: when the deviation of the target trajectory points within the spatio-temporal window is greater than the threshold, and the point is determined to be an abnormal point by the clustering algorithm, then the trajectory point is regarded as a part of the effective target data. Through multiple iterations of screening and optimization, it can be ensured that the dataset only contains trajectory points with significant behavior deviations.
[0148] S502, use the adaptive trajectory similarity matching algorithm to perform similarity comparison on the effective target dataset, compare the effective target dataset with the preset normal action trajectory template, and generate a matching analysis result;
[0149] In some embodiments, the similarity comparison includes feature extraction processing and similarity calculation processing. Specifically, the effective target dataset will be compared with the preset normal action trajectory template to generate a matching analysis result.
[0150] In this embodiment, the feature extraction processing is performed by extracting spatio-temporal trajectory features from the effective target dataset. Set the time series data of the target trajectory as T i ={X i (t 1 ), X i (t 2 ), …, X i (t n )}, where X i (t) represents the position of the target at time t, and n is the length of the trajectory. Feature extraction is performed on the trajectory data, including position, speed, and acceleration information.
[0151] Set the speed of the trajectory as: where Δt is the time interval, and X i (t + Δt) and X i (t) represent the target positions at times t and t + Δt respectively. The acceleration is calculated by the following formula: In some embodiments, the feature extraction may also include the change of the direction angle of the target trajectory and the motion smoothness information. These features can effectively describe the behavior characteristics of the target and will not be elaborated here. On this basis, a high-dimensional feature vector of the target is constructed based on the trajectory features at each time point.
[0152] In this embodiment, the similarity calculation process is processed by an adaptive trajectory similarity matching algorithm, aiming to identify whether there is abnormal behavior by measuring the similarity between the effective target data set and the normal action trajectory template. Set the preset normal action trajectory template as T norm ={X norm (t 1 ), X norm (t 2 ), …, X norm (t m )}, where m is the length of the template trajectory.
[0153] The adaptive trajectory similarity matching is calculated by the dynamic time warping algorithm to measure the similarity between two trajectories. The DTW algorithm finds the optimal matching path by minimizing the cumulative distance between the trajectories. Set D(i, j) as the distance between the trajectory point X i (t) and the template point X norm (t), which is represented by the Euclidean distance: D(i, j)=||X i (t)-X norm (t)||. The goal of DTW is to calculate the minimum cumulative distance through dynamic programming, and then compare the overall similarity between the target trajectory and the template trajectory. The matching result of DTW is represented by the following optimal path P: In some embodiments, the DTW algorithm assigns different weights to the trajectory features in different periods by introducing a weighting strategy, so as to better reflect the difference between the target trajectory and the normal behavior trajectory. Set the weighting factor as w(t), and modify the distance metric formula as: D(i, j)=w(t)||X i (t)-X norm (t)||, where w(t) is adaptively adjusted according to the dynamic changes of the target trajectory to highlight the trajectory features in key periods.
[0154] In this embodiment, the similarity between the target trajectory and the normal behavior trajectory template is calculated by the DTW algorithm, and whether there is behavior abnormality is determined through a threshold. Set the similarity metric as: where S is the similarity value of the trajectory match, and D(i, j) is the distance between the trajectory points. The smaller the similarity value S, the more similar the target trajectory is to the normal trajectory; the larger the similarity value S, the greater the difference between the target trajectory and the normal trajectory.
[0155] When the similarity value exceeds the preset threshold α, it is determined that the target trajectory has a deviation, and a matching analysis result is generated, indicating that the target may have abnormal behavior. Through this step, the target deviating from the normal trajectory can be accurately identified, and data support is provided for subsequent abnormal behavior determination.
[0156] S503 processes the matching results based on the spatio-temporal deviation fusion analysis model, identifies the targets whose deviation from the preset trajectory exceeds the threshold, outputs an abnormal behavior alarm signal, and triggers an associated response action.
[0157] In some embodiments, the spatio-temporal deviation fusion analysis model includes deviation value extraction processing and spatio-temporal fusion calculation processing.
[0158] In this embodiment, the deviation value extraction processing is performed by calculating the difference between the target trajectory and the normal trajectory from the matching results. The similarity metric values between the target trajectory and the normal trajectory are set as follows: Where X i (t) and X norm (t) are the positions of the target trajectory and the normal trajectory at time t respectively. Thus, the difference metric from the normal trajectory. On this basis, by calculating the deviation values between each target trajectory point and the template trajectory point, a set of deviation data sets D diff is formed, representing the degree to which the target trajectory deviates from the normal trajectory. In this embodiment, the spatio-temporal fusion calculation processing is used to further process the deviation data, considering the comprehensive influence of time and space factors to identify abnormal targets. For this purpose, a spatio-temporal deviation fusion algorithm is adopted. The spatial position of the target trajectory at time t is set as X i (t), its deviation value in the time dimension is Δt i , and the spatial deviation value is Δs i . The spatio-temporal deviation fusion calculation is carried out through the following formula: Δ t,s (t) = α·Δt i + β·Δs i , where Δ t,s (t) is the comprehensive deviation, and α and β are weighting factors representing the weights of time deviation and spatial deviation in the spatio-temporal deviation fusion. In some embodiments, the weights α and β are dynamically adjusted according to the actual application scenario to highlight the behavioral characteristics under specific spatio-temporal conditions. In this way, the spatio-temporal deviation of the target trajectory is effectively fused, and the overall difference between the trajectory and the normal trajectory is further quantified.
[0159] In this embodiment, after the spatio-temporal deviation fusion calculation, a deviation threshold ζ is further introduced to set a threshold interval. When the comprehensive deviation Δ t,s (t) exceeds this threshold, it indicates that the degree to which the target trajectory deviates from the normal trajectory is large, thereby identifying the abnormal behavior target. The specific judgment formula is as follows: If AbnormalityScore > ζ, then it is determined that the target behavior is abnormal, and an abnormal behavior alarm signal is generated.
[0160] In this embodiment, when the generated abnormal behavior alarm signal is triggered, the response action will be judged based on the spatio-temporal characteristics of the target trajectory and preset rules. The response action may include but is not limited to starting safety devices, issuing alarms, recording trajectory information, etc.
[0161] Based on the description of the above embodiments of the video capture-based airport crossing safety management method, the embodiments of the present application also disclose a video capture-based airport crossing safety management system. The video capture-based airport crossing safety management system may be a computer program (including program code) that runs the above-mentioned video capture-based airport crossing safety management method. Please refer to the appendix Figure 7 As shown, the video capture-based airport crossing safety management system may run the following units:
[0162] An acquisition unit 110, configured to acquire real-time data to be processed, where the real-time data to be processed is video crossing picture data captured by a video monitoring device in real time;
[0163] A preprocessing unit 120, configured to preprocess the real-time data to be processed to obtain real-time preprocessed data;
[0164] An identification unit 130, configured to perform identification processing on the first object and the second object for the real-time preprocessed data, generate dynamic data of the first object and the second object, and generate their respective motion trajectories based on the dynamic data, and assign independent IDs to each target;
[0165] A detention detection unit 140, configured to perform trajectory filtering processing on the moving trajectory of the first object under the unique ID, identify abnormal stationary targets, and output a dynamic detention warning signal;
[0166] An abnormality detection unit 150, configured to perform screening and matching analysis on the moving trajectory of the second object when the dynamic detention warning signal is output, identify trajectory abnormalities, and trigger an abnormal behavior alarm signal and an associated response action.
[0167] The above is only the preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be within the scope of the concept described herein, through the above teachings or the technology or knowledge in the relevant field. And the changes and modifications made by those skilled in the art that do not depart from the spirit and scope of the present invention should all be within the protection scope of the appended claims of the present invention.
Claims
1. A method for airport crossing safety management based on video capture, characterized in that: The method comprises the following steps: S1, obtaining real-time data to be processed, wherein the real-time data to be processed is video crossing picture data captured in real time by a video surveillance device; S2, preprocessing the real-time data to be processed to obtain real-time preprocessed data; S3 performs recognition processing of the first object and the second object on the real-time preprocessed data, generates dynamic data of the first object and the second object, performs modeling based on the dynamic data to generate respective motion trajectories, and assigns an independent ID to each target; S4, performing trajectory filtering processing on the movement trajectory of the first object under the unique ID, identifying abnormal stationary targets and outputting a dynamic retention warning signal; S5, when outputting the dynamic retention warning signal, screening and matching analysis are performed on the movement trajectory of the second object, the trajectory anomaly is identified and the abnormal behavior alarm signal and associated response action are triggered.
2. The method for airport crossing safety management based on video capture according to claim 1, characterized in that: The S2 step includes the following sub-steps: S201, performing a frame extraction operation on the real-time data to be processed, decomposing the continuous video images into discrete frame sequences and performing normalization processing to obtain discrete normalized data, performing Gaussian filtering to remove noise on the discrete normalized data, and generating denoised frame data; S202, performing background modeling and foreground segmentation processing on the denoised frame data, extracting the motion foreground area using a convolutional neural network segmentation algorithm to generate foreground segmentation data, adding a timestamp mark to each frame of the foreground segmentation data, and generating real-time preprocessing data.
3. The method for airport crossing safety management based on video capture according to claim 1, characterized in that: The first object recognition process in step S3 includes the following sub-steps: S301, performing spatiotemporal domain separation on the foreground segmentation data in the real-time preprocessing data, filtering the motion features of the foreground area by an adaptive spatiotemporal filtering algorithm, and generating motion pattern data, wherein the spatiotemporal domain separation is performed by combining a convolutional neural network and LSTM; S302, applying a partition processing algorithm based on tensor decomposition based on the motion pattern data to dynamically construct a target motion attribute model in the area, thereby generating first object dynamic data, wherein the partition processing algorithm is implemented by local maximum clustering and adaptive threshold algorithm; S303, performing motion trajectory modeling on the first object dynamic data using an optimized Kalman filter algorithm based on time-frequency joint analysis, generating a moving trajectory of the first object, and assigning an independent ID to it, wherein the trajectory modeling adopts a fusion of weighted mean filtering and a multi-dimensional Gaussian model.
4. The method for airport crossing safety management based on video capture according to claim 1, characterized in that: The recognition process of the second object in step S3 includes the following sub-steps: S311, generating a local saliency map for the foreground segmentation data in the real-time preprocessing data, and expanding and contracting the object region by a boundary tracking method based on a space-time composite optimization algorithm to extract second object candidate region data; S312, performing hierarchical feature extraction on the candidate region data of the second object in combination with a deep autoencoder neural network, generating second object dynamic data through a dynamic flow field model and a motion clustering algorithm, wherein the feature extraction is jointly processed through a graph convolutional network and a motion prediction algorithm; S313: Apply a multi-scale convolutional neural network based on spatiotemporal correlation analysis to the dynamic data of the second object to perform time series modeling, generate a movement trajectory of the second object, and assign an independent ID to it.
5. The method for airport crossing safety management based on video capture according to any one of claims 1 to 4, characterized in that: The S4 step includes the following sub-steps: S401, performing time series analysis on the movement trajectory of the first object under the unique ID, calculating the static duration in each time period, and comparing it with a preset reference motion model, thereby obtaining static behavior deviation data; S402, accumulating the stationary behavior deviation data under the unique ID, calculating the accumulated deviation values of multiple time periods, and generating a deviation change rate according to the accumulated deviation changes of adjacent time periods, wherein the deviation change rate is generated by differential analysis and weighted moving average algorithm; S403, constructing a spatiotemporal change trend model of the target stationary behavior according to the accumulated deviation value and the deviation change rate under the unique ID, wherein the model is modeled by a spatiotemporal convolutional network to obtain a change trend of the target stationary behavior; S404, performing causal reasoning analysis on the spatiotemporal variation trend model under the unique ID, generating target abnormal behavior determination data, and combining a multi-objective decision network to perform comprehensive analysis and output a dynamic retention warning signal.
6. The method for airport crossing safety management based on video capture according to claim 5 is characterized in that: The S5 step includes the following sub-steps: S501, when a dynamic detention warning signal is output, the movement trajectory of the second object is screened based on the independent ID of the second object and the spatiotemporal constraint model, wherein the spatiotemporal constraint model screens out the difference data set between the trajectory abnormal points and the normal behavior trajectory through dynamic spatiotemporal window and spatial neighborhood clustering analysis to generate a valid target data set; S502, using an adaptive trajectory similarity matching algorithm to perform similarity comparison on a valid target data set, comparing the valid target data set with a preset normal action trajectory template, and generating a matching analysis result; S503, processing the matching results based on the spatiotemporal deviation fusion analysis model, identifying that the deviation from the preset trajectory exceeds the threshold, outputting an abnormal behavior alarm signal, and triggering an associated response action.
7. The method for airport crossing safety management based on video capture according to claim 4 is characterized in that: The temporal modeling process in S313 includes cross-frame motion feature extraction, spatiotemporal correlation modeling, trajectory generation and ID allocation.
8. The method for airport crossing safety management based on video capture according to claim 5, characterized in that: The time series analysis in step S401 adopts a dynamic time warping algorithm.
9. The method for airport crossing safety management based on video capture according to claim 6, characterized in that: The similarity comparison in step S502 includes feature extraction processing and similarity calculation processing. The similarity calculation processing is performed using an adaptive trajectory similarity matching algorithm. The adaptive trajectory similarity matching is calculated using a dynamic time warping algorithm. The spatiotemporal deviation fusion analysis model in step S503 includes deviation value extraction processing and spatiotemporal fusion calculation processing.
10. An airport crossing safety management system based on video capture, characterized in that: The system comprises: An acquisition unit is used to acquire real-time data to be processed, wherein the real-time data to be processed is video crossing picture data captured in real time by a video surveillance device; A preprocessing unit, used for preprocessing the real-time data to be processed to obtain real-time preprocessed data; an identification unit, configured to perform identification processing of the first object and the second object on the real-time preprocessed data, generate dynamic data of the first object and the second object, perform modeling based on the dynamic data to generate respective motion trajectories, and assign an independent ID to each target; A detention detection unit, configured to perform a trajectory filtering process on the movement trajectory of the first object under the unique ID, identify abnormal stationary targets and output a dynamic detention warning signal; The anomaly detection unit is used to screen and match the movement trajectory of the second object when outputting the dynamic retention warning signal, identify the trajectory anomaly and trigger the abnormal behavior alarm signal and the associated response action.
Citation Information
Patent Citations
Video synopsis generation method capable of solving multi-target collision and occlusion problem
CN106856577A
Vision and IMU sensor fusion positioning system based on dynamic object semantic segmentation
CN113223045A
Multi-degree-of-freedom mechanical arm obstacle avoidance control method based on machine vision
CN119188770A
Intelligent machine vision detection method and system based on image processing and storage medium
CN119205719A
Multithreat safety and security system and specification method thereof
US20090207020A1
Cited By
Video data processing method and system based on digital twinborn scene
CN120321433A