Security check action recognition method and system based on dynamic time warping
Patent Information
- Application Number
- CN202611104775.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-24
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-24
AI Technical Summary
[0006]本发明旨在提供基于动态时间规整的安检动作识别方法及系统,通过状态机前向时序导数触发,结合首行归零的子序列动态时间规整增量计算与后向路径回溯的协同措施,解决了现有技术在不间断的流式安检监控视频中,由于无法预知动作起止点而导致无法对未分割连续视频流进行实时、高精度自动分割与合规性识别的问题
传统动态时间规整(DTW)算法极度依赖预先人工裁剪的离散动作片段,面对未分割的连续监控视频流时往往因找不到起止点而失效。本发明通过前向粗触发与后向精回溯的双向钳制机制,先利用轻量级状态机的一阶导数自适应卡下动态起始帧,再通过首行归零的子序列DTW结合最优路径反向回溯,在不间断的源源不断视频流中,自动、精准地剥离出目标动作的法定起止帧,实现了流式数据的吞吐式自动分割与合规识别。
Smart Images

Figure CN122657976B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent security and pattern recognition technology, specifically to a method and system for security inspection action recognition based on dynamic time warping. Background Technology
[0002] Traditional security checks heavily rely on manual monitoring or on-site supervision. With the development of deep learning and computer vision, utilizing video surveillance for automated action recognition—such as detecting missed checks, violations, and verifying compliance with standard security procedures—has become a crucial means of improving security efficiency and safety. In this context, Dynamic Time Warping (DTW) has been introduced into action recognition. DTW is an algorithm used to measure the similarity between two time series, and its core advantage lies in its ability to handle non-linear time scaling.
[0003] In security screening scenarios, different security personnel often perform the same standard actions, such as metal detector scanning and baggage handling, with varying speeds, rhythms, and durations. DTW (Time-Warping Wave) uses "time axis warping" to find the optimal alignment path between two sequences of actions of different lengths and speeds, thereby calculating the similarity of the actions and achieving accurate identification.
[0004] While DTW-based action recognition methods have unique advantages in handling temporal misalignments, they struggle to effectively handle the "multi-segmentation" and "start-endpoint identification" of actions in practical applications, especially in complex security inspection environments. In actual security checks, security personnel's actions are continuous and uninterrupted. Existing technologies often require pre-editing clean action segments before DTW matching. When faced with unsegmented continuous video streams, traditional DTW struggles to automatically and accurately extract the target security action from uncertain start-endpoints, easily leading to overall matching failure.
[0005] Therefore, to address the problem of being unable to handle unsegmented continuous video streams, a security inspection action recognition method and system based on dynamic time warping is proposed to solve the above problem. Summary of the Invention
[0006] This invention aims to provide a security inspection action recognition method and system based on dynamic time warping. By triggering the forward temporal derivative of the state machine, combined with the collaborative measures of dynamic time warping increment calculation of the subsequence with the first row returning to zero and backward path backtracking, it solves the problem that existing technologies cannot perform real-time, high-precision automatic segmentation and compliance recognition of unsegmented continuous video streams in uninterrupted streaming security inspection monitoring videos due to the inability to predict the start and end points of actions.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: Security check action recognition methods based on dynamic time warping include: S1. Obtain the continuous monitoring video stream of the security check area, and extract the frame-by-frame action features of the target personnel in the continuous monitoring video stream in time sequence to construct a real-time action feature sequence; S2. Input the real-time action feature sequence into a preset security inspection action state machine, calculate the temporal first derivative of the frame-by-frame action features between the current frame and the previous frame to characterize the action change rate, and when the temporal first derivative satisfies the preset state transition condition, activate the matching trigger signal in the state machine and determine the dynamic start frame of the sliding window. S3. Using the dynamic start frame as the boundary, intercept the real-time action feature sequence to form the current test window sequence, and perform subsequence dynamic time warping incremental calculation on the current test window sequence and the pre-stored standard security inspection action template sequence to construct the cumulative distance matrix of the open endpoint. S4. In the latest calculated column of the cumulative distance matrix, search for the minimum cumulative distance value that satisfies the preset similarity threshold; if it exists, determine that the target security check action has been captured in the continuous monitoring video stream, and backtrack according to the matrix position where the minimum cumulative distance value is located, automatically extract the end frame of the target security check action, and determine the dynamic start frame as the start frame of the target security check action, so as to complete the automatic segmentation and action recognition of the continuous video stream.
[0008] Preferably, the frame-by-frame motion features in S1 are multimodal fusion features, and their extraction and construction process includes: Extract the spatiotemporal trajectory of key points of the target person's two-dimensional or three-dimensional human skeleton to form the first feature vector; Extract the spatiotemporal motion trajectory and speed changes of the security inspection equipment held by the target personnel to form a second feature vector; The first feature vector and the second feature vector are adaptively weighted and concatenated to generate the frame-by-frame action features.
[0009] Preferably, when generating the frame-by-frame motion features, the weighting coefficients of the first feature vector and the second feature vector are dynamically adjusted according to the real-time flow of people in the security checkpoint; when the flow of people is greater than a preset density threshold, the weight of the second feature vector is increased to suppress noise interference caused by human skeletal points occlusion.
[0010] Preferably, the security inspection action state machine in S2 includes an idle state, a ready state, and a matching state; the activation of the matching trigger signal in the state machine when the first derivative of the timing satisfies the preset state transition condition specifically includes: When the state machine is in the idle state, if the first derivative of the timing exceeds the set static threshold and the duration reaches the preset number of frames, the state machine will switch from the idle state to the ready state. In the ready state, if the spatiotemporal motion trajectory of the security inspection device in the second feature vector meets the preset start-up acceleration feature, and the key points of the human skeleton represented by the first feature vector enter the preset matching trigger area, the state machine jumps to the matching state, activates the matching trigger signal, and marks the current frame as the dynamic start frame of the sliding window.
[0011] Preferably, the constraints for calculating the dynamic time warping increment of the subsequence in S3 include: The first row of the cumulative distance matrix is initialized to zero to allow the starting point of the standard security action template sequence to adaptively align with any frame in the current test window sequence; A Sakoe-Chiba constraint band is introduced to limit the search space of the cumulative distance matrix. The width of the constraint band is dynamically adjusted according to the action execution rate reflected by the fluctuation period of the temporal change rate of the real-time action feature sequence within a preset time window.
[0012] Preferably, the cumulative distance matrix for constructing open endpoints in step S3 employs a forward one-dimensional rolling update mechanism: Each time a new action feature is input, the system calculates the single-column distance data corresponding to the latest frame only for the cumulative distance matrix, and overwrites or releases the memory of the invalid historical columns, so that the space complexity remains O(M), where M is the length of the standard security check action template sequence.
[0013] Preferably, the preset similarity threshold in S4 adopts a dynamic adaptive adjustment strategy, specifically including: The thickness characteristics of the clothing of the passenger being inspected are extracted and analyzed by infrared imaging and edge contour recognition. The preset similarity threshold is adjusted based on the thickness level of the clothing. When the thickness level of the clothing is greater than the preset thickness threshold, the preset similarity threshold is increased by a preset ratio based on the original basic threshold. Conversely, the preset ratio is decreased or the basic threshold is kept unchanged to compensate for the deformation of the security inspection action caused by the thickness of the passenger's clothing.
[0014] Preferably, S4 is followed by S5: Extract the optimal twist path slope of the automatically segmented target security inspection action segment in the cumulative distance matrix; The slope of the twisted path is compared with the standard compliance rate. If the slope deviates from the preset compliance range, a compliance evaluation warning signal for abnormal action rate is output.
[0015] Furthermore, a security check action recognition system based on dynamic time warping includes: Feature extraction module: used to acquire continuous surveillance video stream of the security check area, and extract frame-by-frame action features of the target personnel in the continuous surveillance video stream in time sequence to construct a real-time action feature sequence; State control module: used to input the real-time action feature sequence into a preset security check action state machine, calculate the temporal first derivative of the frame-by-frame action features between the current frame and the previous frame to characterize the action change rate, and when the temporal first derivative meets the preset state transition conditions, activate the matching trigger signal in the state machine and determine the dynamic start frame of the sliding window. Incremental alignment module: used to intercept the real-time action feature sequence with the dynamic start frame as the boundary to form the current test window sequence, and to perform subsequence dynamic time warping incremental calculation on the current test window sequence and the pre-stored standard security inspection action template sequence to construct the cumulative distance matrix of the open endpoint; The segmentation and recognition module is used to search for the minimum cumulative distance value that satisfies a preset similarity threshold in the latest calculated column of the cumulative distance matrix. If it exists, it determines that a target security check action has been captured in the continuous monitoring video stream. Based on the matrix position of the minimum cumulative distance value, it backtracks backward to automatically extract the end frame of the target security check action and determines the dynamic start frame as the start frame of the target security check action, so as to complete the automatic segmentation and action recognition of the continuous video stream.
[0016] Furthermore, when the computer program is executed by the processor, it implements the security check action recognition method based on dynamic time warping as described above.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: Traditional Dynamic Time Warping (DTW) algorithms heavily rely on pre-cut discrete action segments, often failing when faced with unsegmented continuous surveillance video streams due to the inability to identify start and end points. This invention employs a bidirectional clamping mechanism of forward coarse triggering and backward fine backtracking. It first uses the first derivative of a lightweight state machine to adaptively determine the dynamic start frame, and then combines the first-line zeroed subsequence DTW with optimal path backtracking. This automatically and accurately extracts the legally defined start and end frames of the target action from the continuous video stream, achieving throughput-based automatic segmentation and compliance identification of streaming data.
[0018] This invention employs a forward one-dimensional rolling update algorithm, updating only one column per frame and overwriting invalid memory in real time, thus locking the space complexity to a constant level of O(M). Simultaneously, it introduces a Sakoe-Chiba constraint band based on adaptive widening of action rate, confining the computational load of a single frame to a local narrow channel, significantly reducing the power consumption of the edge industrial control chip, and perfectly eliminating the latency of streaming processing at the security inspection site.
[0019] This approach overcomes the shortcomings of traditional methods, which are prone to false alarms in congested and deformed scenarios. At the feature layer, an adaptive weighting mechanism based on the dual-modal features of posture and equipment, based on pedestrian flow density, is designed to achieve anti-occlusion compensation by observing posture in smooth pedestrian flow and tool trajectory in severe congestion. At the decision layer, infrared and contour technology are used to dynamically analyze the thickness of passengers' clothing, and the decision threshold is adjusted up or down in segments to eliminate motion deformation noise caused by seasonal changes and clothing obstruction, ensuring high reliability in all weather and all seasons.
[0020] This invention mines the geometric value of the optimal twisted path and uses a slope derivative discrimination mechanism. By comparing the path slope with the standard compliant speed range, it can instantly identify and intercept high-risk violations such as security personnel intentionally speeding up or scrambling to complete inspections due to fatigue or perfunctory work. It outputs accurate compliance assessment and audit reports, thereby improving the actual security level of intelligent security systems. Attached Figure Description
[0021] Figure 1 The flowchart shows a security check action recognition method based on dynamic time warping. Figure 2 A geometric diagram illustrating the adaptive dynamic stitching principle for multimodal feature space. Detailed Implementation
[0022] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0023] Reference Figure 1 As shown, the security check action recognition method based on dynamic time warping includes: S1. Obtain the continuous monitoring video stream of the security check area, and extract the frame-by-frame action features of the target personnel in the continuous monitoring video stream in time sequence to construct a real-time action feature sequence; S2. Input the real-time action feature sequence into a preset security inspection action state machine, calculate the temporal first derivative of the frame-by-frame action features between the current frame and the previous frame to characterize the action change rate, and when the temporal first derivative satisfies the preset state transition condition, activate the matching trigger signal in the state machine and determine the dynamic start frame of the sliding window. S3. Using the dynamic start frame as the boundary, intercept the real-time action feature sequence to form the current test window sequence, and perform subsequence dynamic time warping incremental calculation on the current test window sequence and the pre-stored standard security inspection action template sequence to construct the cumulative distance matrix of the open endpoint. S4. In the latest calculated column of the cumulative distance matrix, search for the minimum cumulative distance value that satisfies the preset similarity threshold; if it exists, determine that the target security check action has been captured in the continuous monitoring video stream, and backtrack according to the matrix position where the minimum cumulative distance value is located, automatically extract the end frame of the target security check action, and determine the dynamic start frame as the start frame of the target security check action, so as to complete the automatic segmentation and action recognition of the continuous video stream.
[0024] It should be noted that this invention constructs a streaming closed loop from the overall perspective, consisting of forward state prediction triggering → local incremental matrix alignment → backward endpoint backtracking locking. Its underlying mathematical operating mechanism is as follows: Step S1: The input continuous monitoring video stream is abstracted temporally into an infinite-dimensional feature stream. In the formula, This represents the multimodal frame-by-frame action feature vector extracted at time t. The pre-stored standard security check action template sequence is represented as: In the formula, M is the fixed frame length of the standard action.
[0025] In step S2, the system does not directly run a complex regularization algorithm on the global feature stream X, but instead uses a lightweight first-order temporal derivative as a "filter" for the initial coarse segmentation. At time t, the rate of change of the action... The discrete calculation formula is: when When the rate of change exceeds a preset threshold for k consecutive frames, the state machine transitions and activates a trigger signal. At this point, the system truncates the dynamic start frame from the continuous stream and records its corresponding timestamp as... This step transforms an unbounded time series into a bounded sliding window starting point.
[0026] Step S3, with As a hard boundary, with new video frames ( With the continuous inflow of data, the test window sequence dynamically expands to: The system will The standard template Y is fed into the Subsequence Dynamic Time Warping (DTW) engine. To allow the starting point of the standard action to adaptively align with any frame within the test window (solving the multi-frame time error caused by coarse state machine segmentation), the first row of the cumulative distance matrix D is given special boundary condition constraints: For other positions in the matrix, a forward one-dimensional rolling update mechanism is used, updating each new frame. Only calculate and update the latest column (column t) of the cumulative distance matrix: In the formula, It uses Euclidean distance as the measure. This mechanism reduces the traditional space complexity by... Completely locked in The linear constant level ensures real-time performance at the security checkpoint from the underlying mathematical architecture.
[0027] Step S4, when the bottom element of the latest calculated column... satisfy: Judgment in The target security check has been completed. At this point, the system uses this coordinate... Using the anchor point, perform backtracking along the negative gradient direction of matrix D to find the optimal path in reverse:
[0028] The endpoint of this backtracking path It automatically corrects and locks the precise legal starting point of the action in the continuous flow. and The two-way clamping mechanism enables automatic segmentation of the action.
[0029] To visually demonstrate the evolution of the video stream's state during its flow across various modules, the data lifecycle transformation designed in this invention is shown in the following table: Table 1 Data Lifecycle Transformation
[0030] The frame-by-frame action features in S1 are multimodal fusion features, and their extraction and construction process includes: Extract the spatiotemporal trajectory of key points of the target person's two-dimensional or three-dimensional human skeleton to form the first feature vector; Extract the spatiotemporal motion trajectory and speed changes of the security inspection equipment held by the target personnel to form a second feature vector; The first feature vector and the second feature vector are adaptively weighted and concatenated to generate the frame-by-frame action features.
[0031] When generating the frame-by-frame motion features, the weighting coefficients of the first feature vector and the second feature vector are dynamically adjusted according to the real-time flow of people in the security checkpoint; when the flow of people is greater than a preset density threshold, the weight of the second feature vector is increased to suppress noise interference caused by human skeletal points occlusion.
[0032] It should be noted that this invention constructs highly discriminative multimodal frame-by-frame motion features by jointly representing the "posture of the person being inspected" (first feature vector) and the "trajectory of the handheld security inspection device" (second feature vector). However, security checkpoints are often accompanied by partial or complete occlusion of skeletal key points due to passenger congestion. Therefore, this invention introduces a dynamic weighted control matrix based on real-time passenger flow density. Its underlying mathematical model is as follows: Standardized representation of feature vectors: At time t, the spatiotemporal trajectory feature vector of the target person's skeletal key points (such as shoulder, elbow, and wrist joints) extracted from the surveillance video frame is denoted as the first feature vector. The extracted feature vector of the three-dimensional spatiotemporal motion trajectory and instantaneous velocity change of the center point of the target person's handheld security inspection device (such as a handheld metal detector) is denoted as the second feature vector. .
[0033] Adaptive weighted stitching mapping: fused frame-by-frame action feature vector It is not a simple fixed splicing, but rather a linear or non-linear adjustment through dynamic weighting coefficients, as defined in the following formula: In the formula, The concatenation operator for vectors; and The first and second eigenvectors were respectively processed. Regularization functions are used to eliminate the absolute value deviation of distance measurements caused by different physical dimensions (pixel displacement and velocity); and These are respectively based on real-time pedestrian flow density Dynamically changing adaptive weighting coefficients, satisfying boundary normalization conditions: Design of a Weight Adaptive Function Based on a Sigmoid Variant: To achieve smooth state transitions and avoid drastic oscillations in the feature space caused by minor fluctuations in pedestrian traffic, this invention designs a dynamic weight allocation function based on an inverse Sigmoid variant. Let the real-time pedestrian density (e.g., the number of people per unit area or pixel percentage) within the current security checkpoint, as measured by cameras or infrared sensors, be... The preset density threshold is Then the weight function of the first feature vector (skeleton pose) Defined as: Correspondingly, the weighting function of the second feature vector (security inspection equipment) Defined as: in, This is the sensitivity steepness factor for adaptive adjustment. When the flow of people is extremely low ( )hour, The system primarily relies on high-fidelity human skeletal postures for motion regularization; however, when there is a surge in pedestrian traffic and severe congestion... When the human skeleton is in a state of high noise, the noise level increases dramatically because the skeletal points of the body are easily blocked by luggage or passengers in front and behind. 0 and The system automatically masks or weakens easily distorted skeletal features, and significantly increases and relies on the spatiotemporal motion trajectory features of security inspection equipment with high anti-interference capabilities (because the equipment is usually waving above the human body and is rarely completely obscured).
[0034] To visually demonstrate the system's adaptive adjustment behavior under varying crowd pressures, the present invention is designed as follows: Table 2 Feature-weighted configuration and evolution of technical effects
[0035] From the perspective of geometric projection, it intuitively demonstrates how this invention works at different pedestrian densities ( Under this condition, the physical principle of dynamically compressing or amplifying the feature space by controlling the weighting coefficients to achieve noise interference resistance is derived from... Figure 2 As shown.
[0036] The security check action state machine in S2 includes an idle state, a ready state, and a matching state; when the first derivative of the timing sequence satisfies the preset state transition condition, the matching trigger signal is activated in the state machine, specifically including: When the state machine is in the idle state, if the first derivative of the timing exceeds the set static threshold and the duration reaches the preset number of frames, the state machine will switch from the idle state to the ready state. In the ready state, if the spatiotemporal motion trajectory of the security inspection device in the second feature vector meets the preset start-up acceleration feature, and the key points of the human skeleton represented by the first feature vector enter the preset matching trigger area, the state machine jumps to the matching state, activates the matching trigger signal, and marks the current frame as the dynamic start frame of the sliding window.
[0037] It should be noted that the security check action state machine designed in this invention includes three core normal states: Idle, Ready, and Matching. The state machine, through dual temporal and spatial constraints, filters out irrelevant random disturbances (such as passenger movement or slight adjustments by the security personnel) and accurately captures the actual starting point of the security check action. Its mathematical control model is as follows: Abrupt transition constraint from IDLE to READY state: When the state machine is in the IDLE state, the system is in a low-power listening mode and continuously calculates the first-order temporal derivative (rate of change of action) at the current time t. To achieve a transition to the READY state, the characteristic change rate must exceed a set static threshold for N consecutive frames. Its formalized decision criterion formula is: In the formula, This is an indicator function; the output is 1 when the condition within the parentheses is true, and 0 otherwise. This condition, by introducing a time duration window N (preset frame number), effectively filters out transient isolated noise points in the video stream caused by light transients or occasional camera shake.
[0038] The transition from the READY state to the MATCHING state is subject to both spatial and dynamic constraints. In the READY state, the system determines that the security personnel have broken from their stationary state and begun to accumulate energy or move. To further trigger the MATCHING state and initiate subsequent high-load incremental DTW calculations, the dual intersection constraint (i.e., AND logic) of the equipment's dynamic characteristics and the human body's spatial position characteristics must be satisfied simultaneously: Dynamic characteristics: Indicated based on the second feature vector The calculated spatiotemporal acceleration of the handheld security inspection device at the current moment. This is a preset threshold for the startup acceleration characteristic.
[0039] Spatial characteristics: This indicates that the first eigenvector is represented by... Key points representing the human skeleton (such as wrist joint coordinates). A pre-defined "matching trigger area" (i.e., the compliant core working space in front of the passenger) in a three-dimensional or two-dimensional image coordinate system.
[0040] Only when the system detects that the security equipment is being waved at an accelerated speed ( ), and the wrist simultaneously cuts into the sensitive area of the operation ( When the current frame timestamp is set to the MATCHING state, the state machine transitions to the MATCHING state, and simultaneously sets the current frame timestamp as the dynamic start frame of the sliding window. To clearly illustrate the response mechanism of the state machine under different external stimuli, the state transition decision and action execution matrix table is as follows: Table 3 State Transition Decision and Action Execution Matrix
[0041] The constraints for calculating the dynamic time warping increment of the subsequence in S3 include: The first row of the cumulative distance matrix is initialized to zero to allow the starting point of the standard security action template sequence to adaptively align with any frame in the current test window sequence; A Sakoe-Chiba constraint band is introduced to limit the search space of the cumulative distance matrix. The width of the constraint band is dynamically adjusted according to the action execution rate reflected by the fluctuation period of the temporal change rate of the real-time action feature sequence within a preset time window.
[0042] The cumulative distance matrix for constructing open endpoints in S3 adopts a forward one-dimensional rolling update mechanism: Each time a new action feature is input, the system calculates the single-column distance data corresponding to the latest frame only for the cumulative distance matrix, and overwrites or releases the memory of the invalid historical columns, so that the space complexity remains O(M), where M is the length of the standard security check action template sequence.
[0043] It should be noted that in continuous security check action recognition, the speed and rhythm (temporal scaling) of actions vary significantly due to different security personnel's operating habits, fatigue levels, or unexpected situations on-site. This invention achieves high fault tolerance and low computational cost real-time alignment by introducing a subsequence topology with the first row set to zero, an adaptive variable-width constraint band, and one-dimensional rolling streaming computation in step S3. Its mathematical model is defined as follows: Initialize the boundary conditions for the open endpoints of the subsequence, assuming the standard security check action template sequence is... The length is M; the current window sequence to be tested is The length is N.
[0044] Traditional global DTW requires strict endpoint alignment (i.e., the bottom left corner of the matrix aligns with the top right corner). To allow the starting point of template Y to adaptively match any position within the sliding window of a continuous flow, this invention subsequences the cumulative distance matrix D, and its boundary condition initialization formula is defined as: This geometric boundary design allows the accumulated distance to slide laterally without penalty in the first row (template start point), perfectly eliminating the start frame time error caused by the coarse segmentation of the preceding state machine.
[0045] The present invention introduces a dynamically adaptive Sakoe-Chiba constraint band to prevent excessive ill-conditioned curvature of temporally regularized paths without physical meaning and to further compress the search space. Unlike traditional fixed bandwidth, the constraint band width R(t) of the present invention is dynamically evolving.
[0046] First, the system extracts the preset time window. The temporal rate of change (i.e., the first derivative) of the real-time action feature sequence within the data. The main fluctuation period can be extracted using the autocorrelation function or the Fast Fourier Transform (FFT). The fluctuation cycle directly reflects the speed at which the security inspector waves the detector or turns around (a short cycle indicates a fast action, and a long cycle indicates a slow action).
[0047] The adaptive mapping formula for dynamic bandwidth R(t) is defined as follows: in, Based on the basic constraint bandwidth threshold, This is the standard motion cycle constant. Therefore, at matrix coordinates (i,j), only mesh elements satisfying the following absolute value constraint will be activated and calculated; the remaining positions are directly assigned... .
[0048] Forward one-dimensional rolling update recursive control, at time t, with the new frame features The system no longer recalculates the historical matrix; instead, it uses a forward incremental recursive approach. When updating the latest column (column t), the mathematical recursive equation for matrix element D(i,t) is: in, It is a Euclidean distance measure for multimodal features.
[0049] To visually demonstrate the significant technical advantages of this invention in streaming data processing and on-site industrial control computer memory management, the table below provides a quantitative comparison of memory and computing power: Table 4. Quantitative Comparison of Memory and Computing Power
[0050] The preset similarity threshold in S4 employs a dynamic adaptive adjustment strategy, specifically including: The thickness characteristics of the clothing of the passenger being inspected are extracted and analyzed by infrared imaging and edge contour recognition. The preset similarity threshold is adjusted based on the thickness level of the clothing. When the thickness level of the clothing is greater than the preset thickness threshold, the preset similarity threshold is increased by a preset ratio based on the original basic threshold. Conversely, the preset ratio is decreased or the basic threshold is kept unchanged to compensate for the deformation of the security inspection action caused by the thickness of the passenger's clothing.
[0051] It should be noted that the system can maintain a very high accuracy rate in recognizing actions when dealing with passengers of different seasons and clothing thicknesses, avoiding false alarms caused by distortion of action amplitude due to physical obstruction by clothing. The dynamic adaptive adjustment strategy of similarity threshold is explained below, combining specific feature quantification, mathematical formulas, parameter mapping tables and spatial geometric principles.
[0052] In real-world security screening scenarios, winter passengers wear heavy clothing such as down jackets and coats, while summer passengers wear light clothing such as short sleeves. The thickness of clothing introduces two types of interference: firstly, visually, skeletal key points drift outwards and become distorted; secondly, physical movement is restricted, forcing security personnel to expand the radius of their physical trajectory when performing standard detection actions (such as close-up scanning). This results in an overall higher calculated cumulative distance value (dissimilarity) between the current sequence being tested and the standard template. Without dynamic compensation, traditional fixed thresholds are highly likely to misjudge compliant security screening actions in winter as "non-compliant or undetected."
[0053] This invention quantifies clothing thickness and dynamically adjusts the similarity threshold using a combination of infrared imaging and edge contour analysis. Its underlying mathematical model is as follows: Quantitative extraction of clothing thickness features The system uses infrared thermal imaging sensors to acquire the spatiotemporal temperature field distribution of passengers and employs computer vision algorithms to extract the external edge contours of passengers within a standard surveillance video stream. Because heavy clothing such as down jackets has excellent thermal insulation properties, its infrared radiation outer contour corresponds to the core skeletal structure of the human body (defined by the first feature vector). There is a clear spatial gap between them.
[0054] Let the pixel boundary of the extracted passenger's outer edge contour be defined in the image coordinate system. The corresponding human core skeletal trunk centerline is The quantitative value of the clothing thickness characteristic of the currently inspected passenger. Defined as the average Euclidean distance between the two: Adaptive piecewise control function for similarity threshold: The system will quantify the clothing thickness characteristics With the preset thickness threshold Perform dynamic comparison. Assume the initial standard action matching threshold set by the system is... Finally, the dynamic similarity threshold is applied in step S4. The adaptive piecewise function relation is defined as follows: In the formula, The preset thickness threshold (e.g., the thickness of a regular autumn jacket). This is a preset threshold for thin clothing (such as the thickness of a summer T-shirt). ; A preset scaling factor (relaxation factor) is used to proportionally amplify the threshold tolerance when clothing becomes thicker, the range of motion changes, or the cumulative distance matrix value increases overall, ensuring that no motion is missed. The preset scaling factor (tightening factor) is adjusted to proportionally tighten the threshold range when clothing is extremely thin in summer and movement trajectories are extremely precise, thereby improving the system's sensitivity in identifying disguised or irregular movements.
[0055] To visually demonstrate the system's adaptive control strategy under different seasons or clothing scenarios, the dynamic threshold decision and compensation mapping table is as follows: Table 5 Dynamic Threshold Decision and Compensation Mapping
[0056] S4 is followed by S5: Extract the optimal twist path slope of the automatically segmented target security inspection action segment in the cumulative distance matrix; The slope of the twisted path is compared with the standard compliance rate. If the slope deviates from the preset compliance range, a compliance evaluation warning signal for abnormal action rate is output.
[0057] It should be noted that the formal representation of the optimal twisted path, denoted as W, is derived from the optimal twisted path stripped away by backtracking in step S4. This path consists of K consecutive matrix grid coordinate points: In the formula, each path point Each represents a two-dimensional coordinate. This represents the frame index of the standard security check action template sequence Y. Represents the actual action segments segmented from a continuous flow. Frame index.
[0058] The calculation of the global / local slope of the optimal twist path is crucial, as the tangent slope of the optimal twist path in geometric space directly reflects the relative execution rate of the actual action compared to the standard action. This invention calculates the overall slope of the path using least squares linear regression or the two endpoint difference method. : More precisely, in order to capture rate anomalies that occur locally during the action (e.g., slow at the beginning and fast at the end, or pauses in the middle), this invention uses a sliding derivative window to calculate the local real-time slope on path W. : In the formula, This is the preset number of frames for the timing micro-element window.
[0059] A security check action recognition system based on dynamic time warping includes: Feature extraction module: used to acquire continuous surveillance video stream of the security check area, and extract frame-by-frame action features of the target personnel in the continuous surveillance video stream in time sequence to construct a real-time action feature sequence; State control module: used to input the real-time action feature sequence into a preset security check action state machine, calculate the temporal first derivative of the frame-by-frame action features between the current frame and the previous frame to characterize the action change rate, and when the temporal first derivative meets the preset state transition conditions, activate the matching trigger signal in the state machine and determine the dynamic start frame of the sliding window. Incremental alignment module: used to intercept the real-time action feature sequence with the dynamic start frame as the boundary to form the current test window sequence, and to perform subsequence dynamic time warping incremental calculation on the current test window sequence and the pre-stored standard security inspection action template sequence to construct the cumulative distance matrix of the open endpoint; The segmentation and recognition module is used to search for the minimum cumulative distance value that satisfies a preset similarity threshold in the latest calculated column of the cumulative distance matrix. If it exists, it determines that a target security check action has been captured in the continuous monitoring video stream. Based on the matrix position of the minimum cumulative distance value, it backtracks backward to automatically extract the end frame of the target security check action and determines the dynamic start frame as the start frame of the target security check action, so as to complete the automatic segmentation and action recognition of the continuous video stream.
[0060] It should be noted that the identification system involved in this invention is not a superposition of four independent, static functional blocks, but rather a real-time pipelined system with high concurrency capabilities, connected in series via a shared memory buffer (RingBuffer) and an event-driven state machine bus. The macroscopic architecture and data interaction topology of the system are defined as follows: The feature extraction module runs in the underlying high-frequency image acquisition and processing thread. It continuously outputs fixed-length feature vectors and writes them as time-series producers into the dynamically maintained circular buffer.
[0061] The state control module performs high-frequency polling of the latest features in the circular buffer with extremely low computational overhead (involving only first-order differential and logical jump judgment). Once the state jump condition is met, it broadcasts a MATCHING_START signal to the entire system via the asynchronous event bus and passes the physical memory pointer address of the current frame to the incremental alignment module.
[0062] The incremental alignment module is in a blocked and suspended (Sleep) state before receiving a trigger signal. Once awakened by the state control module, it immediately locks the memory pointer address as the dynamic start frame boundary of the sliding window and calls on the hardware's GPU / NPU computing resources to carry out high-concurrency incremental matrix updates for the subsequently incoming feature columns.
[0063] The segmentation and recognition module shares the same cumulative distance matrix storage space as the incremental alignment module. After each frame update, it performs a global minimum search and backtracking evaluation on the latest column of data in a pipelined parallel mode. Once the judgment action is completed, it sends a CYCLE_RESET signal to the state control module, allowing the system to smoothly return to a low-power listening state.
[0064] To accurately describe the control logic and data interface conversion relationships during the evolution of a continuous video stream within the system, this invention designs the following temporal topology relationships for the interaction of internal system modules. The table below details the underlying component partitioning and hardware resource configuration examples of each functional module on an industrial-grade edge computing chip (such as an embedded GPU or a high-performance industrial control computer): Table 6. Underlying Component Division and Hardware Resource Configuration
[0065] Example: This example uses a standard body check at an airport security checkpoint (using a handheld metal detector to perform an S-shaped close-up scan) as an application scenario to illustrate the specific implementation steps of the present invention in detail.
[0066] Step 1: High-definition surveillance cameras are deployed above and to the sides of the security checkpoint. The system continuously captures a live video stream of the security personnel's work area. When a passenger enters the security checkpoint, the feature extraction module begins online streaming processing of the video frames. Step 2: The system uses an embedded pose estimation algorithm to extract key points of the security inspector's 2D or 3D human skeleton in real time. It focuses on tracking the spatiotemporal motion trajectories of the security inspector's shoulder, elbow, and wrist joints, and concatenates these 3D coordinates sequentially to construct a first feature vector representing the overall posture. Meanwhile, the system uses a target tracking algorithm to perform high-frequency trajectory tracking on the portable metal detector held by the security inspector, capturing the spatiotemporal motion trajectory, positional offset, and instantaneous speed changes of the security device, thereby constructing a second feature vector; Step 3: When generating the final frame-by-frame motion features, the system does not use a fixed stitching ratio, but instead uses the feature extraction module to read the passenger flow density in the security checkpoint in real time. When the channel is unobstructed, the system assigns higher weights to the key points of the human skeleton (first feature vector) to ensure the precision of action detail recognition. During peak hours and when severe queuing congestion occurs in the passageway, the torso or lower limb skeletal points of security inspectors are easily lost due to physical obstruction by passengers in front and behind, generating significant noise. At this time, the system detects that the flow of people exceeds the preset density threshold and automatically triggers the anti-obstruction control mechanism, reducing the weight of skeletal posture and significantly increasing the motion trajectory and speed characteristics of the handheld security inspection device (second feature vector), which is mainly dependent on the handheld detector, because the handheld detector is usually waved on the outside and above the passenger's body and is extremely difficult to be completely obstructed. Step 4: Through this adaptive weighted stitching, the system generates highly discriminative multimodal frame-by-frame motion features for the current video frame and continuously pushes them into the circular feature buffer to form a real-time motion feature sequence.
[0067] Step 5: The generated real-time action feature sequence is synchronously input into the preset security check action state machine in the state control module. To avoid the system running a high-load algorithm for an extended period when no one is conducting security checks, the state machine is initialized to an idle state; The state machine begins to calculate the temporal first derivative of the feature vector between the current input frame and the previous frame at high frequency, and uses this derivative to qualitatively characterize the rate of change of the security inspector's actions. Step 6: When the security inspector is in a stationary state, waiting for passengers or standing with hands hanging down, the first derivative is extremely small, the state machine is locked in the idle state, and the subsequent matrix calculations remain blocked and suspended, maintaining the system's low-power operation. Step 7: When a new passenger steps onto the security checkpoint and the security officer raises their arm to break the static state, the first derivative of the characteristic instantaneously exceeds the set static threshold. If this threshold-breaking behavior continues for a preset number of frames (e.g., five consecutive frames), the state machine eliminates instantaneous noise interference and formally transitions from the idle state to the ready state. Step 8: In the ready state, the system initiates dual spatial and dynamic cross-verification. The state machine closely monitors the second feature vector. If the spatiotemporal motion velocity and acceleration curve of the handheld security inspection device reaches the activation characteristics of a charged-up wave, and the coordinates of the security inspector's wrist joint represented by the first feature vector officially enter the preset core matching trigger area in front of the passenger, the state machine instantly jumps to the matching state; Step 9: The instant the state machine transitions to the matching state, the system immediately activates the matching trigger signal and firmly marks the timestamp of the current video frame that triggered this state transition as the dynamic start frame of the sliding window. Step 10: Once the trigger signal to be matched is activated, the incremental alignment module is awakened. The system uses the dynamically marked start frame as a hard boundary to truncate and intercept the continuously flowing real-time action feature sequence, forming the current test window sequence in memory; Step 11: The incremental alignment module then performs subsequence dynamic time warping incremental calculations between the current test window sequence and the system's pre-stored standard security check action template sequence. To eliminate the multi-frame time error that may be caused by coarse state machine segmentation, the system applies a special mathematical constraint when constructing the two-dimensional cumulative distance matrix: the first row of the cumulative distance matrix is initialized to zero. This design allows the beginning of the standard template action to be adaptively aligned to any frame in the current test window sequence without penalty, exhibiting extremely high streaming fault tolerance. Simultaneously, the system introduces a dynamically adaptive constraint band mechanism. The system analyzes the fluctuation period of the feature sequence within a time window in real time; this period directly reflects the speed of the security inspector's current actions. Based on this fluctuation period, the system dynamically shrinks or expands the search space of the constraint band. Step 12: To ensure real-time performance, the entire cumulative distance matrix is constructed using a forward one-dimensional rolling update mechanism. Each time a new motion feature is input, the system does not recalculate the historical global matrix; instead, it only updates the matrix by calculating the distance data for the latest single column. After calculation, the memory of invalid historical columns is immediately overwritten or released, ensuring that the hardware memory space allocated by the system is always locked at a constant level equal to the length of the standard motion template. Even if the continuous video stream lasts for several hours, the system memory will never expand. Step 13: While the incremental alignment module dynamically refreshes the latest column of data in the matrix, the segmentation and recognition module continuously searches for the current minimum cumulative distance value in the bottom element of the latest calculated column. Step Fourteen: When the security officer has not yet completed a full set of scanning actions, the cumulative distance value at the bottom of the latest column will be large. Only when the security officer smoothly completes the last frame of the standard passenger scanning action will the minimum cumulative distance value instantly drop below the preset similarity threshold. Once this threshold condition is met, the segmentation and recognition module determines that a complete target security check action has been successfully captured in the continuous monitoring video stream; Step 15: At this point, the system uses the latest column matrix position that meets the threshold as the anchor point and starts the reverse path backtracking algorithm. The algorithm searches backward along the negative gradient direction of the cumulative distance matrix, automatically extracting and locating the end frame of the action in the continuous stream; Simultaneously, when the backtracking path reaches the top of the matrix, it automatically corrects and locks its corresponding most precise true starting point. The system directly determines the dynamic starting frame initially locked by the state machine as the legal starting frame for the security check action. Thus, without any manual pre-editing of the video, the system automatically "extracts" and segments the security check action from the continuous video stream, completing adaptive automatic segmentation and action compliance recognition of the continuous stream. Step 16: The system determines that the current security check action cycle has ended, sends a reset signal to the state control module, the state machine returns to the idle state, the current matrix is cleared, and it waits for the next passenger to arrive; Step 17: At the moment when the target security check action is captured and video segmentation is completed in Step 4, the segmentation and recognition module seamlessly cascades the data of the optimal backtracking path to the subsequent evaluation process; Step 18: The system extracts the optimal twist path left by the automatically segmented target security inspection action segment in the cumulative distance matrix, and uses a linear regression method to calculate the globally optimal twist path slope. This slope geometrically and intuitively reflects the relative execution rate between the security inspector's actual actions and the standard template actions; Step 19: The system compares the slope of the optimal twist path with the industry standard compliance rate range stored in the background; If a security inspector hastily shakes the detector around at an extremely fast speed to cope with queuing pressure, the number of video frames will be severely insufficient, and the twisted path will be severely deflected in the matrix space. The calculated optimal twisted path slope will deviate significantly and exceed the preset compliance range limit. Step 20: The system determines that the action is "perfunctory and suspected missed inspection due to excessive speed". The status control module then pops up a red high-risk alarm on the industrial control computer screen at the security checkpoint and outputs a compliance evaluation warning signal for abnormal action speed to the backend, prompting the on-site commander to manually intercept and re-inspect the passenger.
[0068] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A security inspection action recognition method based on dynamic time warping, characterized in that, include: S1. Obtain the continuous monitoring video stream of the security check area, and extract the frame-by-frame action features of the target personnel in the continuous monitoring video stream in time sequence to construct a real-time action feature sequence; S2. Input the real-time action feature sequence into a preset security inspection action state machine, calculate the temporal first derivative of the frame-by-frame action features between the current frame and the previous frame to characterize the action change rate, and when the temporal first derivative satisfies the preset state transition condition, activate the matching trigger signal in the state machine and determine the dynamic start frame of the sliding window. S3. Using the dynamic start frame as the boundary, intercept the real-time action feature sequence to form the current test window sequence, and perform subsequence dynamic time warping incremental calculation on the current test window sequence and the pre-stored standard security inspection action template sequence to construct the cumulative distance matrix of the open endpoint. S4. In the latest calculated column of the cumulative distance matrix, search for the minimum cumulative distance value that satisfies the preset similarity threshold; If it exists, it is determined that the target security check action has been captured in the continuous monitoring video stream. Based on the matrix position where the minimum cumulative distance value is located, the end frame of the target security check action is automatically extracted and the dynamic start frame is determined as the start frame of the target security check action, so as to complete the automatic segmentation and action recognition of the continuous video stream.
2. The security inspection action recognition method based on dynamic time warping according to claim 1, characterized in that, The frame-by-frame action features in S1 are multimodal fusion features, and their extraction and construction process includes: Extract the spatiotemporal trajectory of key points of the target person's two-dimensional or three-dimensional human skeleton to form the first feature vector; Extract the spatiotemporal motion trajectory and speed changes of the security inspection equipment held by the target personnel to form a second feature vector; The first feature vector and the second feature vector are adaptively weighted and concatenated to generate the frame-by-frame action features.
3. The security inspection action recognition method based on dynamic time warping according to claim 2, characterized in that, When generating the frame-by-frame motion features, the weighting coefficients of the first feature vector and the second feature vector are dynamically adjusted according to the real-time flow of people in the security checkpoint; when the flow of people is greater than a preset density threshold, the weight of the second feature vector is increased to suppress noise interference caused by human skeletal points occlusion.
4. The security inspection action recognition method based on dynamic time warping according to claim 2, characterized in that, The security check action state machine in S2 includes an idle state, a ready state, and a matching state; when the first derivative of the timing sequence satisfies the preset state transition condition, the matching trigger signal is activated in the state machine, specifically including: When the state machine is in the idle state, if the first derivative of the timing exceeds the set static threshold and the duration reaches the preset number of frames, the state machine will switch from the idle state to the ready state. In the ready state, if the spatiotemporal motion trajectory of the security inspection device in the second feature vector meets the preset start-up acceleration feature, and the key points of the human skeleton represented by the first feature vector enter the preset matching trigger area, the state machine jumps to the matching state, activates the matching trigger signal, and marks the current frame as the dynamic start frame of the sliding window.
5. The security inspection action recognition method based on dynamic time warping according to claim 1, characterized in that, The constraints for calculating the dynamic time warping increment of the subsequence in S3 include: The first row of the cumulative distance matrix is initialized to zero to allow the starting point of the standard security action template sequence to adaptively align with any frame in the current test window sequence; A Sakoe-Chiba constraint band is introduced to limit the search space of the cumulative distance matrix. The width of the constraint band is dynamically adjusted according to the action execution rate reflected by the fluctuation period of the temporal change rate of the real-time action feature sequence within a preset time window.
6. The security inspection action recognition method based on dynamic time warping according to claim 1, characterized in that, The cumulative distance matrix for constructing open endpoints in S3 adopts a forward one-dimensional rolling update mechanism: Each time a new action feature is input, the system calculates the single-column distance data corresponding to the latest frame only for the cumulative distance matrix, and overwrites or releases the memory of the invalid historical columns, so that the space complexity remains O(M), where M is the length of the standard security check action template sequence.
7. The security inspection action recognition method based on dynamic time warping according to claim 1, characterized in that, The preset similarity threshold in S4 employs a dynamic adaptive adjustment strategy, specifically including: The thickness characteristics of the clothing of the passenger being inspected are extracted and analyzed by infrared imaging and edge contour recognition. The preset similarity threshold is adjusted based on the thickness level of the clothing. When the thickness level of the clothing is greater than the preset thickness threshold, the preset similarity threshold is increased by a preset ratio based on the original basic threshold. Conversely, the preset ratio is decreased or the basic threshold is kept unchanged to compensate for the deformation of the security inspection action caused by the thickness of the passenger's clothing.
8. The security inspection action recognition method based on dynamic time warping according to claim 1, characterized in that, S4 is followed by S5: Extract the optimal twist path slope of the automatically segmented target security inspection action segment in the cumulative distance matrix; The slope of the twisted path is compared with the standard compliance rate. If the slope deviates from the preset compliance range, a compliance evaluation warning signal for abnormal action rate is output.
9. A security inspection action recognition system based on dynamic time warping, characterized in that, include: Feature extraction module: used to acquire continuous surveillance video stream of the security check area, and extract frame-by-frame action features of the target personnel in the continuous surveillance video stream in time sequence to construct a real-time action feature sequence; State control module: used to input the real-time action feature sequence into a preset security check action state machine, calculate the temporal first derivative of the frame-by-frame action features between the current frame and the previous frame to characterize the action change rate, and when the temporal first derivative meets the preset state transition conditions, activate the matching trigger signal in the state machine and determine the dynamic start frame of the sliding window. Incremental alignment module: used to intercept the real-time action feature sequence with the dynamic start frame as the boundary to form the current test window sequence, and to perform subsequence dynamic time warping incremental calculation on the current test window sequence and the pre-stored standard security inspection action template sequence to construct the cumulative distance matrix of the open endpoint; Segmentation and recognition module: used to search for the smallest cumulative distance value that satisfies a preset similarity threshold in the latest calculated column of the cumulative distance matrix; If it exists, it is determined that the target security check action has been captured in the continuous monitoring video stream. Based on the matrix position where the minimum cumulative distance value is located, the end frame of the target security check action is automatically extracted and the dynamic start frame is determined as the start frame of the target security check action, so as to complete the automatic segmentation and action recognition of the continuous video stream.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the security check action recognition method based on dynamic time warping as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Image processing method and system for rehabilitation training action analysis
CN120563743A
Ball action recognition system and method based on time sequence posture coding
CN122391966A