Intelligent lock monitoring method and system based on machine vision

By capturing and analyzing 3D video frame sequences in a smart lock system, extracting the user's 3D key point coordinate structure, and performing posture change tracking and motion feature recognition, the security risks of existing smart lock systems in complex environments are solved, achieving effective defense against malicious attacks and efficient identity verification.

CN120564272BActive Publication Date: 2025-10-21GUANGDONG SAKURA INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511054175.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-10-21
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

Existing smart lock systems have significant security vulnerabilities when facing advanced attack methods such as malicious probing and video playback attacks. Furthermore, they struggle to accurately extract users' 3D motion features in complex environments, leading to low recognition accuracy, false positives, and false negatives, which negatively impact user experience and system security.

Method used

By capturing a continuous sequence of video frames in the smart lock installation area, extracting the coordinate structure of three-dimensional key points, performing posture change tracking modeling, identifying target interaction behavior events, and detecting multiple repeated attempts to trigger the door lock opening mechanism based on action feature subsequences, determining whether the user access request meets the security policy rules, and generating door lock control commands.

Benefits of technology

It effectively defends against malicious probing and video playback attacks, improves the security and reliability of the smart lock system, accurately identifies abnormal attempts, and enhances user experience and security protection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564272B_ABST
    Figure CN120564272B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of intelligent lock monitoring method and system based on machine vision, including the following steps, the method is by capturing the continuous video frame of intelligent lock area and extracting three-dimensional key point coordinates, the posture of user is tracked modeling, obtains action trajectory curve, identifies target interactive behavior event of approaching operation interface. Subsequently, the event is key frame extraction, obtains the action feature subsequence of candidate identity authentication stage, and whether there is multiple repeated attempt to trigger door lock behavior mode is detected. Finally, whether the access request is in accordance with the safety policy based on the behavior mode is judged, generates door lock control instruction and sends to controller module and executes corresponding operation. The technical problem that most current intelligent lock systems exist greater security risks when facing malicious probe, video playback attack and other advanced attack means is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart locks, and in particular to a smart lock monitoring method and system based on machine vision. Background Art

[0002] With the rapid development of smart door lock technology, traditional authentication methods based on passwords, fingerprints, or IC cards are increasingly unable to meet the growing demand for security and convenience. In recent years, machine vision technology has demonstrated tremendous potential in behavioral recognition and identity authentication, offering new insights into enhancing the security of smart lock systems. By analyzing user behavioral characteristics when operating smart locks, such as posture changes and movement trajectories, dynamic assessment of access requests can be achieved, thereby enhancing prevention of unauthorized intrusions.

[0003] However, most current smart lock systems still rely primarily on static authentication methods and lack the ability to monitor and analyze user interactions in real time. This poses significant security risks when facing advanced attacks such as malicious probing and video playback. Furthermore, existing technologies often neglect the effective use of three-dimensional spatial information when extracting behavioral features, resulting in limited motion recognition accuracy and prone to misjudgments and missed detections, impacting user experience and system security.

[0004] More importantly, accurately extracting key user motion features and effectively identifying unusual attempts in complex environments (such as those with fluctuating lighting, obstructions, and multiple people) remains a challenge. Existing methods for behavior modeling and pattern recognition still suffer from high response latency and low recognition rates, making it difficult to achieve efficient and stable security protection in practical applications. Therefore, a smart lock monitoring method that combines three-dimensional motion features with behavioral pattern analysis is urgently needed to address increasingly complex usage scenarios and security threats. Summary of the Invention

[0005] The main purpose of the present invention is to provide a smart lock monitoring method based on machine vision, which solves the technical problem that most current smart lock systems have major security risks when facing advanced attack methods such as malicious probing and video playback attacks.

[0006] To achieve the above objectives, the present invention provides a smart lock monitoring method based on machine vision, comprising the following steps:

[0007] Capturing a continuous video frame sequence within the smart lock installation area, and extracting a three-dimensional key point coordinate structure from the continuous video frame sequence;

[0008] Based on the three-dimensional key point coordinate structure, the user's posture change tracking model is performed to obtain a user motion trajectory curve, and the target interactive behavior event approaching the operation interface in the smart lock is identified in combination with the motion trajectory curve;

[0009] Extract key frames from the target interactive behavior event to obtain a subsequence of action features in the candidate identity verification stage;

[0010] Detecting whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism based on the action feature subsequence;

[0011] Based on the behavior pattern, it is determined whether the user's current access request meets the preset security policy rules, a door lock control command signal is obtained, and the door lock control command signal is sent to the door lock controller module in the smart lock to complete the corresponding operation execution.

[0012] Furthermore, the capturing of a continuous video frame sequence within the smart lock installation area and the extraction of a three-dimensional key point coordinate structure from the continuous video frame sequence include:

[0013] Perform multi-angle image sampling on the ambient light field information in the area where the smart lock is installed to obtain a continuous video frame sequence including the access control channel area;

[0014] Performing a background difference operation based on the continuous video frame sequence to obtain a dynamic foreground mask layer, and performing a connected domain labeling process on the dynamic foreground mask layer to obtain a candidate human body contour area;

[0015] Boundary curve fitting and extreme point analysis are performed on each region in the candidate human body contour area to obtain local significant feature points including the positions of the head, shoulders, and hands, and a three-dimensional key point coordinate structure for characterizing the user's spatial posture is constructed based on the local significant feature points.

[0016] Furthermore, the tracking modeling of the user's posture changes based on the three-dimensional key point coordinate structure to obtain the user's motion trajectory curve includes the following steps:

[0017] Analyzing the temporal features in the three-dimensional key point coordinate structure to obtain a key point motion trajectory sequence, and extracting posture change features in the key point motion trajectory sequence to obtain a multi-dimensional spatiotemporal feature matrix including joint angle changes, limb movement directions, and speed acceleration;

[0018] Performing dynamic spatiotemporal segmentation on the user's limb movements in a continuous video frame sequence based on the multidimensional spatiotemporal feature matrix to obtain a sequence of action segment boundary moments, and performing kinematic feature clustering analysis on the sequence of action segment boundary moments to obtain posture change features;

[0019] Performing action combination pattern matching on the posture change features through hierarchical spatiotemporal correlation analysis technology to obtain a user interaction intention feature vector, and performing temporal probability reasoning on the user interaction intention feature vector to obtain a temporal state transition diagram;

[0020] A dynamic programming search is performed on the timing state transition diagram to obtain an optimal state transition path sequence, and a user motion trajectory curve is constructed based on the optimal state transition path sequence, wherein the user motion trajectory curve includes interactive behavior characteristic parameters such as hand contact position, contact duration, and contact force change.

[0021] Furthermore, the dynamic spatiotemporal segmentation of the user's limb movements in the continuous video frame sequence based on the multi-dimensional spatiotemporal feature matrix to obtain the action segment boundary moment sequence includes the following steps:

[0022] Performing temporal gradient analysis on the multidimensional spatiotemporal feature matrix to obtain a feature change rate curve, and performing adaptive threshold segmentation based on the feature change rate curve to obtain initial action boundary candidate points, wherein the initial action boundary candidate points include time indexes at which features change significantly;

[0023] Constructing a spatiotemporal feature map based on the initial action boundary candidate points, and applying non-local mean filtering to the spatiotemporal feature map to obtain a smoothed feature map, wherein the smoothed feature map retains key action transition information while suppressing small noise disturbances;

[0024] A hierarchical clustering operation is performed on the smoothed feature graph to obtain an action fragment tree structure, and multi-scale boundary optimization is performed based on the action fragment tree structure to obtain an action fragment boundary moment sequence, wherein the action fragment boundary moment sequence reflects the key action sub-nodes in the process of user interaction with the smart lock.

[0025] Furthermore, extracting key frames from the target interactive behavior event to obtain an action feature subsequence in the candidate identity verification stage includes:

[0026] Performing spatiotemporal analysis on the video sequence of the target interactive behavior event to obtain an action change amplitude curve, and setting an adaptive threshold based on the action change amplitude curve to obtain a key moment index set, wherein the key moment index set includes timestamps of the action start, action apex, and action end;

[0027] Performing frame-level sampling on a continuous video frame sequence based on the key moment index set to obtain candidate key frames, and performing multi-scale image pyramid decomposition on the candidate key frames to obtain a multi-resolution feature map;

[0028] Performing spectral clustering analysis on the multi-resolution feature map to obtain a feature similarity matrix, constructing an inter-frame association network based on the feature similarity matrix, and determining a key frame representativeness score for each of the continuous video frame sequences based on the inter-frame association network, wherein the key frame representativeness score reflects the contribution of each frame image to the overall behavior representation;

[0029] The candidate key frames are screened and sorted based on the key frame representativeness scores to obtain a key frame sequence, and a candidate identity verification stage action feature subsequence is extracted from the key frame sequence.

[0030] Furthermore, the spatiotemporal analysis of the video sequence of the target interactive behavior event to obtain the action change amplitude curve includes:

[0031] Performing optical flow calculation on the video sequence of the target interactive behavior event to obtain a motion vector map, and performing motion gradient decomposition on the motion vector map to obtain a three-dimensional motion component tensor including horizontal displacement, vertical displacement, and angular rotation;

[0032] Performing multi-scale time-frequency analysis on the three-dimensional motion component tensor through wavelet transform to obtain a motion feature spectrum, and performing energy density calculation on the motion feature spectrum to obtain a motion energy distribution sequence, wherein the motion energy distribution sequence describes the frequency characteristics of motion changes;

[0033] Performing time domain segmentation processing based on the motion energy distribution sequence to obtain action segment boundary points, and performing smooth interpolation operation on the action segment boundary points to obtain a continuous motion state curve;

[0034] The continuous motion state curve is amplitude normalized to obtain a standardized motion feature sequence, and a motion change amplitude curve is constructed based on the standardized motion feature sequence, wherein the motion change amplitude curve includes time series feature parameters of motion intensity, motion duration, and motion continuity.

[0035] Furthermore, the detecting, based on the action feature subsequence, whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism includes:

[0036] Performing spatiotemporal pattern decomposition on the motion feature subsequence to obtain periodic motion segments, and performing phase alignment processing on the periodic motion segments to generate a standardized motion trajectory template;

[0037] Performing sequence similarity matching on the standardized motion trajectory template based on dynamic time warping to obtain a repeated attempt behavior matching matrix, and detecting whether there is a suspicious behavior marker sequence in the repeated attempt behavior matching matrix;

[0038] Modeling the state transition probability of the suspicious behavior mark sequence through a hidden Markov model to obtain a behavior abnormality score, and performing a time series cumulative calculation of the behavior abnormality score based on Bayesian inference to generate a comprehensive risk coefficient;

[0039] Adaptive threshold segmentation is used to divide the comprehensive risk coefficient into risk levels to obtain a behavior pattern determination result, and a logical check is performed on the behavior pattern determination result to generate a door lock triggering behavior analysis report. Based on the door lock triggering behavior analysis report, it is determined whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism; wherein, the door lock triggering behavior analysis report includes the number of repeated attempts, behavior consistency and risk confidence.

[0040] The present invention also provides a smart lock monitoring system based on machine vision, comprising:

[0041] A capture module is used to capture a continuous video frame sequence within the smart lock installation area and extract a three-dimensional key point coordinate structure from the continuous video frame sequence;

[0042] A modeling module is used to track and model the user's posture changes based on the three-dimensional key point coordinate structure to obtain a user motion trajectory curve, and identify target interactive behavior events approaching the operation interface in the smart lock in combination with the motion trajectory curve;

[0043] An extraction module, configured to extract key frames from the target interactive behavior event to obtain an action feature subsequence in the candidate identity verification stage;

[0044] A triggering module, configured to detect whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism based on the action feature subsequence;

[0045] The judgment module is used to judge whether the user's current access request meets the preset security policy rules based on the behavior pattern, obtain the door lock control command signal, and send the door lock control command signal to the door lock controller module in the smart lock to complete the corresponding operation execution.

[0046] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.

[0047] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0048] The present invention provides a machine vision-based smart lock monitoring method, comprising the following steps: capturing a continuous video frame sequence within the smart lock installation area and extracting a three-dimensional key point coordinate structure from the continuous video frame sequence; tracking and modeling a user's posture changes based on the three-dimensional key point coordinate structure to obtain a user motion trajectory curve, and identifying target interaction behavior events approaching an operation interface in the smart lock in combination with the motion trajectory curve; extracting key frames from the target interaction behavior events to obtain a motion feature subsequence in a candidate identity verification stage; detecting whether a behavior pattern of multiple repeated attempts to trigger a door lock opening mechanism is present based on the motion feature subsequence; determining whether the user's current access request satisfies a preset security policy rule based on the behavior pattern, obtaining a door lock control command signal, and sending the door lock control command signal to a door lock controller module in the smart lock to execute the corresponding operation. This method solves the technical problem that most current smart lock systems have significant security risks when facing advanced attack methods such as malicious probing and video playback attacks. By detecting whether a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism is present based on the motion feature subsequence, the method can effectively identify attack behaviors such as brute force cracking and gesture imitation, further enhancing the technical effect of the defense capability against illegal intrusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 1 is a schematic diagram of the steps of a smart lock monitoring method based on machine vision in one embodiment of the present invention;

[0050] Figure 2 This is a block diagram of a smart lock monitoring system based on machine vision in one embodiment of the present invention;

[0051] Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.

[0052] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0054] like Figure 1 As shown, Figure 1 A method for monitoring a smart lock based on machine vision in one embodiment of the present invention includes the following steps:

[0055] Step S1, capturing a continuous video frame sequence within the smart lock installation area, and extracting a three-dimensional key point coordinate structure from the continuous video frame sequence.

[0056] Specifically, the system captures a continuous sequence of video frames within the smart lock installation area and extracts a 3D keypoint coordinate structure from this sequence. This step relies on the combined application of machine vision and 3D pose estimation technology. Specifically, the system continuously collects video data from the smart lock installation area using multi-view cameras or RGB-D cameras with depth perception (such as Kinect or Intel RealSense) deployed around the smart lock, generating a continuous sequence of video frames that provides raw visual input for subsequent analysis. Furthermore, the system employs deep learning-based 3D human pose estimation models (such as VIBE and VideoPose3D) to detect and track key body parts (such as wrists, shoulders, and head) in each frame, extracting a 3D keypoint coordinate structure that describes the user's spatial position changes. For example, as a user approaches a smart lock and prepares to authenticate, the system can continuously capture and 3D-model the user's hand movements, obtaining a time-varying sequence of keypoint coordinates in 3D space. This provides a high-precision spatial information foundation for subsequent pose modeling and behavior recognition.

[0057] Step S2: tracking and modeling the user's posture changes based on the three-dimensional key point coordinate structure to obtain a user motion trajectory curve, and identifying target interactive behavior events approaching the operation interface in the smart lock in combination with the motion trajectory curve.

[0058] Specifically, tracking and modeling the user's posture changes based on the three-dimensional key point coordinate structure refers to using the three-dimensional key point coordinate data extracted in the previous step, combined with time series information, to construct a dynamic posture evolution model of the user within the smart lock installation area. By continuously tracking the position changes of each key point in three-dimensional space and mapping it into a continuous motion trajectory curve, the system can accurately depict the user's behavior pattern during operation. For example, when a user approaches the smart lock and prepares to enter a password or perform gesture verification, the movement trajectory of their hands and other body parts will be modeled in real time and quantified in the form of a motion trajectory curve. Subsequently, the system analyzes the user's movement trends and spatial behavior characteristics based on the motion trajectory curve to identify whether there is a target interaction behavior event approaching the operating interface of the smart lock, that is, to determine whether the user is performing an action that may trigger the identity verification process. For example, when the system detects that the key points of the user's hand gradually approach the touch panel of the smart lock, accompanied by a brief pause and click action, it can be determined that a target interaction behavior event has occurred, thereby providing behavioral basis for the subsequent identity verification stage.

[0059] Step S3: extract key frames from the target interactive behavior event to obtain an action feature subsequence of the candidate identity authentication stage.

[0060] Specifically, performing keyframe extraction on the target interactive behavior event means, after identifying the interactive behavior of the user approaching the smart lock operation interface, selecting representative keyframes from the continuous video frame sequence corresponding to the behavior event to construct a subsequence of action features for the candidate identity authentication stage. The specific implementation method is to locate the time interval in which the user performs key operations based on the time alignment information of the action trajectory curve, and select several keyframes within this interval using strategies such as time sampling, motion amplitude change or posture significance, thereby forming a compact subsequence containing core action features. For example, when the user approaches the smart lock and prepares to enter the password, the system accurately captures the period when the user actually touches or simulates the operation interface by analyzing the changes in the three-dimensional key points of his hand, and extracts several keyframes such as the gesture start, gesture expansion and gesture end from it to form a subsequence of action features for the candidate identity authentication stage. This process not only reduces the data redundancy of subsequent processing, but also retains the core action information related to identity authentication, providing an efficient and accurate data basis for subsequent detection of whether there are behavioral patterns of repeated attempts to trigger the door lock opening mechanism.

[0061] Step S4: detecting whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism based on the action feature subsequence.

[0062] Specifically, detecting whether there is a behavioral pattern of repeated attempts to trigger the door lock opening mechanism based on the action feature subsequence refers to analyzing the behavioral timing characteristics and action similarity of the user's continuous operations based on the action feature subsequence in the candidate identity authentication stage to identify whether the user exhibits abnormal behavior of repeatedly attempting to unlock the door. In the specific implementation process, the system compares the action feature subsequences corresponding to different operations, extracts features such as posture similarity, motion trajectory consistency, and time interval regularity between each operation, and evaluates whether there is a "multiple repeated attempts" behavioral pattern based on a preset judgment threshold. For example, when a user attempts to enter a password through gestures or virtual keys, if the system detects that the user performs a highly similar hand movement sequence multiple times in a row, and the time interval between each movement is short and regular, it can be determined that there is a behavioral pattern of repeated attempts to trigger the door lock opening mechanism. This analysis method can effectively identify brute force or imitation attacks performed by malicious attackers without authorization, thereby providing a key basis for subsequent security policy decisions.

[0063] Step S5, based on the behavior pattern, determines whether the user's current access request meets the preset security policy rules, obtains a door lock control command signal, and sends the door lock control command signal to the door lock controller module in the smart lock to complete the corresponding operation execution.

[0064] Specifically, determining whether a user's current access request satisfies pre-set security policy rules based on the behavioral pattern involves detecting a user's repeated attempts to trigger the door lock unlocking mechanism. The system then compares this behavior pattern with the pre-set security policy to assess the legitimacy of the current access behavior and generates a corresponding door lock control command signal accordingly. This is achieved by setting behavioral feature thresholds, operation frequency limits, and an abnormal behavior scoring mechanism to determine whether the user's behavior exceeds the reasonable range of normal authentication. If the user's behavior is identified as abnormal, the system will deny the access request and trigger an alarm mechanism; otherwise, the unlocking operation will be allowed. For example, if a user performs highly similar gestures repeatedly with a short interval between them, the system will determine this as malicious probing behavior and generate a "Do not unlock" door lock control command signal. This signal is sent to the door lock controller module in the smart lock, preventing the door from opening and ensuring the safety of the device and the premises. Through this closed-loop control process, the system achieves a complete linkage from behavior perception, security decision-making, to physical control, improving the security and reliability of smart locks in complex application scenarios.

[0065] In a specific embodiment, capturing a continuous video frame sequence within the smart lock installation area and extracting a three-dimensional key point coordinate structure from the continuous video frame sequence includes:

[0066] Perform multi-angle image sampling on the ambient light field information in the area where the smart lock is installed to obtain a continuous video frame sequence including the access control channel area;

[0067] Performing a background difference operation based on the continuous video frame sequence to obtain a dynamic foreground mask layer, and performing a connected domain labeling process on the dynamic foreground mask layer to obtain a candidate human body contour area;

[0068] Boundary curve fitting and extreme point analysis are performed on each region in the candidate human body contour area to obtain local significant feature points including the positions of the head, shoulders, and hands, and a three-dimensional key point coordinate structure for characterizing the user's spatial posture is constructed based on the local significant feature points.

[0069] Specifically, capturing a continuous sequence of video frames within the smart lock installation area and extracting the three-dimensional coordinate structure of key points from this continuous sequence of video frames is a crucial first step in implementing a machine vision-based smart lock monitoring method. This step, through a series of image acquisition and processing operations, captures user motion information from the physical space and converts it into structured data that can be used for posture modeling and behavior analysis, providing a solid data foundation for subsequent identity verification, interaction recognition, and abnormal behavior detection. Specifically, the system first constructs a continuous sequence of video frames encompassing the access control passage area by sampling the ambient light field information from multiple angles within the smart lock installation area. This step typically relies on multiple cameras deployed around the smart lock, distributed at specific angles around the lock to ensure coverage of all possible directions from which a user might approach and operate the smart lock. For example, in a typical application scenario, the system can utilize three 1080P high-definition cameras, positioned above, to the lower left, and to the lower right of the smart lock, providing an effective field of view of approximately 120 degrees in front of and to the sides of the smart lock, thereby enabling full-view recording of the user entering the access control passage and approaching the smart lock. Each camera continuously captures video data at 30 frames per second, generating a raw video frame stream that forms a continuous video frame sequence. After obtaining the continuous video frame sequence, the system further performs a background subtraction operation on this sequence to extract dynamic foreground objects. This process is typically implemented using a Gaussian mixture model (GMM) or a deep learning-based background modeling algorithm (such as BackgroundSubtractorMOG2 in OpenCV). The system pre-establishes a stable background model and compares each image frame with the background model at the pixel level to isolate the areas of change, which is the dynamic foreground mask layer. For example, as a user approaches a smart lock, their body outline gradually stands out from the static background. The system can accurately capture this moving object through the background subtraction operation, providing basic data for subsequent human detection. Next, the system performs connected component labeling on the dynamic foreground mask layer to obtain candidate human outline regions. Connected component labeling is a typical image segmentation technique that focuses on identifying connected pixel regions in an image and assigning them unique label numbers. For example, in a real-world test, when the system detected two people simultaneously in front of a smart lock, two independent foreground regions appeared in the dynamic foreground mask layer generated by the background difference operation. Using connected domain labeling, the system identified these two regions and marked them as candidate human contour regions. The system then selected the primary target region most likely associated with the current interaction based on characteristics such as size, position, and shape, for subsequent key point extraction.The system then performs boundary curve fitting and extreme point analysis on each region within the candidate human contour to extract local salient feature points, including the head, shoulders, and hands. This process typically involves image processing techniques such as contour extraction, polynomial fitting, and curvature calculation. For example, after successfully locating a candidate human contour in a given frame, the system first performs Canny edge detection on its edges to extract the complete contour curve. It then fits the contour curve using a polynomial function to obtain a smooth mathematical representation. Based on this, the system further calculates the curvature variation along the contour curve, identifying the points of curvature maximum. These points often correspond to key body parts. For example, a top maximum point may represent the top of the user's head, two larger extreme points in the middle may correspond to the shoulders, and a lower extreme point that shifts significantly over time is likely to represent the user's hand movements. Finally, based on these extracted local salient feature points, the system constructs a three-dimensional keypoint coordinate structure to characterize the user's spatial posture. Because the image sampling is performed from multiple angles, the system can map key points in the 2D image to 3D space through stereo matching or depth estimation techniques. For example, the system uses the parallax difference generated when two cameras capture the same target object, combined with known camera intrinsic and extrinsic parameters, to estimate the coordinates of each key point in 3D space using a triangulation algorithm. For example, suppose that at a certain moment, the position of the user's right index finger is at (520, 340) in the left camera image and (610, 335) in the right camera image. Using parallax calculation, the system can infer the precise coordinates of this point in 3D space (x=1.2m, y=0.4m, z=0.9m). As video frames are continuously captured, these 3D key points are continuously updated over time, forming a time series trajectory that describes the user's posture changes, thereby forming a 3D key point coordinate structure.

[0070] In a specific embodiment, the step of performing posture change tracking modeling on the user based on the three-dimensional key point coordinate structure to obtain a user motion trajectory curve includes the following steps:

[0071] Analyzing the temporal features in the three-dimensional key point coordinate structure to obtain a key point motion trajectory sequence, and extracting posture change features in the key point motion trajectory sequence to obtain a multi-dimensional spatiotemporal feature matrix including joint angle changes, limb movement directions, and speed acceleration;

[0072] Performing dynamic spatiotemporal segmentation on the user's limb movements in a continuous video frame sequence based on the multidimensional spatiotemporal feature matrix to obtain a sequence of action segment boundary moments, and performing kinematic feature clustering analysis on the sequence of action segment boundary moments to obtain posture change features;

[0073] Performing action combination pattern matching on the posture change features through hierarchical spatiotemporal correlation analysis technology to obtain a user interaction intention feature vector, and performing temporal probability reasoning on the user interaction intention feature vector to obtain a temporal state transition diagram;

[0074] A dynamic programming search is performed on the timing state transition diagram to obtain an optimal state transition path sequence, and a user motion trajectory curve is constructed based on the optimal state transition path sequence, wherein the user motion trajectory curve includes interactive behavior characteristic parameters such as hand contact position, contact duration, and contact force change.

[0075] Specifically, the tracking and modeling of user posture changes based on the three-dimensional key point coordinate structure to obtain a user motion trajectory curve is the core technical link for achieving behavior perception and interaction intention recognition in this smart lock monitoring method. This step conducts an in-depth analysis of the three-dimensional key point coordinate structure extracted in the previous step, combines the motion characteristics and spatial distribution information in the time series, and constructs a motion trajectory curve that can characterize the dynamic evolution of the user's operation behavior, providing an accurate behavioral basis for the subsequent target interaction behavior event recognition. Specifically, after obtaining the user's three-dimensional key point coordinate structure, the system first analyzes its temporal characteristics, that is, tracking the position, direction, and speed of each key point over time in a continuous video frame sequence, thereby forming a key point motion trajectory sequence. For example, in an actual test scenario, the system obtains the position data of the user's hand key points in three-dimensional space at a sampling frequency of 30 frames per second, and arranges this data in chronological order to form a time series containing hundreds of coordinate points, which is used to describe the entire process of the hand from moving away from the smart lock to approaching the operation interface and then performing the input action. Based on this, the system further extracts posture change features from the key point motion trajectory sequence, including multi-dimensional spatiotemporal features such as joint angle changes, limb movement direction, and velocity and acceleration. These features are ultimately combined into a multi-dimensional spatiotemporal feature matrix. For example, when a user attempts to unlock their phone, their wrist angle changes from 90 degrees to 150 degrees within one second, while their hand movement speed rapidly increases from 0.2 m / s to 0.8 m / s. These numerical features are integrated into the multi-dimensional spatiotemporal feature matrix and serve as important input for posture modeling. The system then performs dynamic spatiotemporal segmentation of the user's limb movements in a continuous video frame sequence based on this multi-dimensional spatiotemporal feature matrix to identify the boundary moments between different actions. This process relies on a combination of a sliding window algorithm and cluster analysis techniques. For example, the system uses a fixed-length sliding window (e.g., each window contains 10 frames) to segment the entire video sequence and calculate a feature similarity index within each window. When the feature difference between a window and an adjacent window exceeds a set threshold, the system marks it as a potential action segment boundary. In this way, the system can accurately identify key time nodes such as when the user switches from a standing state to a hand-raising action and then to a click operation, thereby obtaining a set of action segment boundary moment sequences. Next, the system performs kinematic feature clustering analysis on these boundary moment sequences, that is, groups segments with similar motion patterns into one category, thereby extracting more representative posture change features. For example, in multiple tests, it was found that when a user performs a "gesture input" operation, their hand usually goes through four stages of "lifting-hovering-swiping-falling back". The system can identify these four stages separately through a clustering algorithm and establish corresponding feature models. Furthermore, the system uses hierarchical spatiotemporal correlation analysis technology to match the posture change features with action combination patterns to identify the user's interaction intention.This process combines sequence modeling methods from deep learning (such as LSTM networks) with traditional rule-matching mechanisms to capture the contextual relationships between actions. For example, in a typical application scenario, the system detects a user's hand moving in a sequence of "approaching, briefly pausing, and quickly pressing down." Combining this with a library of action patterns trained using historical data, the system determines that this action is highly likely a password entry or gesture verification operation and generates a corresponding user interaction intention feature vector. This feature vector not only contains information about the action type but also records key parameters such as the temporal sequence and duration of the actions. To more accurately understand the temporal evolution of user behavior, the system further performs temporal probabilistic reasoning on the user interaction intention feature vector to construct a temporal state transition graph that describes the state transitions of user behavior. This graph uses nodes to represent different behavior states (such as "still," "hand raised," "touch," and "input completed"), and edges to represent the transition probabilities between states. For example, based on training with a large number of real-world samples, the system determined that the probability of transitioning from the "still" state to the "hand raised" state is 75%, while the probability of transitioning from the "hand raised" state to the "touch" state is 60%. This statistical modeling approach enables the system to predict the user's next possible action and adjust the monitoring strategy accordingly. Finally, the system performs a dynamic programming search on the time-series state transition diagram to find an optimal state transition path sequence, that is, a behavioral evolution path that best matches the current observation data. For example, during an actual use, the system evaluated all possible paths using a dynamic programming algorithm and found that the path of "rest → raise hand → hover → click → return to rest" had the highest probability, so it was determined as the current user's action trajectory. On this basis, the system constructs a user action trajectory curve based on this optimal state transition path sequence, which includes not only the hand's motion trajectory in three-dimensional space, but also key interactive behavior characteristic parameters such as the hand contact position (such as a point on the smart lock panel), contact duration (such as the finger's dwell time of 0.4 seconds), and contact force changes (indirectly estimated through acceleration changes).

[0076] In a specific embodiment, the dynamic spatiotemporal segmentation of the user's limb movements in a continuous video frame sequence based on the multidimensional spatiotemporal feature matrix to obtain a sequence of action segment boundary moments includes the following steps:

[0077] Performing temporal gradient analysis on the multidimensional spatiotemporal feature matrix to obtain a feature change rate curve, and performing adaptive threshold segmentation based on the feature change rate curve to obtain initial action boundary candidate points, wherein the initial action boundary candidate points include time indexes at which features change significantly;

[0078] Constructing a spatiotemporal feature map based on the initial action boundary candidate points, and applying non-local mean filtering to the spatiotemporal feature map to obtain a smoothed feature map, wherein the smoothed feature map retains key action transition information while suppressing small noise disturbances;

[0079] A hierarchical clustering operation is performed on the smoothed feature graph to obtain an action fragment tree structure, and multi-scale boundary optimization is performed based on the action fragment tree structure to obtain an action fragment boundary moment sequence, wherein the action fragment boundary moment sequence reflects the key action sub-nodes in the process of user interaction with the smart lock.

[0080] Specifically, the dynamic spatiotemporal segmentation of user body movements in a continuous video frame sequence based on the multidimensional spatiotemporal feature matrix to obtain a sequence of action segment boundary moments is a key technical step in achieving structured modeling of user behavior in this smart lock monitoring method. This step uses a series of time series analysis and image processing techniques to identify semantically meaningful action transition nodes from complex three-dimensional keypoint motion data, providing a temporally accurate basis for action segmentation for subsequent behavior recognition and security assessment. After obtaining the user's multidimensional spatiotemporal feature matrix, the system first performs temporal gradient analysis to capture the changing trends of the user's movements. This multidimensional spatiotemporal feature matrix contains information such as the position, angle, velocity, and acceleration of multiple body parts (such as the head, shoulders, and hands) at different time frames, forming a high-dimensional temporal feature space. For example, in a real-world test, the system collected 12 dimensions of data, including the three-dimensional coordinates, joint angles, and movement velocity of a user's hand, over a continuous 10-second period while attempting to enter a gesture password. This data was organized into a 300×12 matrix (each row representing a frame). The system then performs a frame-by-frame difference calculation on this matrix, comparing the magnitude of change in each eigenvalue between adjacent frames to generate a feature change rate curve. This curve reflects the temporal intensity of the user's action, typically manifested as steep or flat areas on the curve. For example, the feature change rate curve shows a distinct peak when a user quickly raises their hand to approach the smart lock panel, while it tends to plateau during the stationary phase. Based on this feature change rate curve, the system performs adaptive threshold segmentation to identify initial action boundary candidate points. Specifically, the system dynamically adjusts the segmentation threshold based on historical data to ensure reliable detection of key action transition moments even in complex environments such as lighting fluctuations and background interference. For example, within a certain time period, if the feature change rate exceeds a set threshold (e.g., a rate of change greater than 0.8 for five consecutive frames), the system marks those frames as potential action boundary points, forming an initial set of action boundary candidate points. These candidate points typically correspond to transitions in user behavior, such as key moments like "starting to raise your hand," "finger touching the panel," and "completing the action and retracting it." Furthermore, the system constructs a spatiotemporal feature map based on the initial action boundary candidate points, that is, maps the local features around each candidate point into a two-dimensional image form for the subsequent application of image processing algorithms. For example, the system takes 10 frames of data before and after each candidate point to form a window containing 21 frames, and converts the multi-dimensional feature vector of each frame into pixel values ​​in a grayscale image, and finally forms a spatiotemporal feature map that reflects the changes in user movements. Then, the system applies non-local mean filtering technology to the spatiotemporal feature map to suppress small disturbances caused by factors such as camera shake and environmental noise. Non-local mean filtering is a denoising algorithm based on similar block matching, which can effectively remove random noise while retaining image edge details.For example, in one experiment, the system discovered a large number of small noise points in the original spatiotemporal feature map. After non-local means filtering, the image became clearer, the outlines of action transitions became more distinct, and key action transition information was not lost. The system then performed hierarchical clustering on the smoothed feature map to identify action structures at different granularities. Hierarchical clustering is an unsupervised learning method that does not require a preset number of clusters and automatically constructs a tree-like structure based on similarities between data. For example, in one test, the system divided the smoothed feature map into several subregions and performed cluster analysis by calculating the Euclidean distance between each region. This ultimately resulted in a three-layered action segment tree structure. The first layer represents the overall action cycle, the second layer represents the main action phases (e.g., "prepare - execute - end"), and the third layer is broken down into more specific action units (e.g., "raise - hover - click - retract"). This hierarchical structure allows the system to analyze user behavior from a macro to micro level, enhancing the robustness of action recognition. Finally, the system performs multi-scale boundary optimization based on the action segment tree structure to accurately locate the start and end moments of the action segments. This process combines coarse-grained and fine-grained boundary information and selects the optimal timestamp combination through a voting mechanism or dynamic programming. For example, in a certain operation, the system recognized that the entire gesture input process lasted about 3 seconds at a coarse-grained level, and was further divided into 4 sub-actions at a fine-grained level, with each sub-action lasting an average of 0.75 seconds. By fusing and optimizing the results at different scales, the system ultimately obtains a set of accurate action segment boundary moment sequences, where each boundary point corresponds to a key action transition node in the user's interaction with the smart lock, such as "hand leaving the static state", "first touch of the operation interface", "completion of input and retraction of hand", etc.

[0081] In a specific embodiment, extracting key frames from the target interactive behavior event to obtain an action feature subsequence in the candidate identity verification stage includes:

[0082] Performing spatiotemporal analysis on the video sequence of the target interactive behavior event to obtain an action change amplitude curve, and setting an adaptive threshold based on the action change amplitude curve to obtain a key moment index set, wherein the key moment index set includes timestamps of the action start, action apex, and action end;

[0083] Performing frame-level sampling on a continuous video frame sequence based on the key moment index set to obtain candidate key frames, and performing multi-scale image pyramid decomposition on the candidate key frames to obtain a multi-resolution feature map;

[0084] Performing spectral clustering analysis on the multi-resolution feature map to obtain a feature similarity matrix, constructing an inter-frame association network based on the feature similarity matrix, and determining a key frame representativeness score for each of the continuous video frame sequences based on the inter-frame association network, wherein the key frame representativeness score reflects the contribution of each frame image to the overall behavior representation;

[0085] The candidate key frames are screened and sorted based on the key frame representativeness scores to obtain a key frame sequence, and a candidate identity verification stage action feature subsequence is extracted from the key frame sequence.

[0086] Specifically, extracting keyframes from the target interactive behavior event to obtain a candidate authentication phase motion feature subsequence is a key technical step in achieving behavioral information compression and core action recognition in this smart lock monitoring method. This step constructs a concise yet comprehensive subset of behavioral features by filtering the most representative time points and image frames from the continuous video frame sequence of the target interactive behavior event, thereby providing efficient and accurate data support for subsequent abnormal behavior detection. Specifically, after the system identifies the target interactive behavior event, in which a user approaches the smart lock and begins interacting with the user interface, it first performs spatiotemporal analysis on the video sequence to obtain a motion change amplitude curve. This process relies on a joint analysis of the motion trajectories of three-dimensional key points and the content of their corresponding image frames. For example, in one actual test, a user's hand went through four phases: "lift-hover-swipe-retract" while attempting to enter a gesture password. The system comprehensively evaluated the displacement, angle change, and velocity fluctuation of the hand key points in each frame to generate a motion change amplitude curve reflecting the intensity of the movement. This curve typically exhibits a pattern of alternating peaks and troughs, with the peak region representing the moment of most significant motion change. Subsequently, the system sets an adaptive threshold based on the curve, automatically identifies the timestamps of the start, peak and end of the action, and groups these time points into a key moment index set. For example, in a video containing 120 frames, the system successfully identifies the 15th frame as "hand raising start", the 45th frame as "gesture completion", and the 78th frame as "hand retraction", forming a set of accurate action time nodes. On this basis, the system performs frame-level sampling on the continuous video frame sequence based on the key moment index set, that is, selects several frames near the key moment as candidate key frames. For example, the system takes 3 frames before and after each key moment (a total of 7 frames) to form a preliminary set of candidate key frames. In order to further improve the quality and semantic expression ability of the key frames, the system performs multi-scale image pyramid decomposition on the candidate key frames, that is, uses methods such as Gaussian pyramid or wavelet transform to decompose each frame into multiple feature maps at different resolutions. For example, in a key frame, the original image resolution is 640×480 pixels. After three-layer pyramid decomposition, the system obtains low-resolution images with resolutions of 320×240, 160×120, and 80×60, respectively, while retaining important visual features such as edges and textures in the image. This multi-resolution representation helps the system capture user action details at different scales and improves the representation capabilities of key frames. Next, the system performs spectral clustering analysis on the multi-resolution feature maps to identify image frames with similar visual features. Spectral clustering is an unsupervised learning method based on graph theory that can group frames based on their similarities.For example, in one experiment, the system input the image feature vectors of all candidate keyframes into a spectral clustering algorithm, ultimately generating four clusters: the first cluster contained frames during the initial hand-lift phase, the second contained frames during the middle phase of the gesture, the third contained frames during the touch panel phase, and the fourth contained frames during the hand-retraction phase. This approach effectively distinguished the visual representations of different gesture phases and further constructed an inter-frame association network. This network uses image frames as nodes and the similarity between frames as edge weights, describing the strength of associations between frames. For example, the system found a high similarity between frames 30 and 32 (a similarity of 0.92), while the similarity between frames 15 and 70 was lower (only 0.35), establishing a connection between the frames. Based on this inter-frame association network, the system further calculated a keyframe representativeness score for each candidate keyframe. This score reflects the frame's contribution to the overall representation of the target interactive behavior event. A higher score indicates that the frame more effectively represents the core features of the entire action. For example, in one test, the system discovered that frame 45 coincided with the moment a gesture was completed and had the highest inter-frame connectivity, thus assigning it the highest representativeness score (e.g., 0.98). Frame 10, however, received a score of 0.45 because it merely indicated that the user had not yet entered the action state. This scoring mechanism effectively identifies frames with genuine behavioral semantic value and prevents redundant frames from interfering with subsequent analysis. Finally, the system screens and ranks candidate keyframes based on their representativeness scores, generating an optimized keyframe sequence. For example, given a video with 100 candidate frames, the system selects the top 20 most representative frames based on their score rankings and arranges them in chronological order to form the final keyframe sequence. Based on this, the system further extracts candidate action feature subsequences from this sequence, specifically those frames that focus on the user performing the authentication operation. For example, in a real-world scenario, the system identified frames 25 to 50 as containing the complete gesture input process. These frames were then extracted as candidate action feature subsequences for the authentication phase for subsequent behavioral pattern analysis.

[0087] In a specific embodiment, performing spatiotemporal analysis on the video sequence of the target interactive behavior event to obtain the action change amplitude curve includes:

[0088] Performing optical flow calculation on the video sequence of the target interactive behavior event to obtain a motion vector map, and performing motion gradient decomposition on the motion vector map to obtain a three-dimensional motion component tensor including horizontal displacement, vertical displacement, and angular rotation;

[0089] Performing multi-scale time-frequency analysis on the three-dimensional motion component tensor through wavelet transform to obtain a motion feature spectrum, and performing energy density calculation on the motion feature spectrum to obtain a motion energy distribution sequence, wherein the motion energy distribution sequence describes the frequency characteristics of motion changes;

[0090] Performing time domain segmentation processing based on the motion energy distribution sequence to obtain action segment boundary points, and performing smooth interpolation operation on the action segment boundary points to obtain a continuous motion state curve;

[0091] The continuous motion state curve is amplitude normalized to obtain a standardized motion feature sequence, and a motion change amplitude curve is constructed based on the standardized motion feature sequence, wherein the motion change amplitude curve includes time series feature parameters of motion intensity, motion duration, and motion continuity.

[0092] Specifically, performing spatiotemporal analysis of the video sequence of the target interactive behavior event to generate a motion amplitude curve is a key technical step in extracting dynamic user behavior features within this smart lock monitoring method. This step uses a series of image processing and signal analysis techniques to extract temporal feature parameters reflecting the intensity, duration, and coherence of the user's motion from the continuous video frame sequence of the target interactive behavior event. Ultimately, a motion amplitude curve is constructed that accurately describes the changing trends in the user's operational behavior. After the system identifies the target interactive behavior event, in which a user approaches the smart lock and begins interacting with the interface, it first performs optical flow calculations on the video sequence to obtain a motion vector diagram between each frame. Optical flow is a classic motion estimation technique that estimates the direction and speed of an object's motion by comparing pixel displacements between adjacent frames. For example, in a real-world test, a user attempting to enter a gesture password experienced four stages of hand movement: raising their hand, hovering, swiping, and retracting. The system used the Horn-Schunck optical flow algorithm to compare each frame, generating a motion vector diagram containing the direction and magnitude of motion for each pixel. The system then further performs motion gradient decomposition on these motion vector maps, breaking down the motion information within the two-dimensional plane into three-dimensional motion component tensors: horizontal displacement, vertical displacement, and angular rotation. For example, in one operation, the user's right hand moved approximately 15 centimeters to the upper right, accompanied by a clockwise rotation of approximately 20 degrees. The system records this information as a data structure showing an increase in horizontal displacement, an increase in vertical displacement, and a positive angular rotation. Next, the system performs multi-scale time-frequency analysis on these three-dimensional motion component tensors using wavelet transforms to extract motion frequency characteristics at different time scales. Wavelet transforms are a mathematical tool suitable for analyzing non-stationary signals. They decompose the original signal into multiple sub-signals with different frequency bandwidths, facilitating the capture of periodic and sudden changes in motion. For example, in one experiment, the system decomposed the user's motion signal into four scale layers: layer 1 reflects high-frequency, rapid movements (such as finger taps), layers 2-3 reveal medium-frequency movements (such as hand movements), and layer 4 characterizes low-frequency, overall posture changes (such as leaning forward). By statistically analyzing the energy distribution at each scale, the system further calculates the motion energy distribution sequence, which describes the degree of energy concentration of the action in different frequency bands. For example, in a certain test, it was found that when the user performs gesture input, the action is mainly concentrated in the 2nd and 3rd layers, accounting for 75% of the energy, while the 1st layer only accounts for 15%, indicating that the operation is mainly medium-speed movement and lacks drastic sudden changes. On this basis, the system performs time domain segmentation processing based on the motion energy distribution sequence, that is, identifying the different stages of the action and its boundary points. This process usually relies on a method that combines sliding window detection with threshold judgment.For example, the system sets a sliding window of 10 frames and scans the energy sequence frame by frame. When the energy change rate in the window exceeds the set threshold (such as the change rate for 5 consecutive frames is greater than 0.6), it is considered that a change in the action state has occurred and is marked as an action segment boundary point. Suppose that in a video containing 120 frames, the system successfully identifies 3 obvious action transition nodes: the 20th frame is "hand raising to start", the 55th frame is "gesture completion", and the 90th frame is "hand retraction". These time nodes constitute a preliminary set of action segment boundary points. In order to improve the smoothness and accuracy of the action boundary, the system further performs smooth interpolation operations on the action segment boundary points. Specifically, the system uses cubic spline interpolation or linear interpolation to fit the discontinuous action boundary points to generate a continuous motion state curve. For example, during a certain operation, the originally detected boundary points exhibited jumps (e.g., at frames 18, 53, and 89). After interpolation, the system adjusted these to more reasonable frames 20, 55, and 90, making the entire curve more natural and smooth, consistent with the temporal evolution of human motion. Finally, the system normalized the amplitude of the continuous motion state curve to eliminate the influence of individual differences and environmental interference. Normalization involves scaling the maximum value of the curve to a fixed range (e.g., [0, 1]) to ensure comparability across different operations. For example, in one test, the system set the maximum value of the original motion state curve to 1 and scaled the remaining values ​​proportionally to ensure that the amplitude of motion changes for all users was consistent. The normalized data was organized into a standardized motion feature sequence, which served as the basis for constructing the motion change amplitude curve. The movement change amplitude curve not only includes movement intensity (such as the peak value reaches 0.9) and movement duration (such as the high energy state lasts for 2 seconds), but also covers key timing characteristic parameters such as movement continuity (such as small curve fluctuations, indicating smooth movement).

[0093] In a specific embodiment, the detecting whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism based on the action feature subsequence includes:

[0094] Performing spatiotemporal pattern decomposition on the motion feature subsequence to obtain periodic motion segments, and performing phase alignment processing on the periodic motion segments to generate a standardized motion trajectory template;

[0095] Performing sequence similarity matching on the standardized motion trajectory template based on dynamic time warping to obtain a repeated attempt behavior matching matrix, and detecting whether there is a suspicious behavior marker sequence in the repeated attempt behavior matching matrix;

[0096] Modeling the state transition probability of the suspicious behavior mark sequence through a hidden Markov model to obtain a behavior abnormality score, and performing a time series cumulative calculation of the behavior abnormality score based on Bayesian inference to generate a comprehensive risk coefficient;

[0097] Adaptive threshold segmentation is used to divide the comprehensive risk coefficient into risk levels to obtain a behavior pattern determination result, and a logical check is performed on the behavior pattern determination result to generate a door lock triggering behavior analysis report. Based on the door lock triggering behavior analysis report, it is determined whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism; wherein, the door lock triggering behavior analysis report includes the number of repeated attempts, behavior consistency and risk confidence.

[0098] Specifically, detecting whether there are repeated attempts to trigger the door lock's opening mechanism based on the action feature subsequence is a key step in achieving abnormal behavior identification and security decision-making in this smart lock monitoring method. This step systematically determines whether the user is engaging in illegal probing or brute force attacks by conducting in-depth analysis of the action feature subsequences from the candidate authentication phase, combining various technical approaches such as periodic action modeling, dynamic time matching, probabilistic reasoning, and risk assessment, thereby providing a basis for subsequent security response. Specifically, after obtaining the action feature subsequences from the candidate authentication phase, the system first performs spatiotemporal pattern decomposition to identify possible periodic action segments. For example, in an actual test, a user performed four similar raise-swipe-retract actions in succession while attempting to enter an incorrect gesture password, with each action occurring approximately 1.5 seconds apart and lasting approximately 2 seconds. Using a combination of sliding window scanning and spectral analysis, the system successfully identified these four highly similar action segments and labeled them as potential "repeated attempts." The system then performs phase alignment on these periodic motion segments. This involves shifting and scaling the timeline to align the start, peak, and end points of the motions within different periods, thereby generating a standardized motion trajectory template. For example, the original start times of the four motions in a given operation are frames 10, 35, 60, and 85, respectively. After phase alignment, these are uniformly mapped to the interval between frames 0 and 20, forming a unified time base for subsequent comparison and analysis. Based on this, the system performs sequence similarity matching on these standardized motion trajectory templates using the dynamic time warping (DTW) algorithm, calculating the degree of motion similarity between each attempt and generating a matching matrix for repeated attempts. DTW is a sequence matching method suitable for nonlinear time alignment and is particularly well-suited for comparing temporally distinct but semantically identical motion sequences. For example, in the aforementioned test scenario, the system uses the first motion as the reference template and performs DTW comparisons with the other three motions, resulting in similarity scores of 0.92, 0.89, and 0.87 (out of a maximum score of 1), indicating that the three attempts are highly similar in terms of motion morphology. These scores are organized into a 4×4 matching matrix to quantify the similarity between different attempts. The system further detects whether there is a suspicious behavior marker sequence in the matrix, that is, it marks the action combination with a similarity higher than a set threshold (such as 0.85) and more than 3 occurrences as a suspected repeated attempt behavior. For example, in a certain video, the system found that the similarity of 4 consecutive gesture input actions was higher than 0.85, so it was determined to be a suspicious behavior marker sequence. Next, the system uses the hidden Markov model (HMM) to model the state transition probability of the suspicious behavior marker sequence to capture the evolution of user behavior in the time dimension.The HMM is a classic temporal probability model that effectively describes the transition from a user's "normal attempt" to an "abnormal attempt." For example, during the training phase, the system leverages extensive historical data to construct an HMM model with two hidden states ("normal" and "abnormal") and learns the transition probabilities between the states (e.g., 0.8 for "normal→normal" and 0.2 for "normal→abnormal"). In practice, the system inputs the current user's action sequence into the model and calculates the probability of being in the "abnormal" state at each moment, ultimately generating a time series reflecting the behavioral abnormality score. For example, in one operation, a user performed four highly similar gestures in succession. The system calculated behavioral abnormality scores of 0.91, 0.93, 0.95, and 0.96, indicating that the user's behavior increasingly deviated from normal patterns. To enhance the robustness of the assessment results, the system further accumulates these behavioral abnormality scores over time using Bayesian inference to generate a comprehensive risk factor. Bayesian inference is a method for updating posterior probabilities based on prior knowledge. It can combine historical behavior with current observations to improve prediction accuracy. For example, in one test, the system initially set a user's prior risk probability to 0.3 (considering them normal). With four consecutive anomaly scores, the system gradually updated the posterior risk probability, ultimately determining a risk factor of 0.97 for the user, indicating that their behavior was extremely high risk. This comprehensive risk factor not only considers the degree of abnormality of individual behaviors but also incorporates the temporal evolution of the entire sequence of behaviors, providing enhanced discriminative power. Finally, the system uses adaptive threshold segmentation to categorize this comprehensive risk factor, automatically identifying high-risk behaviors and generating a door lock triggering behavior analysis report. For example, the system sets three risk levels: low (<0.5), medium (0.5–0.8), and high (>0.8). When a user's comprehensive risk factor reaches 0.97, the system categorizes the behavior as high risk and generates a detailed door lock triggering behavior analysis report. This report includes key metrics such as the number of repeated attempts (e.g., 4), behavioral consistency (e.g., an average similarity of 0.89 between attempts), and risk confidence (e.g., 0.97), which are used by the subsequent logic verification module. The system further performs logic verification on this analysis report, such as determining whether there are any misidentifications and whether it is physically feasible (e.g., whether the interval between actions is reasonable). Ultimately, it confirms whether there is a behavioral pattern that repeatedly triggers the door lock's opening mechanism.

[0099] The above describes the smart lock monitoring method based on machine vision in the embodiment of the present invention. The following describes the smart lock monitoring system based on machine vision in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a smart lock monitoring system based on machine vision includes:

[0100] Capturing module 21, used to capture a continuous video frame sequence within the smart lock installation area and extract a three-dimensional key point coordinate structure from the continuous video frame sequence;

[0101] A modeling module 22 is configured to track and model the user's posture changes based on the three-dimensional key point coordinate structure to obtain a user motion trajectory curve, and identify target interaction behavior events approaching the operation interface in the smart lock in combination with the motion trajectory curve;

[0102] An extraction module 23 is used to extract key frames from the target interactive behavior event to obtain an action feature subsequence of the candidate identity verification stage;

[0103] A triggering module 24 is configured to detect whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism based on the action feature subsequence;

[0104] The judgment module 25 is used to judge whether the user's current access request meets the preset security policy rules based on the behavior pattern, obtain the door lock control command signal, and send the door lock control command signal to the door lock controller module in the smart lock to complete the corresponding operation execution.

[0105] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to the above method embodiment, which will not be repeated here.

[0106] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided, wherein the internal structure of the computer device can be as follows: Figure 3 As shown. The computer device includes a processor, memory, display screen, input device, network interface and database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.

[0107] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0108] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-described method when executed by a processor. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0109] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM.

[0110] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.

[0111] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A smart lock monitoring method based on machine vision, characterized in that: The following steps are involved: Capturing a continuous video frame sequence within the smart lock installation area, and extracting a three-dimensional key point coordinate structure from the continuous video frame sequence; Based on the three-dimensional key point coordinate structure, the user's posture change tracking model is performed to obtain a user motion trajectory curve, and the target interactive behavior event approaching the operation interface in the smart lock is identified in combination with the motion trajectory curve; Extract key frames from the target interactive behavior event to obtain a subsequence of action features in the candidate identity verification stage; Detecting whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism based on the action feature subsequence; Based on the behavior pattern, it is determined whether the user's current access request meets the preset security policy rules, a door lock control command signal is obtained, and the door lock control command signal is sent to the door lock controller module in the smart lock to complete the corresponding operation execution; The method of performing posture change tracking modeling on the user based on the three-dimensional key point coordinate structure to obtain a user motion trajectory curve includes the following steps: Analyzing the temporal features in the three-dimensional key point coordinate structure to obtain a key point motion trajectory sequence, and extracting posture change features in the key point motion trajectory sequence to obtain a multi-dimensional spatiotemporal feature matrix including joint angle changes, limb movement directions, and speed and acceleration; Performing dynamic spatiotemporal segmentation on the user's limb movements in a continuous video frame sequence based on the multidimensional spatiotemporal feature matrix to obtain a sequence of action segment boundary moments, and performing kinematic feature clustering analysis on the sequence of action segment boundary moments to obtain posture change features; Performing action combination pattern matching on the posture change features through hierarchical spatiotemporal correlation analysis technology to obtain a user interaction intention feature vector, and performing temporal probability reasoning on the user interaction intention feature vector to obtain a temporal state transition diagram; Performing a dynamic programming search on the time-series state transition graph to obtain an optimal state transition path sequence, and constructing a user motion trajectory curve based on the optimal state transition path sequence, wherein the user motion trajectory curve includes interactive behavior characteristic parameters such as hand contact position, contact duration, and contact force change; The method of performing dynamic spatiotemporal segmentation on the user's limb movements in a continuous video frame sequence based on the multidimensional spatiotemporal feature matrix to obtain a sequence of action segment boundary moments includes the following steps: Performing temporal gradient analysis on the multidimensional spatiotemporal feature matrix to obtain a feature change rate curve, and performing adaptive threshold segmentation based on the feature change rate curve to obtain initial action boundary candidate points, wherein the initial action boundary candidate points include time indexes at which features change significantly; Constructing a spatiotemporal feature map based on the initial action boundary candidate points, and applying non-local mean filtering to the spatiotemporal feature map to obtain a smoothed feature map, wherein the smoothed feature map retains key action transition information while suppressing small noise disturbances; A hierarchical clustering operation is performed on the smoothed feature graph to obtain an action fragment tree structure, and multi-scale boundary optimization is performed based on the action fragment tree structure to obtain an action fragment boundary moment sequence, wherein the action fragment boundary moment sequence reflects the key action sub-nodes in the process of user interaction with the smart lock.

2. The method for monitoring a smart lock based on machine vision according to claim 1, characterized in that: The method of capturing a continuous video frame sequence within the smart lock installation area and extracting a three-dimensional key point coordinate structure from the continuous video frame sequence includes: Perform multi-angle image sampling on the ambient light field information in the area where the smart lock is installed to obtain a continuous video frame sequence including the access control channel area; Performing a background difference operation based on the continuous video frame sequence to obtain a dynamic foreground mask layer, and performing a connected domain labeling process on the dynamic foreground mask layer to obtain a candidate human body contour area; Boundary curve fitting and extreme point analysis are performed on each region in the candidate human body contour area to obtain local significant feature points including the positions of the head, shoulders, and hands, and a three-dimensional key point coordinate structure for characterizing the user's spatial posture is constructed based on the local significant feature points.

3. The method for monitoring a smart lock based on machine vision according to claim 1, characterized in that: The step of extracting key frames from the target interactive behavior event to obtain an action feature subsequence in the candidate identity verification stage includes: Performing spatiotemporal analysis on the video sequence of the target interactive behavior event to obtain an action change amplitude curve, and setting an adaptive threshold based on the action change amplitude curve to obtain a key moment index set, wherein the key moment index set includes timestamps of the action start, action apex, and action end; Performing frame-level sampling on a continuous video frame sequence based on the key moment index set to obtain candidate key frames, and performing multi-scale image pyramid decomposition on the candidate key frames to obtain a multi-resolution feature map; Performing spectral clustering analysis on the multi-resolution feature map to obtain a feature similarity matrix, constructing an inter-frame association network based on the feature similarity matrix, and determining a key frame representativeness score for each of the continuous video frame sequences based on the inter-frame association network, wherein the key frame representativeness score reflects the contribution of each frame image to the overall behavior representation; The candidate key frames are screened and sorted based on the key frame representativeness scores to obtain a key frame sequence, and a candidate identity verification stage action feature subsequence is extracted from the key frame sequence.

4. The method for monitoring a smart lock based on machine vision according to claim 3, characterized in that: The performing of spatiotemporal analysis on the video sequence of the target interactive behavior event to obtain an action change amplitude curve includes: Performing optical flow calculation on the video sequence of the target interactive behavior event to obtain a motion vector map, and performing motion gradient decomposition on the motion vector map to obtain a three-dimensional motion component tensor including horizontal displacement, vertical displacement, and angular rotation; Performing multi-scale time-frequency analysis on the three-dimensional motion component tensor through wavelet transform to obtain a motion feature spectrum, and performing energy density calculation on the motion feature spectrum to obtain a motion energy distribution sequence, wherein the motion energy distribution sequence describes the frequency characteristics of motion changes; Performing time domain segmentation processing based on the motion energy distribution sequence to obtain action segment boundary points, and performing smooth interpolation operation on the action segment boundary points to obtain a continuous motion state curve; The continuous motion state curve is amplitude normalized to obtain a standardized motion feature sequence, and a motion change amplitude curve is constructed based on the standardized motion feature sequence, wherein the motion change amplitude curve includes time series feature parameters of motion intensity, motion duration, and motion continuity.

5. The method for monitoring a smart lock based on machine vision according to claim 1, characterized in that: The detecting whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism based on the action feature subsequence includes: Performing spatiotemporal pattern decomposition on the motion feature subsequence to obtain periodic motion segments, and performing phase alignment processing on the periodic motion segments to generate a standardized motion trajectory template; Performing sequence similarity matching on the standardized motion trajectory template based on dynamic time warping to obtain a repeated attempt behavior matching matrix, and detecting whether there is a suspicious behavior marker sequence in the repeated attempt behavior matching matrix; Modeling the state transition probability of the suspicious behavior mark sequence through a hidden Markov model to obtain a behavior abnormality score, and performing a time series cumulative calculation of the behavior abnormality score based on Bayesian inference to generate a comprehensive risk coefficient; Adaptive threshold segmentation is used to divide the comprehensive risk coefficient into risk levels to obtain a behavior pattern determination result, and a logical check is performed on the behavior pattern determination result to generate a door lock triggering behavior analysis report. Based on the door lock triggering behavior analysis report, it is determined whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism; wherein, the door lock triggering behavior analysis report includes the number of repeated attempts, behavior consistency and risk confidence.

6. A smart lock monitoring system based on machine vision, characterized in that: include: A capture module is used to capture a continuous video frame sequence within the smart lock installation area and extract a three-dimensional key point coordinate structure from the continuous video frame sequence; A modeling module is used to track and model the user's posture changes based on the three-dimensional key point coordinate structure to obtain a user motion trajectory curve, and identify target interactive behavior events approaching the operation interface in the smart lock in combination with the motion trajectory curve; An extraction module, configured to extract key frames from the target interactive behavior event to obtain an action feature subsequence in the candidate identity verification stage; A triggering module, configured to detect whether there is a behavior pattern of multiple repeated attempts to trigger the door lock opening mechanism based on the action feature subsequence; A judgment module is used to judge whether the user's current access request meets the preset security policy rules based on the behavior pattern, obtain a door lock control command signal, and send the door lock control command signal to the door lock controller module in the smart lock to complete the corresponding operation execution; The method of performing posture change tracking modeling on the user based on the three-dimensional key point coordinate structure to obtain a user motion trajectory curve includes the following steps: Analyzing the temporal features in the three-dimensional key point coordinate structure to obtain a key point motion trajectory sequence, and extracting posture change features in the key point motion trajectory sequence to obtain a multi-dimensional spatiotemporal feature matrix including joint angle changes, limb movement directions, and speed and acceleration; Performing dynamic spatiotemporal segmentation on the user's limb movements in a continuous video frame sequence based on the multidimensional spatiotemporal feature matrix to obtain a sequence of action segment boundary moments, and performing kinematic feature clustering analysis on the sequence of action segment boundary moments to obtain posture change features; Performing action combination pattern matching on the posture change features through hierarchical spatiotemporal correlation analysis technology to obtain a user interaction intention feature vector, and performing temporal probability reasoning on the user interaction intention feature vector to obtain a temporal state transition diagram; Performing a dynamic programming search on the time-series state transition graph to obtain an optimal state transition path sequence, and constructing a user motion trajectory curve based on the optimal state transition path sequence, wherein the user motion trajectory curve includes interactive behavior characteristic parameters such as hand contact position, contact duration, and contact force change; The method of performing dynamic spatiotemporal segmentation on the user's limb movements in a continuous video frame sequence based on the multidimensional spatiotemporal feature matrix to obtain a sequence of action segment boundary moments includes the following steps: Performing temporal gradient analysis on the multidimensional spatiotemporal feature matrix to obtain a feature change rate curve, and performing adaptive threshold segmentation based on the feature change rate curve to obtain initial action boundary candidate points, wherein the initial action boundary candidate points include time indexes at which features change significantly; Constructing a spatiotemporal feature map based on the initial action boundary candidate points, and applying non-local mean filtering to the spatiotemporal feature map to obtain a smoothed feature map, wherein the smoothed feature map retains key action transition information while suppressing small noise disturbances; A hierarchical clustering operation is performed on the smoothed feature graph to obtain an action fragment tree structure, and multi-scale boundary optimization is performed based on the action fragment tree structure to obtain an action fragment boundary moment sequence, wherein the action fragment boundary moment sequence reflects the key action sub-nodes in the process of user interaction with the smart lock.

7. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Video event identification method based on audio-visual mode fusion

    CN116797976A