Lightweight Real-Time Action Recognition Method Based on State Decomposition and State Machine
By disassembling the action into atomic state and combining deep learning and vector retrieval technology, using a state machine to make exception judgments, the problems of low accuracy and efficiency of action recognition in the prior art are solved, and high-precision action recognition on resource-constrained devices are achieved.
Patent Information
- Application Number
- CN202510233515.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The existing action recognition methods have significant shortcomings in recognition accuracy and recognition efficiency, especially in resource-constrained devices, and it is difficult to accurately capture the state relationship between targets in complex scenarios.
A lightweight real-time action recognition method based on state decomposition and state machine is adopted. By disassembling the action into an atomic state, combining dynamic detection of deep learning and static detection of vector retrieval, the motion state machine is used to perform abnormal judgment and early warning signal transmission.
It improves the accuracy and reliability of action recognition, can operate stably on devices with resource-constrained resources, and significantly improves the accuracy of action recognition and the stability of algorithms in complex scenarios.
Smart Images

Figure CN119723680B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of action detection, and relates to deep learning and state machine technologies. Specifically, it is a lightweight real-time action recognition method based on state decomposition and state machines. Background Art
[0002] At present, with the acceleration of the intelligent and automated processes, the importance of action recognition technology has become increasingly prominent in many fields such as production line monitoring, smart home, and human-computer interaction. It can achieve real-time analysis and automatic response in complex dynamic scenarios, and is one of the core elements to improve efficiency and intelligence level. However, the existing action recognition methods currently have significant deficiencies in terms of recognition accuracy and efficiency.
[0003] The existing video sequence action recognition methods based on deep learning highly rely on large-scale labeled datasets, have complex network architectures, and high computational complexity. In edge devices or resource-constrained scenarios, such as some old devices on industrial production lines or low-power smart home sensor terminals, it is difficult to effectively operate due to lack of computing resources, severely restricting the recognition efficiency. At the same time, these methods focus on overall action classification and lack effective means to analyze action details and intermediate states, resulting in difficulty in accurately capturing the state relationships between targets in complex scenarios and a significant reduction in recognition accuracy. For example, in scenarios with multi-target interaction, occlusion, or high action similarity, such methods based on trajectory analysis cannot directly detect object state changes and only rely on indirect trajectory inference, often leading to recognition errors.
[0004] The methods based on key point detection and pose estimation are relatively mature in human action pose analysis, but have weak action recognition capabilities for actions involving object interaction, and are limited in practical application scenarios such as component assembly in manufacturing and cargo handling in the logistics industry. Moreover, they have strict requirements for the accuracy of key point detection. Once key points are lost or misdetected, the overall action recognition effect will be severely damaged, the recognition accuracy will drop sharply, and the recognition efficiency will also decrease due to frequent error handling. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; for this purpose, the present invention proposes a lightweight real-time action recognition method based on state decomposition and state machines to solve the technical problems of low recognition accuracy and poor method stability of existing methods.
[0006] To achieve the above object, the present invention provides a lightweight real-time action recognition method based on state decomposition and state machines, including:
[0007] S1. Collect historical human motion videos, extract key frames from the historical human motion videos and perform preprocessing to obtain a number of historical key frames, and extract continuous action segments from the number of historical key frames in chronological order to obtain an action sequence;
[0008] S2. Decompose each action in the action sequence into atomic states to obtain an atomic state information library;
[0009] S3. Build an object detection model based on a deep learning algorithm, and use the object detection model to perform dynamic detection on the real-time motion video to obtain a dynamic region detection result;
[0010] S4. Build a vector retrieval library based on a number of historical key frames, and use the vector retrieval algorithm to perform static detection on the real-time motion video according to the vector retrieval library to obtain a static region detection result;
[0011] S5. Use a motion state machine to perform anomaly judgment on the dynamic region detection result and the static region detection result, and send a warning signal according to the anomaly judgment result; wherein,
[0012] The motion state machine uses its preset state transition rule template and the number of atomic states determined according to the atomic state information library to match the dynamic region detection result and the static region detection result with the atomic state information library to determine the atomic state of the current key frame in the real-time motion video; according to the frame buffer and the time threshold mechanism, monitor the duration and sequence change of the atomic state;
[0013] The frame buffer is used to store the atomic states of consecutive key frames with consistent states before and after. When the size of the frame buffer is equal to the set frame buffer size threshold, pop the atomic state in the frame buffer and compare it with the last atomic state in the sequence to be detected; the time threshold mechanism is used to increase the value of the time counter when the comparison result is consistent, otherwise add the atomic state popped from the frame buffer to the sequence to be detected and reset the time counter to 0; if the value of the time counter reaches the preset time threshold, or the sequence to be detected is inconsistent with the predefined state transition rule module in the motion state machine, it is determined as an anomaly.
[0014] Based on the above technical process, the present invention provides a lightweight real-time action recognition method by combining state decomposition and a state machine, reducing the complexity and resource consumption of the system, which is beneficial to running on resource-constrained devices; and combining dynamic detection based on deep learning and static detection based on vector retrieval to analyze human motion videos from different dimensions, improving the accuracy and reliability of action recognition, and being able to more accurately judge anomalies and send warning signals in a timely manner.
[0015] Further, the construction method of the atomic state information library includes:
[0016] S2-1. Divide the detection regions in the historical human motion video into dynamic regions and static regions according to whether there are displacement changes in several objects in the action sequence.
[0017] S2-2. Constrain the dynamic regions and static regions through a semantic rule base to obtain logical conditions of several atomic states; wherein, an atomic state represents the smallest state unit of an action and is used to reflect the position state and interaction situation between an object and a human body part during the action process.
[0018] S2-3. Integrate the logical conditions of several atomic states through database technology to obtain an atomic state information library.
[0019] Furthermore, the stored information of the dynamic region and the static region includes the region name, the upper left coordinate of the region, and the lower right coordinate of the region.
[0020] Human motion actions are relatively complex. Decomposing the action sequence into atomic states can analyze the action at the smallest state unit, simplifying the understanding and processing difficulty of the action. At the same time, the atomic state can reflect the position state and interaction situation between an object and a human body part during the action process, which helps to understand and recognize the action more carefully and accurately.
[0021] Furthermore, the construction of the object detection model based on the deep learning algorithm includes:
[0022] S3-1. Use annotation software to perform class annotation on the dynamic regions in several historical key frames to obtain an annotation data set; wherein, the annotation content includes the class and position information of the object.
[0023] S3-2. One-to-one correspond the annotation data set with several key frames to obtain an object detection data set, and divide the object detection data set into a training set, a validation set, and a test set according to a preset ratio.
[0024] S3-3. Build an object detection algorithm framework based on the deep learning algorithm and set the training parameters of the object detection algorithm framework.
[0025] S3-4. Input the training set and the validation set into the object detection algorithm framework for iterative training and evaluation optimization to obtain the model with the highest validation accuracy.
[0026] S3-5. Input the test set into the model with the highest validation accuracy for testing to obtain the test accuracy, and determine whether the test accuracy is greater than the preset test accuracy; if yes, mark the model with the highest validation accuracy as the object detection model; if not, adjust the training parameters of the object detection algorithm framework and jump to S3-4.
[0027] Further, constructing a vector retrieval library according to a plurality of historical key frames includes:
[0028] S41-1, cropping out static regions in a plurality of historical key frames to obtain a static image set, and using annotation software to perform category annotation on a plurality of objects in the static image set to obtain a static annotation set;
[0029] S41-2, grouping the static annotation set according to the categories of several atomic states in the semantic rules to obtain a vector retrieval data set;
[0030] S41-3, using a vector retrieval model to extract features from the static image set according to the vector retrieval data set to obtain a static region feature set ; where N represents the number of static regions in a plurality of historical key frames, represents the feature vector of the static region j in a plurality of historical key frames; among them, the vector retrieval model is a publicly available image detection model and does not require additional training and construction;
[0031] S41-4, making the static region feature set and the vector retrieval data set correspond one by one to obtain a vector retrieval library.
[0032] Further, the vector retrieval algorithm includes:
[0033] S42-1, using an annotation tool to mark a plurality of static regions in a real-time motion video to obtain a real-time static region image set;
[0034] S42-2, using a vector retrieval model to extract features from the real-time static region image set to obtain a real-time static feature set ; where M represents the number of static regions in a plurality of key frames, represents the feature vector of the static region i;
[0035] S42-3, according to the formula calculate and several in the static region feature set similarity , to obtain several similarities;
[0036] S42-4, screening out the corresponding to the highest similarity value among several similarities, to obtain , and extracting from the vector retrieval library grouping, and giving grouping to , to obtain the atomic state category of the i-th static region in the real-time motion video.
[0037] Group the static annotation set according to the categories of atomic states in the semantic rules, so that the construction of the vector retrieval library has semantic logic. During actual detection, through the vector retrieval algorithm, the atomic state category of the static region can be quickly and accurately determined, the information contained in the static part during the action process can be mined, and combined with the dynamic detection results, the action can be more comprehensively characterized, thereby more accurately identifying the action, reducing misjudgment and missed judgment situations, and thus improving the overall action recognition accuracy.
[0038] S5. Use the motion state machine to perform anomaly judgment on the dynamic region detection result and the static region detection result, and send a warning signal according to the anomaly judgment result.
[0039] Furthermore, the construction process of the motion state machine includes:
[0040] S51-1. Determine the number of atomic states and the state transition rule template according to the atomic state information library, and set the frame buffer size threshold and time threshold to obtain state parameters.
[0041] S51-2. Construct a state machine, transmit the state parameters to the state machine, and initialize the time counter to obtain the motion state machine; wherein, the time counter is an empty atomic state sequence with an initial value of 0 and a maximum length equal to the number of atomic states.
[0042] Furthermore, the use of the motion state machine to perform anomaly judgment on the dynamic region detection result and the static region detection result includes:
[0043] S52-1. Match the dynamic region detection result and the static region detection result with the atomic state information library to determine whether there is a consistent result; if yes, obtain the atomic state S of the current key frame in the real-time motion video i ; if not, mark the atomic state of the current key frame in the real-time motion video as None and jump to S3;
[0044] S52-2. Determine whether the frame buffer is empty; if yes, add the atomic state S of the current key frame i to the frame buffer; if not, jump to S52-3;
[0045] S52-3. Compare whether the frame buffer size is equal to the preset frame buffer size threshold; if yes, pop the atomic state in the frame buffer and mark it as S k , and clear the buffer, then jump to S52-2; if not, jump to S52-4;
[0046] S52-4. Determine whether all the atomic states stored in the frame buffer are S i ; if yes, the atomic state S of the current key frame iAdd to the frame buffer; if not, clear the frame buffer and set the atomic state S of the current key frame i Add to the frame buffer;
[0047] S52-5. Perform anomaly detection according to the sequence to be detected of the motion state machine; wherein, the sequence to be detected is used to store S k .
[0048] The state machine integrates the dynamic region detection results and the static region detection results, and matches them with the atomic state information library, making full use of information in different dimensions in the video, comprehensively analyzing each key frame of the action, so as to more accurately determine the atomic state of the current key frame. And the frame buffer stores and processes the atomic state according to the preset threshold and conditions, ensuring that the system can continuously and accurately perform anomaly detection when facing continuously changing video frames, avoiding accidental detection errors from contaminating the sequence to be detected and causing unnecessary false detections, and enhancing the reliability of the system in different action scenarios.
[0049] Further, the anomaly detection according to the sequence to be detected of the motion state machine includes:
[0050] S52-51. Judge whether the sequence to be detected is empty; if so, add S k to the sequence to be detected; if not, jump to S52-52;
[0051] S52-52. Compare S k with the last state in the sequence to be detected; if they are the same, increment the time counter by 1 and jump to S52-53; if not, add S k to the sequence to be detected and reset the time counter to 0;
[0052] S52-53. Judge whether the time counter reaches the preset time threshold; if so, mark it as an anomaly and send an alarm signal; if not, jump to S52-54;
[0053] S52-54. Compare whether the sequence to be detected is the same as that in the state transition rule template; if so, jump to S52-1; if not, mark it as an anomaly and send an alarm signal.
[0054] Using the time counter to monitor the state duration, through the preset time threshold and comparison with the time counter, such problems caused by abnormal state duration can be accurately captured, improving the accuracy of anomaly detection.
[0055] Further, the target detection model and the vector retrieval algorithm are lightweighted through model pruning and INT8 quantization technology, and are deployed on the intelligent edge device using a streaming processing architecture, and real-time action recognition is performed in an asynchronous execution manner.
[0056] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0057] By decomposing actions into atomic states and combining lightweight object detection and vector retrieval technologies, the present invention can accurately identify the interaction states of dynamic objects and the states of static regions, comprehensively extract information from single-frame images, and effectively improve the action recognition accuracy in complex scenarios;
[0058] Introducing an action state machine with a frame buffer and a time threshold mechanism, the frame buffer can store and integrate the recognition results in a short time, reducing the impact of occasional misidentifications to achieve jitter elimination; the time threshold mechanism can alarm in time when the expected action sequence is interrupted, avoiding deadlock waiting, and significantly improving the stability and reliability of the algorithm operation;
[0059] Aiming at the problem of limited computing power of edge devices, model pruning and INT8 quantization technologies are adopted to reduce resource occupancy, and a streaming processing architecture and asynchronous execution mode are used to optimize data flow and module cooperation, enabling the algorithm to run stably at high frame rates on intelligent edge devices and enhancing its applicability in actual scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0061] Figure 1 It is a schematic diagram of the technical process of the lightweight real-time action recognition method based on state decomposition and state machine provided by the present invention;
[0062] Figure 2 It is a schematic diagram of the atomic state information library provided by the present invention;
[0063] Figure 3 It is the construction process of the vector retrieval library provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The following will clearly and completely describe the technical solutions of the present invention in combination with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0065] Please refer to Figures 1 - 3 , the first aspect embodiment of the present invention provides a lightweight real-time action recognition method based on state decomposition and state machine, including:
[0066] S1, Collect historical human motion videos, extract key frames from the historical human motion videos and perform preprocessing to obtain a number of historical key frames, and extract continuous action segments from the number of historical key frames in chronological order to obtain an action sequence;
[0067] S2, Decompose each action in the action sequence into atomic states to obtain an atomic state information library;
[0068] S3, Build an object detection model based on a deep learning algorithm, and use the object detection model to perform dynamic detection on the real-time motion video to obtain a dynamic region detection result;
[0069] S4, Build a vector retrieval library according to a number of historical key frames, and perform static detection on the real-time motion video using a vector retrieval algorithm based on the vector retrieval library to obtain a static region detection result;
[0070] S5, Use a motion state machine to perform anomaly judgment on the dynamic region detection result and the static region detection result, and send a warning signal according to the anomaly judgment result; where,
[0071] The motion state machine uses its preset state transition rule template and the number of atomic states determined according to the atomic state information library to match the dynamic region detection result and the static region detection result with the atomic state information library to determine the atomic state of the current key frame in the real-time motion video; According to the frame buffer and the time threshold mechanism, monitor the duration and sequence change of the atomic state;
[0072] The frame buffer is used to store the atomic states of consecutive key frames with consistent states before and after. When the size of the frame buffer is equal to the set frame buffer size threshold, the atomic state in the frame buffer is popped, and it is compared with the last atomic state in the sequence to be detected; The time threshold mechanism is used to increase the value of the time counter when the comparison result is consistent, otherwise the atomic state popped from the frame buffer is added to the sequence to be detected, and the time counter is reset to 0; If the value of the time counter reaches the preset time threshold, or the sequence to be detected is inconsistent with the predefined state transition rule module in the motion state machine, it is determined as an anomaly.
[0073] In the field of object detection, existing action recognition methods are vulnerable to target occlusion, action similarity, and background interference in complex scenarios, resulting in low recognition accuracy. And currently, many action recognition algorithms are prone to problems such as recognition jitter or system deadlock waiting when dealing with short-term misrecognition, action interruption, or long-term non-triggering of expected actions.
[0074] Therefore, by decomposing actions into atomic states such as target interaction states and regional states, and combining lightweight object detection and vector retrieval technologies, the present invention improves the accuracy of single-frame state extraction and enhances the action recognition ability in complex scenarios; by introducing an action state machine with a frame buffer and a time threshold mechanism, the stability of the recognition results is improved, and at the same time, the algorithm is prevented from falling into a long-term unresponsive state in abnormal situations.
[0075] The lightweight real-time action recognition method based on state decomposition and state machine provided in this embodiment includes the following steps:
[0076] Step S1, collect historical human motion videos, extract key frames from the historical human motion videos and perform preprocessing to obtain a number of historical key frames, and extract continuous action segments from the number of historical key frames in chronological order to obtain an action sequence.
[0077] Specifically, first, collect historical human motion videos in real time, and then extract key frames from the videos, which can be extracted based on time intervals or frame differences. Then, preprocess a number of historical key frames, such as using methods such as image enhancement, size normalization, and color space conversion; finally, arrange the preprocessed key frames in chronological order, and extract continuous action segments according to time relationships and action coherence; and use timestamp information to group adjacent key frames according to a certain time window to obtain an action sequence.
[0078] Step S2, decompose each action in the action sequence into atomic states to obtain an atomic state information library.
[0079] Specifically, as Figure 2 shown, in this embodiment, the recognition region is mainly divided based on the decomposition of scene information, and then constrained by semantic rules to decompose complex actions into atomic states, and an atomic state information library containing region information and atomic state information is constructed.
[0080] The decomposition based on scene information maps actions to the combined changes of dynamic regions (such as objects, human body parts) and static regions (such as containers, ground states) in the scene. The information that needs to be stored for each region includes the region name and the coordinates of the upper left and lower right corners of the region. Among them, the static regions need to be initialized before the algorithm runs.
[0081] Based on the decomposition of scene information, constraints are imposed through semantic rules to define the logical conditions for each atomic state. Each action may contain multiple atomic states, and the atomic state information includes the algorithm type for state detection and specific rules. For example, for the atomic state "hand in contact with the device", the algorithm type for state detection is "Detection", and the specific rule is that there is an overlapping area between the "hand" and the "device", which means that the "grasping" action is completed; for the atomic state "device in the specified area", the algorithm type for state detection is "Retrieval", and the specific rule is that the "status of the specified area is 'non-empty'", which means that the "placement" action is completed; then, the logical conditions for several atomic states are obtained; among them, the atomic state represents the smallest state unit of an action, which is used to reflect the position state and interaction between objects and human body parts during the action process;
[0082] Finally, the logical conditions of several atomic states are integrated through database technology to obtain an atomic state information library.
[0083] Step S3: Construct an object detection model based on a deep learning algorithm, and use the object detection model to perform dynamic detection on the real-time motion video to obtain the dynamic area detection result.
[0084] Specifically, the construction process of the object detection model in this embodiment is as follows:
[0085] S3-1: Use annotation software to perform category annotation on the dynamic areas in several historical key frames to obtain an annotated data set; among them, the annotation content includes the category and position information of the object;
[0086] S3-2: One-to-one correspond the annotated data set with several key frames to obtain an object detection data set, and divide the object detection data set into a training set, a validation set, and a test set according to a preset ratio;
[0087] S3-3: Construct an object detection algorithm framework based on a deep learning algorithm and set the training parameters of the object detection algorithm framework;
[0088] S3-4: Input the training set and the validation set into the object detection algorithm framework for iterative training and evaluation optimization to obtain the model with the highest validation accuracy;
[0089] S3-5: Input the test set into the model with the highest validation accuracy for testing to obtain the test accuracy, and determine whether the test accuracy is greater than the preset test accuracy; if yes, mark the model with the highest validation accuracy as the object detection model; if not, adjust the training parameters of the object detection algorithm framework and jump to S3-4.
[0090] The present invention realizes accurate action recognition by disassembling a motion video into a dynamic area and a static area and then detecting them separately. Among them, dynamic detection is to perform real-time detection on the dynamic areas (such as objects and human body parts) in a human motion video through an object detection model constructed based on a deep learning algorithm. By collecting historical motion videos, extracting key frames and annotating them, a high-precision object detection model is constructed, which can obtain information such as the category and position of the dynamic area in real time, such as real-time monitoring of dynamic changes such as the contact state between the hand and the device, so as to capture the dynamic interaction information between objects and human body parts during the action process; the static area (such as the container and the ground state) is initialized before the algorithm runs and complements the dynamic area. It provides relatively stable scene information for action recognition, such as the state of the designated area where the device is placed, etc., which helps to understand the background and environment where the action occurs and assist in determining whether the action is completed. The combination of the two comprehensively captures action information from two dimensions of dynamic changes and static background, making action recognition more accurate and complete.
[0091] Step S4, construct a vector retrieval library according to a number of historical key frames, and use the vector retrieval algorithm to perform static detection on the real-time motion video according to the vector retrieval library to obtain a static area detection result.
[0092] Specifically, in this embodiment, a number of historical key frames are used to construct a vector retrieval library, including:
[0093] First, cut out the static areas in a number of historical key frames to obtain a static image set, and use annotation software to perform category annotation on a number of objects in the static image set to obtain a static annotation set;
[0094] Then, group the static annotation set according to the categories of a number of atomic states in the semantic rules to obtain a vector retrieval data set;
[0095] Next, use the vector retrieval model to extract features from the static image set according to the vector retrieval data set to obtain a static area feature set ; where N represents the number of static areas in a number of historical key frames, represents the feature vector of the static area j in a number of historical key frames;
[0096] Finally, make the static area feature set correspond one by one with the vector retrieval data set to obtain a vector retrieval library.
[0097] Construct a vector retrieval algorithm based on the vector retrieval data set and the vector retrieval model. The specific construction and reasoning steps of the vector retrieval algorithm are:
[0098] S42-1, use an annotation tool to annotate a number of static areas in the real-time motion video to obtain a real-time static area image set;
[0099] S42-2. Extract features from the real-time static region image set using a vector retrieval model to obtain a real-time static feature set. Among them, M represents the number of static regions in several key frames. represents the feature vector of static region i.
[0100] S42-3. Calculate according to the formula Calculate and the similarity with several in the static region feature set , to obtain several similarities.
[0101] S42-4. Select the corresponding to the highest similarity value among several similarities, to obtain , and extract the grouping from the vector retrieval library, and assign the grouping to , to obtain the atomic state category of the i-th static region in the real-time motion video.
[0102] S5. Use a motion state machine to perform anomaly judgment on the dynamic region detection result and the static region detection result, and send a warning signal according to the anomaly judgment result.
[0103] Specifically, first initialize the state machine:
[0104] Determine the number of atomic states and the state transition rule template according to the atomic state information library, and set the frame buffer size threshold and time threshold to obtain state parameters.
[0105] Transmit the state parameters to the state machine and initialize the time counter to obtain a motion state machine; among them, the time counter is an empty atomic state sequence with an initial value of 0 and a maximum length of the number of atomic states.
[0106] Then use the motion state machine to perform anomaly judgment on the dynamic region detection result and the static region detection result, including:
[0107] First, obtain the dynamic region detection result and the static region detection result of the current key frame, and compare them with the atomic state information library to determine whether there is a consistent result; if so, obtain the atomic state S i of the current frame; if not, set the atomic state of the current frame to None and jump to S3 for re-detection.
[0108] Next, when the frame buffer size is less than the set frame buffer size threshold, process the atomic state S i based on the frame buffer state:
[0109] Specifically, if the frame buffer is empty, directly add S i to the frame buffer; if the frame buffer is not empty and all the atomic states stored in the frame buffer are S i , then add S i to the frame buffer; if the frame buffer is not empty but the atomic states stored in the frame buffer are not S i , then clear the frame buffer and add S i ;
[0110] When the size of the frame buffer is equal to the set frame buffer size threshold, pop the atomic state in the frame buffer and record it as S k , and clear the buffer;
[0111] Then construct a sequence to be detected for storing S k ; Before adding S k to the sequence to be detected, first determine whether the sequence to be detected is empty; if it is, directly add S K to the sequence to be detected; if not, compare S k with the last state in the sequence to be detected; if they are the same, increment the time counter by 1, and if the time counter reaches the set time threshold at this time, send an alarm signal; if not, add S k to the sequence to be detected and reset the time counter to 0;
[0112] Finally, it is also necessary to compare whether the sequence to be detected is consistent with the state transition rule template; if it is, obtain the recognition result of the next frame and repeat the above steps; if not, mark it as abnormal and send an alarm signal.
[0113] The state machine uses its preset state transition rule template and the number of atomic states determined according to the atomic state information library to match the dynamic region detection result and the static region detection result with the atomic state information library, so as to accurately determine the atomic state of the current key frame in the real-time motion video. At the same time, by setting the frame buffer size threshold and the time threshold, and cooperating with the time counter, the state machine can monitor the duration and sequence change of the atomic state. When the last state in the sequence to be detected is inconsistent with a specific state, the sequence to be detected can be updated in time; if the time counter reaches the preset time threshold, or the sequence to be detected is inconsistent with the state transition rule template, it can be quickly marked as abnormal and an alarm signal can be sent, realizing the comprehensive monitoring of the action in the time dimension and the state sequence dimension, greatly improving the accuracy and timeliness of the alarm judgment, effectively avoiding misjudgment and missed judgment, and being able to quickly detect abnormal actions in the real-time action recognition process, providing strong support for taking timely measures.
[0114] Some of the data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is obtained by software simulation of a large amount of collected data to get a formula that is closest to the actual situation. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.
[0115] The working principle of the present invention:
[0116] The present invention proposes a real-time action recognition algorithm. By combining a lightweight object detection algorithm, it accurately identifies the interaction state of dynamic objects. At the same time, it uses a vector retrieval algorithm to efficiently judge the state of static regions, and comprehensively extracts the object and region state information from a single-frame image. On this basis, through preset action rules, it realizes the accurate parsing of action events in a video sequence;
[0117] The present invention designs an action state machine with a frame buffer and a time threshold. The frame buffer significantly reduces the impact of occasional misidentifications on the system performance through short-term storage and integration of the recognition results, and realizes the jitter optimization of the recognition results. The time threshold mechanism can trigger an alarm when the expected action sequence is interrupted for a long time, avoiding the system falling into a deadlock waiting state, thus significantly improving the running stability and reliability of the algorithm;
[0118] Through the lightweight optimization of the algorithm, the present invention effectively reduces the occupancy of computing resources and energy consumption, enabling the algorithm to stably run at a high frame rate in an environment with limited computing power such as intelligent edge devices, and significantly improving its applicability in actual scenarios.
[0119] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A lightweight real-time action recognition method based on state decomposition and state machine, characterized in that: include: S1, collecting historical human motion videos, extracting and preprocessing key frames of the historical human motion videos to obtain a number of historical key frames, and extracting continuous action segments from the several historical key frames in chronological order to obtain an action sequence; S2-1, dividing the detection area in the historical human motion video into a dynamic area and a static area according to whether a number of objects in the action sequence have displacement changes; S2-2, constraining the dynamic area and the static area through the semantic rule base to obtain the logical conditions of several atomic states; wherein the atomic state represents the smallest state unit of the action, which is used to reflect the position state and interaction between the object and the human body parts during the action; S2-3, integrating the logical conditions of several atomic states through database technology to obtain an atomic state information library; S3, builds a target detection model based on a deep learning algorithm, uses the target detection model to perform dynamic detection on real-time motion video, and obtains dynamic area detection results; S4, constructing a vector retrieval library according to a number of historical key frames, and performing static detection on the real-time motion video using a vector retrieval algorithm according to the vector retrieval library to obtain a static area detection result; S5, using the motion state machine to make abnormal judgments on the dynamic area detection results and the static area detection results, and sending a warning signal according to the abnormal judgment results; wherein, The motion state machine uses its preset state transition rule template and the number of atomic states determined according to the atomic state information library to match the dynamic area detection result and the static area detection result with the atomic state information library to determine the atomic state of the current key frame in the real-time motion video; monitors the duration and sequence changes of the atomic state according to the frame buffer and time threshold mechanism; The frame buffer is used to store the atomic states of consecutive key frames with consistent states. When the frame buffer size is equal to the set frame buffer size threshold, the atomic state in the frame buffer is popped up and compared with the last atomic state in the sequence to be detected; the time threshold mechanism is used to increase the value of the time counter when the comparison result is consistent, otherwise the atomic state in the popped-up frame buffer is added to the sequence to be detected, and the time counter is reset to 0; if the value of the time counter reaches the preset time threshold, or the sequence to be detected is inconsistent with the state transition rule module predefined in the motion state machine, it is determined to be abnormal.
2. The lightweight real-time action recognition method based on state decomposition and state machine according to claim 1 is characterized in that: The storage information of the dynamic area and the static area includes the area name, the coordinates of the upper left corner of the area, and the coordinates of the lower right corner of the area.
3. The lightweight real-time action recognition method based on state decomposition and state machine according to claim 1 is characterized in that: The target detection model is constructed based on the deep learning algorithm, including: S3-1, using annotation software to classify the dynamic areas in several historical key frames to obtain an annotation data set; wherein the annotation content includes the category and location information of the object; S3-2, one-to-one correspondence between the labeled data set and a number of key frames to obtain a target detection data set, and the target detection data set is divided into a training set, a validation set, and a test set according to a preset ratio; S3-3, build a target detection algorithm framework based on the deep learning algorithm, and set the training parameters of the target detection algorithm framework; S3-4, input the training set and the validation set into the target detection algorithm framework for iterative training and evaluation optimization to obtain the model with the highest validation accuracy; S3-5, input the test set into the model with the highest verification accuracy for testing, obtain the test accuracy, and determine whether the test accuracy is greater than the preset test accuracy; if yes, mark the model with the highest verification accuracy as the target detection model; if not, adjust the training parameters of the target detection algorithm framework and jump to S3-4.
4. The lightweight real-time action recognition method based on state decomposition and state machine according to claim 1 is characterized in that: The vector retrieval library is constructed according to a number of historical key frames, including: S41-1, cutting out static areas in a number of historical key frames to obtain a static image set, and using annotation software to perform category annotation on a number of objects in the static image set to obtain a static annotation set; S41-2, grouping the static annotation set according to the categories of several atomic states in the semantic rules to obtain a vector retrieval data set; S41-3, using the vector retrieval model to extract features from the static image set according to the vector retrieval data set, to obtain a static area feature set ; Where N represents the number of static regions in several historical key frames, Represents the feature vector of the static region j in several historical key frames; wherein the vector retrieval model is a public image detection model; S41-4, make one-to-one correspondence between the static area feature set and the vector retrieval data set to obtain a vector retrieval library.
5. The lightweight real-time action recognition method based on state decomposition and state machine according to claim 4 is characterized in that: The vector retrieval algorithm comprises: S42-1, using a labeling tool to label several static areas in the real-time motion video to obtain a real-time static area image set; S42-2, using the vector retrieval model to extract features from the real-time static area image set, to obtain a real-time static feature set ; Where M represents the number of static regions in a number of key frames, The feature vector representing the static region i; S42-3, according to the formula calculate Several static area feature sets Similarity , and obtain a number of similarities; S42-4, select the one with the highest similarity value among several similarities ,get , and extract it from the vector retrieval library The grouping will Grouping of , get the atomic state category of the i-th static area in the real-time motion video.
6. The lightweight real-time action recognition method based on state decomposition and state machine according to claim 1 is characterized in that: The construction process of the motion state machine includes: S51-1, determining the number of atomic states and the state transition rule template according to the atomic state information library, and setting a frame buffer size threshold and a time threshold to obtain state parameters; S51-2, construct a state machine, transfer state parameters to the state machine, and initialize the time counter to obtain a motion state machine; wherein the time counter is an empty atomic state sequence with an initial value of 0, and the maximum length is the number of atomic states.
7. The lightweight real-time action recognition method based on state decomposition and state machine according to claim 1 is characterized in that: The method of using the motion state machine to perform abnormality judgment on the dynamic area detection result and the static area detection result includes: S52-1, obtain the dynamic region detection result and the static region detection result of the current key frame, and match them with the atomic state information library to determine whether there is a consistent result; if yes, obtain the atomic state S of the current key frame in the real-time motion video. i ; If not, mark the atomic state of the current key frame in the real-time motion video as None and jump to S3; S52-2, determine whether the frame buffer is empty; if so, set the atomic state S of the current key frame i Add to the frame buffer; if not, jump to S52-3; S52-3, compare whether the frame buffer size is equal to the preset frame buffer size threshold; if yes, pop up the atomic state in the frame buffer and mark it as S k , and clear the buffer, jump to S52-2; otherwise, jump to S52-4; S52-4, determine whether the atomic states stored in the frame buffer are all S i ; If yes, then the atomic state S of the current key frame i Add to the frame buffer; otherwise, clear the frame buffer and set the atomic state S of the current key frame. i Add frame buffer; S52-5, making an abnormality judgment based on the sequence to be detected of the motion state machine.
8. The lightweight real-time action recognition method based on state decomposition and state machine according to claim 7 is characterized in that: The abnormality judgment according to the sequence to be detected of the motion state machine includes: S52-51, determine whether the sequence to be detected is empty; if so, set S k Add to the sequence to be detected; if not, jump to S52-52; S52-52, compare S k Is it consistent with the last state in the sequence to be detected? If yes, the time counter is +1 and jumps to S52-53; if no, S k Add the sequence to be detected and reset the time counter to 0; S52-53, determine whether the time counter reaches the preset time threshold; if yes, mark it as abnormal and send an alarm signal; if no, jump to S52-54; S52-54, compare whether the sequence to be detected is consistent with the state transition rule template; if yes, jump to S52-1; if not, mark it as abnormal and send an alarm signal.
9. The lightweight real-time action recognition method based on state decomposition and state machine according to claim 1, characterized in that: The target detection model and the vector retrieval algorithm are lightweight through model pruning and INT8 quantization technology, and are deployed on intelligent edge devices using a streaming processing architecture, and real-time action recognition is performed through asynchronous execution.
Citation Information
Patent Citations
ICU patient body motion video identification method based on human skeleton key point detection
CN118781654A
Interacting with hierarchical clusters of video segments using a video timeline
US20220075513A1