Office scene behavior recognition method based on behavior key points
By introducing a behavior key point recognition method in office scenarios, combined with Swin-Transformer and SEAttention, the problem of missing timing features in single-frame behavior detection is solved, and high-precision behavior recognition and real-time multi-objective tracking is realized, which is suitable for fast response application scenarios.
Patent Information
- Application Number
- CN202510435494.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
AI Technical Summary
In the behavior analysis of office scenes, single-frame behavior detection lacks timing characteristics, resulting in a reduced accuracy of the model's behavior analysis and it is difficult to extract the posture and action characteristics of the behavior targets in detail, especially when clothing changes, the effect is not good.
Using a behavioral key point recognition method, the recognition system performs feature extraction in a single video frame, combined with Swin-Transformer and SEAttention structures, multi-scale feature extraction and behavioral target tracking are carried out, key point information is introduced to improve detection accuracy, and model performance is improved through module pre-training and overall tuning.
It effectively improves the accuracy and real-time performance of behavior recognition, can accurately track multiple targets in complex scenarios, reduce errors and handle noise, and is suitable for fast response application scenarios, improving the detection ability of targets at different scales.
Smart Images

Figure CN120356262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of recognition methods, and specifically to a behavior recognition method for office scenarios based on behavior key points. Background Art
[0002] The behavior analysis method for office scenarios is divided into detecting the behaviors of multiple targets in each frame of the video stream, and only showing the spatial positions of the behavior targets in the visual sense and the possible behavior categories in the current time-series frame of the video stream.
[0003] In the prior art, the behavior analysis in office scenarios mainly detects abnormal behaviors frame by frame for the behaviors of time-series vision. Since the detected behaviors are mainly one possible behavior category at a moment, and in actual business scenarios, the behavior actions are continuous and complex. Relying only on the body posture in a single frame, there is often a lack of information, lacking its time-series characteristics. In addition, using only the characteristics of the spatial region of the target object in a single frame as the category of the detection target, the extraction of the characteristics of the accessories of the posture and actions of the target object is relatively not specific enough. In actual office scenarios, due to the change of clothing with seasons, it will increase the reduction of the accuracy of behavior analysis of the model. Therefore, a behavior recognition method for office scenarios based on behavior key points is proposed. Summary of the Invention
[0004] The purpose of the present invention is to provide a behavior recognition method for office scenarios based on behavior key points to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A behavior recognition method for office scenarios based on behavior key points, including the following steps:
[0006] S1. Through the recognition system, feature extraction is performed in a single video frame;
[0007] S2. After detecting the behaviors of multiple targets in a single frame of the video stream, identify and locate the behavior objects in the image;
[0008] S3. For the obtained trajectory of the behavior target object, intercept the behavior target object frame by frame from the video frame;
[0009] S4. Locate and identify and classify the features at each moment of the time-series vision.
[0010] Preferably, in the above S1, the recognition system includes a behavior detection module, a behavior tracking module, a video generation module, and a behavior recognition module;
[0011] The behavior detection module is connected to the behavior tracking module, the behavior tracking module is connected to the video generation module, and the video generation module is connected to the behavior recognition module;
[0012] The described behavior detection module is used for feature extraction of a single video frame;
[0013] The described behavior tracking module is used for identifying and locating the behavior objects in the image after single-frame multi-object behavior detection in the video stream;
[0014] The described video generation module is used for intercepting the behavior target object frame by frame in the video frame;
[0015] The described behavior recognition module is used for locating and classifying the behaviors in the video.
[0016] Preferably, the above-mentioned behavior detection module includes a Backone unit, a Neck unit and a Head unit;
[0017] The Backone unit is connected to the Neck unit, and the Neck unit is connected to the Head unit;
[0018] The Backone unit is used to solve the feature extraction of small samples, and introduces Swin-Transformer, and performs multi-scale feature extraction through PatchMerging in the way of feature fusion;
[0019] The Neck unit is used to introduce the SEAttention structure to enhance the model expression ability;
[0020] The Head unit is used to add a detection head for small scales, and output the category of behavior detection, the position information of the target object in the visual space and the key point space information;
[0021] The PatchMerging operation in the Swin-Transformer uses a mechanism similar to the pooling operation in the convolutional network for downsampling;
[0022] The SEAttention realizes the modeling and optimization of the dependence between feature channels through two key operations of squeeze and excitation.
[0023] Preferably, the above-mentioned key point space information includes key points in the facial area, key points in the left area of the upper body, key points in the lower body area, and key points in the area of action accessories;
[0024] The key points in the facial area include the left and right ears, the left and right eyes, the tip of the nose, the left and right corners of the mouth, the upper lip, and the lower lip;
[0025] The key points in the upper body area include the left and right shoulders, the left and right elbows, and the left and right wrists;
[0026] The key points in the lower body area include the left and right hips, the left and right knees, and the left and right ankles;
[0027] The key points of the action accessory area include the upper left vertex of the mobile phone, the pen tip, and the pen head.
[0028] Preferably, the above-mentioned behavior tracking module includes a detection box grading unit, a Kalman filter prediction unit, a data association unit, and a trajectory update and management unit;
[0029] The detection box grading unit is connected to the Kalman filter prediction unit, the Kalman filter prediction unit is connected to the data association unit, and the data association unit is connected to the trajectory update and management unit.
[0030] Preferably, the above-mentioned detection box grading unit is used to grade the detection boxes in each frame of the image;
[0031] The Kalman filter prediction unit is used to predict the candidate boxes to estimate the position and movement trajectory of the behavior target object in the next frame;
[0032] The data association unit is used to match the detection boxes in the current frame with the trajectories in the previous frames to establish or update the trajectory of the target object;
[0033] The trajectory update and management unit is used to update and manage the trajectory states of all behavior target objects in each frame of the algorithm;
[0034] The data association unit adopts BYTE data association.
[0035] Preferably, the above-mentioned video generation module includes a behavior trajectory acquisition unit, a single-target behavior cropping unit, a behavior strategy determination unit, and a video production unit;
[0036] The behavior trajectory acquisition unit is connected to the single-target behavior cropping unit, the single-target behavior cropping unit is connected to the behavior strategy determination unit, and the behavior strategy determination unit is connected to the video production unit.
[0037] Preferably, the above-mentioned behavior trajectory acquisition unit is used to obtain the trajectory information of each behavior target object from the behavior tracking module;
[0038] The single-target behavior cropping unit is used to intercept the behavior segments of the target object frame by frame from the video frames;
[0039] The behavior strategy determination unit is used to judge whether the intercepted behavior segments conform to the preset behavior strategies;
[0040] The video production unit is used to synthesize the behavior segments that conform to the preset behavior strategies into a complete video file.
[0041] Preferably, the above-mentioned behavior recognition module includes a video frame feature extraction unit, a target object temporal feature unit, and a temporal behavior recognition output unit;
[0042] The video frame feature extraction unit is connected to the target object temporal feature unit, and the target object temporal feature unit is connected to the temporal behavior recognition output unit.
[0043] Preferably, the above-mentioned video frame feature extraction unit is used to extract useful visual features from video frames;
[0044] The target object temporal feature unit is used to combine and compress the key point features and behavior detection type features extracted frame by frame in the time dimension to form a more compact and useful feature representation;
[0045] The temporal behavior recognition output unit is used to identify the behavior types in the video by the extracted and combined features.
[0046] Preferably, the above.
[0047] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0048] First, in the model construction process of the present invention, the method of module pre-training and then overall tuning is adopted. The Swin-transformer model is mainly used for feature extraction on visual frames, which acts as a role similar to the convolution operation in CNN in image processing. Due to the characteristics of the Transformer structure, it has advantages in processing global relationships, long-distance information transmission, etc., and can effectively improve the model's abstract ability and expression ability for image content. The SEAttention structure is adopted, and its advantages lie in enhancing feature importance, improving model performance, reducing computational complexity, and at the same time, it can be widely applied to various deep learning tasks, which has important significance and application prospects. Four detection heads are adopted to improve the detection ability for targets of different scales to solve the problems of missed detection or poor detection effect. In addition, when extracting features, the key point information of the behavior target is also introduced to obtain more refined behavior attribute features, further improving the accuracy in detection and behavior recognition.
[0049] Second, through the behavior tracking module, the present invention shows real-time performance and high efficiency in target tracking, is applicable to application scenarios that require quick response, can effectively predict the state of a dynamic system, improve tracking accuracy, support multi-target tracking. At the same time, through the feature fusion and information fusion capabilities, it combines the observation data and prediction results to update the target state, reduce errors and handle noise and uncertainty. Its modeling and update mechanism makes it an ideal choice for dealing with complex tracking scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0051] Figure 1 Schematic diagram of the recognition system of the present invention;
[0052] Figure 2 Schematic diagram of the behavior detection module of the present invention;
[0053] Figure 3 Schematic diagram of the behavior tracking module of the present invention;
[0054] Figure 4 Schematic diagram of the video generation module of the present invention;
[0055] Figure 5 Schematic diagram of the behavior recognition module of the present invention. Detailed implementation manners
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0057] It should be noted that the structures, ratios, sizes, etc. shown in the accompanying drawings of this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions that the present application can be implemented. Therefore, they do not have technical substance significance. Any modification of the structure, change of the proportional relationship, or adjustment of the size should still fall within the scope that can be covered by the technical content disclosed in the present application without affecting the effects that the present application can produce and the purposes that can be achieved.
[0058] Embodiment
[0059] Please refer to Figures 1-5 , the present invention provides a technical solution: a method for recognizing office scene behaviors based on behavior key points, including the following steps:
[0060] S1. Through the recognition system, feature extraction is performed in a single video frame;
[0061] S2. After detecting multi-object behaviors in a single frame of the video stream, the behavior objects in the image are recognized and located;
[0062] S3. Extract the behavior target object frame by frame from the video frames based on the obtained behavior target object trajectory;
[0063] S4. Locate, identify, and classify the features at each moment of the temporal vision.
[0064] The recognition system includes a behavior detection module, a behavior tracking module, a video generation module, and a behavior recognition module; the behavior detection module is connected to the behavior tracking module, the behavior tracking module is connected to the video generation module, and the video generation module is connected to the behavior recognition module;
[0065] The behavior detection module is used to extract features from a single video frame;
[0066] The behavior detection module includes a Backone unit, a Neck unit, and a Head unit;
[0067] The Backone unit is connected to the Neck unit, and the Neck unit is connected to the Head unit; after each feature extraction, downsampling is performed once, increasing the receptive field of the next window attention operation on the original image, thereby performing multi-scale feature extraction on the input image;
[0068] The Backone unit is used to solve the feature extraction of small samples and introduces Swin-Transformer, and performs multi-scale feature extraction through PatchMerging in the way of feature fusion; the Neck unit, an attention mechanism in the convolutional neural network (CNN) for enhancing the model's expression ability, introduces the SEAttention structure to enhance the model's expression ability; the Head unit is used to add small-scale detection heads, output the category of behavior detection, the position information of the target object in the visual space, and the key point space information, and learn or output the behavior of the detected behavior. Not only should there be the category of behavior detection and the position information of the behavior target object in the visual space, but also the key point space information during training or output inference process should be learned simultaneously. Four detection heads are used to improve the detection ability for different scale targets to solve the problems of missed detection or poor detection effect. In addition, during feature extraction, the key point information of the behavior target is introduced to obtain the behavior attribute features more precisely, further improving the accuracy in detection and behavior recognition;
[0069] The PatchMerging operation in Swin-Transformer performs downsampling using a mechanism similar to the pooling operation in convolutional networks; SEAttention realizes the modeling and optimization of the dependencies between feature channels through two key operations, namely squeeze and excitation, enabling the network to pay more attention to those feature channels that are more important for the current task.
[0070] The key point spatial information includes key points in the facial area, key points in the left upper body area, key points in the lower body area, and key points in the action accessory area;
[0071] The key points in the facial area include the left and right ears, left and right eyes, the tip of the nose, left and right corners of the mouth, the upper lip, and the lower lip;
[0072] The key points in the upper body area include the left and right shoulders, left and right elbows, and left and right wrists;
[0073] The key points in the lower body area include the left and right hips, left and right knees, and left and right ankles;
[0074] The key points in the action accessory area include the upper left vertex of the mobile phone, the tip of the pen, and the pen head.
[0075] The behavior tracking module is used to identify and locate the behavior objects in the image after single-frame multi-object behavior detection in the video stream; the behavior tracking module includes a detection box grading unit, a Kalman filter prediction unit, a data association unit, and a trajectory update and management unit;
[0076] The detection box grading unit is connected to the Kalman filter prediction unit, the Kalman filter prediction unit is connected to the data association unit, and the data association unit is connected to the trajectory update and management unit.
[0077] The detection box grading unit is used to grade the detection boxes in each frame of the image;
[0078] The Kalman filter prediction unit is used to predict the candidate boxes to estimate the position and movement trajectory of the behavior target object in the next frame;
[0079] The data association unit is used to match the detection boxes in the current frame with the trajectories in the previous frames to establish or update the trajectories of the target objects;
[0080] The trajectory update and management unit is used to update and manage the trajectory states of all behavior target objects in each frame of the algorithm;
[0081] The data association unit adopts BYTE data association.
[0082] First, perform detection box grading based on the scores of the detection box grading unit; then, for the candidate boxes, use the Kalman filtering prediction unit to make predictions to estimate the position of the behavioral target object in the next frame and the motion trajectory of the target object. Secondly, adopt the BYTE data association method to match the detection boxes in the current frame with the trajectories in the previous frames; finally, the algorithm for each frame updates and manages the trajectory states of all behavioral target objects, and ByteTrack can maintain continuous tracking of each behavioral target object, and can resume tracking even when the target temporarily disappears or is occluded.
[0083] A video generation module, used to intercept behavioral target objects frame by frame in video frames;
[0084] The video generation module includes a behavioral trajectory acquisition unit, a single-target behavior cropping unit, a behavioral strategy determination unit, and a video production unit; the behavioral trajectory acquisition unit is connected to the single-target behavior cropping unit, the single-target behavior cropping unit is connected to the behavioral strategy determination unit, and the behavioral strategy determination unit is connected to the video production unit.
[0085] The behavioral trajectory acquisition unit is used to obtain the trajectory information of each behavioral target object from the behavioral tracking module;
[0086] The single-target behavior cropping unit is used to intercept the behavior segments of the target object frame by frame from the video frames;
[0087] The behavioral strategy determination unit is used to determine whether the intercepted behavior segments conform to the preset behavioral strategies;
[0088] The video production unit is used to synthesize the behavior segments that conform to the preset behavioral strategies into a complete video file.
[0089] Obtain the starting time-domain position from the behavioral trajectory and intercept the behavioral target objects frame by frame from the video frames through the mass points of the behavioral target objects at each moment in the time domain and the width and height at the start of the behavior. Then, when intercepting until the end of the behavior, adopt a certain behavioral target detection fault tolerance strategy to ensure the continuity of the action, and then generate candidate behavioral videos.
[0090] If the behavioral action is inconsistent with the action type at the previous moment, still merge the behavior, but mark the action with a suspicion flag. When the number of suspicion flags of the action exceeds the threshold of the total record of the behavioral action, a video is generated.
[0091] A behavior recognition module, used to locate and classify the behaviors in the video.
[0092] The behavior recognition module includes a video frame feature extraction unit, a target object temporal feature unit, and a temporal behavior recognition output unit; the video frame feature extraction unit is connected to the target object temporal feature unit, and the target object temporal feature unit is connected to the temporal behavior recognition output unit.
[0093] The video frame feature extraction unit is used to extract useful visual features from video frames; the model parameters pre-trained with single-frame and single-person behavior data, and the training tasks are still behavior types, the spatial position area (bbox) of the target object in the image, and the key points of the target object.
[0094] The target object temporal feature unit is used to combine and compress the key point features and behavior detection type features extracted frame by frame in the time dimension to form a more compact and useful feature representation.
[0095] The key point features (in the format of B*T,M,new_W,new_H) and behavior detection type features (in the format of B*T,M,new_W,new_H) extracted frame by frame are separately merged in the time dimension to form features (B,T,M,new_W,new_H), and then feature compression is performed in the (M,new_W,new_H) dimension to construct features (B,T,P). Then, the temporal key points and behavior type features respectively perform temporal information compression on the above features at each moment of temporal vision through the deep learning model Mamba based on the state space model (SSM).
[0096] The temporal behavior recognition output unit is used to extract and combine features to identify the behavior types in the video, and output the behavior categories through feature fusion of the compressed temporal key points and behavior type features.
[0097] The behavior recognition module is the same model as the backbone network model before the output of the previous behavior type detection. The difference is that it is a pre-trained model in advance, only used for feature extraction, and its parameters are frozen when performing the behavior recognition task and do not perform backpropagation.
[0098] In summary, in the model construction process of the present invention, the method of module pre-training and then overall tuning is adopted. The Swin-transformer model is mainly used in feature extraction on visual frames, which plays a role similar to the convolution operation in CNN in image processing. Due to the characteristics of the Transformer structure, it has advantages in processing global relationships, long-distance information transmission, etc., and can effectively improve the model's abstract ability and expression ability for image content. The SEAttention structure is adopted, and its advantages lie in enhancing feature importance, improving model performance, and reducing computational complexity. At the same time, it can be widely applied to various deep learning tasks, which has important significance and application prospects. Multiple detection heads are used to improve the detection ability for targets of different scales, and four detection heads are used to solve the problems of missed detection or poor detection effects. In addition, when extracting features, the key point information of the behavioral target is introduced to more precisely obtain the behavioral attribute features, further improving the accuracy in detection and behavior recognition.
[0099] Through the behavior tracking module, the present invention demonstrates real-time performance and high efficiency in target tracking, is applicable to application scenarios that require quick response, can effectively predict the state of a dynamic system, improve tracking accuracy, and support multi-target tracking. At the same time, through the feature fusion and information fusion capabilities, the target state is updated by combining the observation data and the prediction results, reducing errors and handling noise and uncertainty. The modeling and update mechanism therein makes it an ideal choice for dealing with complex tracking scenarios.
[0100] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present invention can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
Claims
1. A method for identifying office scene behaviors based on behavior key points, characterized in that, It includes the following steps: S1. Feature extraction is performed in a single video frame through an identification system; S2. After single-frame multi-object behavior detection in the video stream, the behavior objects in the image are identified and located; S3. The trajectory of the obtained behavior target object is used to intercept the behavior target object frame by frame from the video frame; S4. The features at each moment of the temporal vision are located and identified and classified.
2. The method for identifying office scene behaviors based on behavior key points according to claim 1, wherein: In S1, the identification system includes a behavior detection module, a behavior tracking module, a video generation module, and a behavior recognition module; The behavior detection module is connected to the behavior tracking module, the behavior tracking module is connected to the video generation module, and the video generation module is connected to the behavior recognition module; The behavior detection module is used to extract features from a single video frame; The behavior tracking module is used to identify and locate the behavior objects in the image after single-frame multi-object behavior detection in the video stream; The video generation module is used to intercept the behavior target object frame by frame in the video frame; The behavior recognition module is used to locate and identify and classify the behaviors in the video.
3. A method for identifying office scene behaviors based on behavior key points according to claim 1, characterized in that: The behavior detection module includes a Backone unit, a Neck unit, and a Head unit; The Backone unit is connected to the Neck unit, and the Neck unit is connected to the Head unit; The Backone unit is used to solve the feature extraction of small samples, and introduces Swin-Transformer, and performs multi-scale feature extraction through PatchMerging in the way of feature fusion; The Neck unit is used to introduce the SEAttention structure to enhance the model expression ability; The Head unit is used to add a detection head for small scales, and output the category of behavior detection, the position information of the target object in the visual space, and the key point space information; The PatchMerging operation in the Swin-Transformer performs downsampling by using a mechanism similar to the pooling operation in the convolutional network; The SEAttention realizes the modeling and optimization of the dependence between feature channels through two key operations of squeeze and excitation.
4. The method for identifying office scene behaviors based on behavior key points according to claim 3, wherein: The key point space information includes key points in the facial area, key points in the left area of the upper body, key points in the lower body area, and key points in the area of action accessories; The key points in the facial area include the left and right ears, the left and right eyes, the tip of the nose, the left and right corners of the mouth, the upper lip, and the lower lip; The key points in the upper body area include the left and right shoulders, the left and right elbows, and the left and right wrists; The key points in the lower body area include the left and right hips, the left and right knees, and the left and right ankles; The key points in the area of action accessories include the upper left vertex of the mobile phone, the tip of the pen, and the pen head.
5. The method for identifying office scene behaviors based on behavior key points according to claim 2, wherein: The behavior tracking module includes a detection box grading unit, a Kalman filter prediction unit, a data association unit, and a trajectory update and management unit; The detection box grading unit is connected to the Kalman filter prediction unit, the Kalman filter prediction unit is connected to the data association unit, and the data association unit is connected to the trajectory update and management unit.
6. The method for identifying office scene behaviors based on behavior key points according to claim 5, characterized in that: The detection box grading unit is used to perform grading processing on the detection boxes in each frame of the image; The Kaufman filter prediction unit is used to predict the candidate boxes to estimate the position and movement trajectory of the behavior target object in the next frame; The data association unit is used to match the detection boxes in the current frame with the trajectories in the previous frames to establish or update the trajectories of the target objects; The trajectory update and management unit is used to update and manage the trajectory states of all behavior target objects in each frame of the algorithm; The data association unit adopts BYTE data association.
7. A method for identifying office scene behaviors based on behavior key points according to claim 2, characterized in that: The video generation module includes a behavior trajectory acquisition unit, a single-target behavior cropping unit, a behavior strategy determination unit, and a video production unit; The behavior trajectory acquisition unit is connected to the single-target behavior cropping unit, the single-target behavior cropping unit is connected to the behavior strategy determination unit, and the behavior strategy determination unit is connected to the video production unit.
8. The method for identifying office scene behaviors based on behavior key points according to claim 7, characterized in that: The behavior trajectory acquisition unit is used to obtain the trajectory information of each behavior target object from the behavior tracking module; The single-target behavior cropping unit is used to intercept the behavior segments of the target object frame by frame from the video frames; The behavior strategy determination unit is used to determine whether the intercepted behavior segments conform to the preset behavior strategies; The video production unit is used to synthesize the behavior segments that conform to the preset behavior strategies into a complete video file.
9. The method for identifying office scene behaviors based on behavior key points according to claim 2, wherein: The behavior recognition module includes a video frame feature extraction unit, a target object temporal feature unit, and a temporal behavior recognition output unit; The video frame feature extraction unit is connected to the target object temporal feature unit, and the target object temporal feature unit is connected to the temporal behavior recognition output unit.
10. The method for identifying office scene behaviors based on behavior key points according to claim 9, characterized in that: The video frame feature extraction unit is used to extract useful visual features from the video frames; The target object temporal feature unit is used to combine and compress the key point features and the features of the behavior detection types extracted frame by frame in the time dimension to form a more compact and useful feature representation; The temporal behavior recognition output unit is used to identify the behavior types in the video with the extracted and combined features.