Window service behavior specification detection method and system for video space-time attention
By building a Transformer model based on multi-token query, combining local and global feature extraction modules, the problems of wasted computing resources and insufficient accuracy in window service behavior recognition are solved, and efficient behavioral norm detection is achieved.
Patent Information
- Application Number
- CN202510403110.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
In the window service behavior specification detection, it is difficult for the prior art to effectively distinguish highly similar and subtle differences in service behavior, resulting in waste of computing resources and insufficient recognition accuracy.
The Transformer model based on multi-token query is adopted, combined with the local feature aggregation module and the global spatiotemporal feature extraction module, and through 3D convolution and attention mechanism, the details of behavior and global context information are captured to optimize the feature extraction process.
While reducing the computational complexity, the accuracy and efficiency of behavior recognition are improved, the standardization of window service behavior can be accurately detected, and the interference of inter-frame redundant feature is reduced.
Smart Images

Figure CN120340130A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and specifically relates to a method and system for detecting window service behavior norms based on video spatio-temporal attention. Background Art
[0002] Video behavior recognition is an important research direction in the field of computer vision. Its core task is to accurately identify and classify the behavior actions of people from video sequences. The key to this task lies in extracting effective spatio-temporal features from continuous video frames. Due to the existence of a large amount of repetitive or highly correlated visual information between and within video frames, the model performs repeated calculations on these redundant information, resulting in a waste of computing resources. In the specific application scenario of window service behavior norm detection, there is a significant high similarity between some service behaviors, and their differences are only reflected in subtle action features. Taking the two behaviors of "business guidance" and "indicating to take a seat" as examples, their main visual manifestations are almost exactly the same in most time series, and there are only differences in spatial features in a few key frames. This high similarity makes it difficult for traditional local spatio-temporal feature extraction methods to achieve accurate discrimination. It is necessary to model the global context information of the entire behavior sequence to effectively capture the essential differences of behaviors. Window service behaviors usually manifest as continuous dynamic processes, and often consist of multiple sub-actions with temporal dependence relationships. For example, the behavior of "handling business" may include multiple sub-action stages such as "receiving customers", "checking information", and "signing and confirming". In order to achieve accurate recognition and norm detection of such behaviors, the recognition model needs to have two capabilities: on the one hand, it must be able to accurately capture local detail features, such as the subtle changes in hand movements and the adjustment of body postures; on the other hand, it also needs to effectively model the global context information of the entire behavior sequence to understand the complete semantic information of the behavior. Summary of the Invention
[0003] Based on the problems existing in the prior art, the present invention provides a method and system for detecting window service behavior norms based on video spatio-temporal attention. To achieve the above object, the present invention provides the following technical solutions:
[0004] A method and system for detecting window service behavior norms based on video spatio-temporal attention, comprising the following steps:
[0005] S1: Specify the installation location and hardware conditions of the acquisition device, and construct a training and test data set;
[0006] S2: Construct a behavior recognition method based on a Transformer with multi-token queries for behavior detection;
[0007] S3: After preprocessing the video of the window service staff, input it into the constructed behavior recognition method based on multi-token query Transformer for training;
[0008] S4: Use the trained behavior recognition model for actual window service behavior specification detection.
[0009] Furthermore, the specific steps in step S1 include: setting the installation location and hardware conditions of the acquisition device, such as the camera performance requirements, installation location, and shooting angle, etc., to capture the video frames of the service staff in the window service scenario to meet the requirements of behavior action detection and recognition.
[0010] Furthermore, step S2 includes the following steps:
[0011] S21: Process the input video data through chunking and positional encoding.
[0012] S22: Capture the detailed information of the behavior through the constructed shallow local feature aggregation module, enhance the perception ability of local actions, and reduce the calculation of redundant information at the same time.
[0013] S23: Capture the global context information of the behavior through the constructed global spatio-temporal feature extraction module, enhance the adaptability to complex scenarios, and model the long-distance temporal dependence relationship at the same time.
[0014] S24: Complete the classification of the behavior through the classification module.
[0015] Furthermore, step S21 includes the following steps:
[0016] S211: Assume that the input video data is 3×T×H×W, where T represents the sampling frequency of the video, and H and W represent the height and width of each frame. Map the input video to L tokens (i.e., D×K h ×K w where D is the size of the convolutional kernel in the time direction, and K h and K w are the height and width of the convolutional kernel in the spatial directions respectively). Map the input video to L tokens, which are represented as X in ∈R L×C , where L=(T×H×W) / (D×K h ×K w ), representing the number of mapped tokens, C represents the dimension of the token, and D is the size of the convolutional kernel in the time direction. Perform positional encoding on the L tokens in X in and generate a classification token to obtain the encoded representation X in The encoded representation Xpos ∈R K×C Among them, K = L + 1 represents the total number of tokens after adding the classification token.
[0017] Furthermore, S22 constructs a Local Feature Aggregation Module (LFAM) composed of a Spatiotemporal Feature Attention Module (SFAM), aiming to introduce local inductive bias, improve the robustness of the model, and solve the problem of resource waste caused by the inability of spatial attention to focus on local regions at the initial stage of feature extraction in Vision Transformer. Each SFAM includes the following specific steps:
[0018] Normalize the input data X through batch normalization to stabilize the feature distribution and avoid gradient explosion or gradient disappearance during training. Use 3D convolution to compress the channels of the normalized features to reduce the computational complexity and obtain pos where C is the number of channels after convolution, C d = C / dw_reduction, and dw_reduction is the reduction factor in depthwise separable convolution. d
[0019]
[0020] In the formula, bn represents batch normalization. When i = 1, the input data is X pos , and the input of other layers is the output of the previous layer of SFAM The range of i is from 1 to M, and M is the number of layers of SFAM in LFAM.
[0021] Use 3D depthwise separable convolution DWConv3D to further perform feature extraction on to obtain the output feature This convolution operation decomposes the standard convolution into per-channel spatial convolution and pointwise 1×1×1 convolution, effectively reducing the number of model parameters and computational complexity while retaining the key information of the features.
[0022]
[0023] Introduce a channel attention module CAM to process to generate a channel attention map M T . Multiply M T with channel-wise pointwise to obtain The CAM can adaptively adjust the channel weights, enabling the model to better focus on the key channel information.
[0024]
[0025] In the formula, ⊙ C represents element-wise multiplication channel by channel, and Cj_attn represents passing through the attention module.
[0026] The spatial attention module SAM is introduced to process and generate the spatial attention weight map M S . Multiply M S element-wise with to obtain The spatial attention module SAM can help the model effectively focus on the key regional information within the frame.
[0027]
[0028] In the formula, ⊙ E represents element-wise multiplication, and Sp_attn represents passing through the attention module.
[0029] The features obtained after being processed by SAM are passed through a 3D convolution to increase the channel dimension and are connected residually with the input features to generate the final output
[0030]
[0031] In the formula, SFAM i represents the i-th layer SFAM in the local feature aggregation module LFAM.
[0032] Furthermore, the global spatio-temporal feature aggregation module (Global Spatial Feature Extraction Module, GSFEM) constructed in step S23. The GSFEM is composed of 12 identical spatio-temporal processing units (Spatio-Temporal Processing Unit, STPU) stacked together. Each STPU consists of three modules: the hybrid spatial perception module (Hybrid Spatial Perception Module, HSPM), the temporal attention module for multi-query token selection (Temporal Attention Module for Multi-Query Token Selection, TAMMQTS), and the spatio-temporal feature fusion module (Spatial-Temporal Feature Fusion Module, STFFM). Step S23 specifically includes the following:
[0033] S231: Input the output features from LFAM into the HSPM of STPU. The main task of HSPM is to extract the global context information in the spatial dimension from the input features.
[0034] S232: Input the output features of the HSPM after step S231 into TAMMQTS. The main task of TAMMQTS is to extract the global dependencies in the time dimension from the input features.
[0035] S233: Input the output features of the HSPM in step S231 and the output features of the TAMMQTS in step S232 into STFFM for fusion. The specific calculation formula is as follows:
[0036]
[0037] In the formula, represents the output features of the HSPM of the i th STPU, represents the output features of the multi-TAMMQTS of the i th STPU, represents the output features of the i th STPU.
[0038] Furthermore, in step S231, the specific steps of the said HSPM include:
[0039] After normalizing the input data through batch normalization, the local spatiotemporal multi-head relation aggregator (LS_MHRA) and the enhanced features are used for residual connection with the input features to obtain the locally enhanced features
[0040] LS_MHRA describes the local spatiotemporal correlation between tokens by defining the affinity matrix . To enhance the attention to tokens in the local neighborhood, the affinity matrix of LS_MHRA is limited to the local range and contains the learnable parameter matrix a ls ∈R t×3×3 .
[0041]
[0042] In the formula, represents the affinity matrix of the kth token in the nth head, n = 1, …, N, where N represents the number of LS_MHRAs. k = 1, …, K, where K represents the total number of tokens. xk Represents the k-th token in the input data, x in the input data j Indicates that x k is one of the tokens within the range of a ls By calculating and comparing all x within this range, x j can learn the local spatio-temporal relationships between all x within the affinity limit range. Each token in the input data k has its own corresponding affinity matrix. j The global affinity matrix is composed of the concatenation of the affinity matrices of all tokens in
[0043] The global affinity matrix is composed of the concatenation of the affinity matrices of all tokens in, and its calculation formula is:
[0044]
[0045] In the formula, represents the global affinity matrix of the n-th head in LS_MHRA. To further model the relationships between tokens, the relation aggregator R(·) of the n-th head is defined as:
[0046]
[0047] In the formula, V n (·) represents the linear projection of the input information of this function.
[0048] LS_MHRA concatenates the results of the relation aggregators of N heads and performs a linear transformation, and its calculation formula is as follows:
[0049]
[0050] In the formula, U ∈ R C×C is a learnable fusion matrix.
[0051] Layer normalization is used to process the input data features Introduce the Global Spatial Multi-Head Relation Aggregator (GS_MHRA) to obtain the feature representation after global spatial association This process can enhance the global spatial perception ability of a single-frame image, enabling the model to more effectively model long-range spatial dependence relationships.
[0052]
[0053] Through the feed-forward neural network FFN on Processed to obtain output features This process preserves local details and global context information, further enhancing the network's feature expression ability.
[0054]
[0055] Furthermore, in step S232, TAMMQTS designs two processing paths: one path is the classification token extraction path, which extracts the classification token from the input data by splitting, and the classification token is denoted as The other path is the screening path for spatio-temporal feature tokens. The screening path for spatio-temporal feature tokens includes the following specific steps:
[0056] Introduce temporal information into the input data through the dynamic position encoding DPE. The specific calculation formula is as follows:
[0057]
[0058] In the data filter out the motion features, and extract spatio-temporal tokens from the input data using the method of segmented cutting, denoted as Through (Multi-Query Token Selection, MQTS) screening of the tokens, obtain the key and redundant tokens and denote them as This process introduces a context information vector called History_token to guide the cross-attention calculation. History_token does not come from the actual input data, but is used as a fixed reference vector for the attention mechanism. The specific calculation formula is as follows:
[0059] Perform cross-attention calculation on and History_token to obtain the enhanced features This process integrates the global context information into the input data to enhance the temporal feature representation.
[0060]
[0061] Through the fully connected layer, perform a linear transformation on and introduce a non-linear transformation through the activation function to obtain the importance score of each token.
[0062]
[0063] Where Attn_Weight iIndicates the importance score of each token, and δ represents the Quick_Gelu activation function.
[0064] Take the top k values of Attn_Weight through the Top-k screening algorithm and record the indices corresponding to these k values. A boolean mask matrix is constructed based on these indices to distinguish between key tokens and redundant tokens. In the experiment, by adjusting the value of k, the ratio of key tokens to redundant tokens can be dynamically controlled to explore the impact of different ratios on the model performance. i The above text wraps around here.
[0065] Mask_Weight i = TOP-K(Attn_Weight i )
[0066] In the formula, Mask_Weigjt i represents the boolean mask matrix constructed by k indices.
[0067] Multiply Mask_Weight i element-wise with to obtain the feature matrix Fusion_Weight i .
[0068]
[0069] In the formula, ⊙ represents element-wise multiplication.
[0070] Divide Fusion_Weigjt i into two parts: key tokens and redundant tokens through the boolean mask matrix.
[0071]
[0072] In the formula, the Divide operation takes advantage of the characteristics of the boolean mask matrix. When multiplying the data by 1 in the boolean mask matrix, the original feature remains unchanged, and the corresponding feature is hidden where the value in the boolean mask matrix is 0. The above text wraps around here.
[0073] After the classification token extraction path and the spatio-temporal feature token screening path are processed, TAMMQTS focuses on the changes of key tokens. Through GC_MHRA, and are fused to obtain the output feature This process fuses the information of classification tokens and spatio-temporal feature tokens, captures richer semantic and spatio-temporal dependencies, and reduces the calculation of redundant tokens by focusing on the changes of key tokens.
[0074]
[0075] Process through FFN to obtain the output features of TAMMQTS
[0076]
[0077] Among them, TAMMQTS i represents the STPU in GSFEM i of TAMMQTS.
[0078] Furthermore, in step S24, the features after GSFEM are segmented, classification tokens are extracted and input into a classifier, and behavior classification and specification detection are performed in combination with the window service behavior specification detection method. The specific steps are as follows:
[0079] S241: Segment the input data and extract classification tokens, denoted as where N represents the total number of STPUs in GSFEM.
[0080] S242: Input the from step S41 into the classifier. The calculation formula of the classifier is as follows:
[0081]
[0082] In the formula, FCLayer represents a fully connected layer, the input dimension of which is C (i.e., the feature dimension of each token), and the output dimension is K (i.e., the total number of predefined behavior categories).
[0083] Furthermore, the step S3 includes the following steps:
[0084] S31: Statistically screen the work videos and labels of service personnel in the dataset;
[0085] S32: Normalize the input data to the same size and input it into the behavior recognition method based on multi-token query Transformer for classification and prediction.
[0086] On the other hand, the present invention provides a system for the behavior recognition method based on multi-token query Transformer, including:
[0087] A data input module for transmitting the video data captured by the monitoring device into a pre-trained behavior recognition model based on multi-token query Transformer;
[0088] The feature extraction and behavior classification module uses the behavior recognition network of the Transformer based on multi-token query to extract spatio-temporal features and classify the behaviors in the video.
[0089] The behavior determination and feedback module is used to determine the classification results, identify behaviors that do not meet the specifications, and remind the staff through sound, pop-up windows or log records to ensure that the service behaviors meet the specification standards.
[0090] The beneficial effects of the present invention are as follows: The Transformer behavior recognition model and system based on multi-token query proposed by the present invention can achieve higher recognition accuracy with lower computational complexity compared with other behavior recognition solutions based on Vision Transformer. The system innovatively constructs a local feature aggregation module, which combines 3D convolution with channel and spatial attention mechanisms, enabling spatial attention to focus on shallow local features, effectively retaining key spatial information while reducing the computational amount of shallow feature extraction in the model. The core multi-query token selection temporal attention module of the system is specifically used to extract important temporal features, accurately aggregate key action information, and effectively avoid the interference of redundant inter-frame features. The hybrid spatial attention mechanism adopted by the system improves the traditional ViT feature extraction module, significantly enhancing the attention ability to local neighborhood spatio-temporal features while maintaining the global spatial feature extraction ability. The overall architecture optimization of the system achieves the best balance between computational efficiency and recognition accuracy in the behavior recognition task.
[0091] The present invention aims at the behavior specification detection in the scenario of window service. By calculating the features currently calculated and the features of the examples in the following text through the multi-query token selection temporal attention module, the intelligent distinction between key features and redundant features is realized. This distinction enables the system to accurately focus on key temporal features, effectively improving the accuracy of action feature aggregation. The module can reduce the interference of redundant inter-frame information and demonstrate excellent detection performance in practical application scenarios such as window service specifications.
[0092] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings
[0093] In order to make the objectives, technical solutions and advantages of the present invention clearer, the modules of the present invention will be described in detail below with reference to the drawings, where:
[0094] Figure 1This is the workflow diagram of the window service behavior specification detection method and system of the present invention.
[0095] Figure 2 This is the model architecture diagram of the present invention.
[0096] Figure 3 This is the architecture diagram of each spatio-temporal attention module in the shallow local feature aggregation module of the model architecture of the present invention.
[0097] Figure 4 This is the architecture diagram of the multi-query token selection mechanism MQTS of the time attention module TAMMQTS for multi-query token selection in the model architecture of the present invention. Specific Embodiments
[0098] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be variously modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0099] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0100] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be understood as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0101] S1 includes device installation, setting of installation conditions, and construction of a behavior label library:
[0102] S11: Specify the installation location and conditions of the acquisition device. Install a camera on each service window to capture real-time videos of service personnel. It is required that the cameras arranged on-site have a frame rate of no less than 30fps, and at the same time, the resolution of the cameras is not less than 1080P, so as to transmit and process real-time video information for subsequent calculations.
[0103] S12: When establishing the behavior label library, first identify the start and end times of each behavior from the video data obtained by the acquisition device, and perform dynamic cropping based on the integrity of the behavior to ensure that each video segment precisely corresponds to the time interval when the behavior occurs, while maintaining the temporal coherence of the action sequence. The cropped video segments are classified and stored by behavior name. Subsequently, digital encoding is performed on all behavior labels to convert the text category labels into numerical labels, and the encoded labels are stored in the behavior label library together with the corresponding video segment paths. Each row of the behavior label library includes two parts: the path of the cropped video segment and the numerical label of this segment, providing structured data support for subsequent model training and detection.
[0104] The said S2 includes the following steps:
[0105] S21: The input video data is processed through chunking and positional encoding.
[0106] S22: Capture the detailed information of the behavior through the constructed shallow local feature aggregation module, enhance the perception ability of local actions, and at the same time reduce the calculation of redundant information.
[0107] S23: Capture the global context information of the behavior through the constructed global spatio-temporal feature extraction module, enhance the adaptability to complex scenes, and at the same time model long-range temporal dependencies.
[0108] S24: Complete the classification of behaviors through the classification module.
[0109] Furthermore, the said step S21 includes the following steps:
[0110] S211: Chunking and positional encoding processing of data. Suppose the input video data is 3×8×224×224, where 8 represents the sampling frequency of the video, and 224 represents the height and width of each frame. The input video is mapped into 1568 tokens through three-dimensional convolution (the convolution kernel size is 1×16×16), and they are represented as X in ∈R 1568×768 , 768 represents the dimension of the token. Perform positional encoding on the 1568 tokens in X in , and generate a classification token to obtain X in The encoded representation X pos ∈R 1569×768 .
[0111] Further, step S22 includes the following steps:
[0112] Construct a Local Feature Aggregation Module (LFAM) aimed at introducing local inductive bias, enhancing the robustness of the model, and solving the problem of resource waste caused by the inability of Vision Transformer to focus spatial attention on local regions at the initial stage of feature extraction. LFAM consists of 4 Spatio-Temporal Feature Attention Modules (SFAMs). Each SFAM includes the following specific steps:
[0113] Optionally, the number of Spatio-Temporal Feature Attention Modules (SFAMs) in LFAM can be 2, 3, 5, or 6.
[0114] Normalize the input data X through batch normalization pos to stabilize the feature distribution and avoid gradient explosion or gradient disappearance during training. Use 3D convolution to compress the channels of the normalized features to reduce the computational complexity, obtaining where 384 is the number of channels after convolution.
[0115]
[0116] In the formula, bn represents batch normalization. When i = 1, the input data is X pos , and the input of other layers is the output of the previous layer's SFAM The range of i is from 1 to 4, and 4 is the number of layers of SFAMs in LFAM.
[0117] Use 3D depthwise separable convolution (DWConv3D) to further extract features from to obtain the output features This convolution operation decomposes the standard convolution into a per-channel spatial convolution and a pointwise 1×1×1 convolution, effectively reducing the number of model parameters and computational complexity while retaining the key information of the features.
[0118]
[0119] Introduce a Channel Attention Module (CAM) to process to generate a channel attention map M T . Multiply M T element-wise with to obtain CAM can adaptively adjust the channel weights, enabling the model to better focus on key channel information.
[0120]
[0121] In the formula, ⊙ CThe symbol ⊙ represents element-wise multiplication, and Ch_attn represents passing through the attention module.
[0122] The spatial attention module SAM is introduced to process and generate the spatial attention weight map M S . Multiply M S element-wise with to obtain The spatial attention module SAM can help the model effectively focus on the key region information within the frame.
[0123]
[0124] In the formula, ⊙ E represents element-wise multiplication, and Sp_attn represents passing through the attention module.
[0125] The features obtained after being processed by SAM are used to enhance the channel dimension through 3D convolution and perform a residual connection with the input features to generate the final output
[0126]
[0127] In the formula, SFAM i represents the i-th layer SFAM in the local feature aggregation module LFAM.
[0128] Furthermore, the global spatio-temporal feature aggregation module (Global Spatial Feature Extraction Module, GSFEM) constructed in step S23. GSFEM is stacked by 12 identical spatio-temporal processing units (Spatio-Temporal Processing Unit, STPU). Each STPU consists of three modules: the hybrid spatial perception module (Hybrid Spatial Perception Module, HSPM), the temporal attention module for multi-query token selection (Temporal Attention Module for Multi-Query Token Selection, TAMMQTS), and the spatio-temporal feature fusion module (Spatial-Temporal Feature Fusion Module, STFFM). Step S23 specifically includes the following:
[0129] S231: Input the output features after passing through LFAM into the HSPM of STPU. The main task of HSPM is to extract the global context information in the spatial dimension from the input features.
[0130] S232: Input the output features of the HSPM after step S231 into TAMMQTS. The main task of TAMMQTS is to extract the global dependencies in the time dimension from the input features.
[0131] S233: Input the output features of the HSPM in step S231 and the output features of the TAMMQTS in step S232 into STFFM for fusion. The specific calculation formula is as follows:
[0132]
[0133] In the formula, represents the output features of the HSPM of the i th STPU, represents the output features of the multi-TAMMQTS of the STPU i , represents the output features of the STPU i .
[0134] Furthermore, in step S231, the specific steps of the described HSPM include:
[0135] After normalizing the input data through batch normalization, the local spatiotemporal multi-head relation aggregator (LS_MHRA) and the enhanced features are used for residual connection with the input features to obtain the locally enhanced features
[0136] LS_MHRA describes the local spatiotemporal correlation between tokens by defining an affinity matrix . To enhance the attention of tokens in the local neighborhood, the affinity matrix of LS_MHRA is limited to the local range, including a learnable parameter matrix a ls ∈R t×3×3 .
[0137]
[0138] In the formula, represents the affinity matrix of the kth token in the nth head, where n = 1, …, 12 represents the number of LS_MHRAs. k = 1, …, 1569 represents the total number of tokens. x k represents the kth token in the input data , x j represents a certain token within the range of x k in a ls range. By summing up all x within this rangej By performing computational comparisons, x can be made k learn the local spatio-temporal relationships among all x within the affinity limit range. Each token in the input data j has its own corresponding affinity matrix. The global affinity matrix
[0139] is composed of the affinity matrices of all tokens in and its calculation formula is:
[0140]
[0141] In the formula, represents the global affinity matrix of the nth head in LS_MHRA. To further model the relationships between tokens, the relation aggregator R(·) of the nth head is defined as:
[0142]
[0143] In the formula, V n (·) represents the linear projection of the input information of this function.
[0144] LS_MHRA concatenates the results of the relation aggregators of 12 heads and performs a linear transformation to obtain Its calculation formula is as follows:
[0145]
[0146] In the formula, U ∈ R 768×768 is a learnable fusion matrix.
[0147] The input data features are processed using layer normalization and a Global Spatial Multi-Head Relation Aggregator (GS_MHRA) that introduces global spatial associations is employed to obtain the feature representation after global spatial associations This process can enhance the global spatial perception ability of a single-frame image, enabling the model to more effectively model long-range spatial dependencies.
[0148]
[0149] The is processed through a feed-forward neural network FFN to obtain the output features This process preserves local details and global context information, further enhancing the feature expression ability of the network.
[0150]
[0151] Further, in step S232, TAMMQTS designs two processing paths: one path is the classification token extraction path, which extracts the classification tokens from the input data by splitting, and the classification tokens are denoted as The other path is the screening path for spatio-temporal feature tokens. The screening path for spatio-temporal feature tokens includes the following specific steps:
[0152] Introduce temporal information into the input data through the dynamic position encoding DPE to obtain The specific calculation formula is as follows:
[0153]
[0154] Screen out the motion features from the data and extract the spatio-temporal tokens from the input data using the segmented cutting method, denoted as Screen the tokens of through (Multi-Query Token Selection, MQTS) to obtain the key and redundant tokens and denote them as This process introduces a context information vector called History_token to guide the cross-attention calculation. History_token does not come from the actual input data but serves as a fixed reference vector for the attention mechanism. The specific calculation formula is as follows:
[0155] Perform cross-attention calculation on and History_token to obtain the enhanced feature This process incorporates the global context information into the input data to enhance the temporal feature representation.
[0156]
[0157] Perform a linear transformation on through the fully connected layer and introduce a non-linear transformation using the activation function to obtain the importance score of each token.
[0158]
[0159] In the formula, Attn_Weight i represents the importance score of each token, and δ represents the Quick_Gelu activation function.
[0160] Get Attn_Weight by Top-k filtering algorithm i The first k values of are obtained and the corresponding indices of these k values are recorded. A Boolean mask matrix is constructed based on these indices to distinguish key tokens from redundant tokens. In the experiment, by adjusting the k value, the ratio of key tokens to redundant tokens can be dynamically controlled to explore the impact of different ratios on model performance.
[0161] Mask_Weight i =TOP-K(Attn_Weight i )
[0162] Where, Mask_Weight i Represents a Boolean mask matrix constructed by k indices. The optional k can be any number between 0 and 1.
[0163] Mask_Weight i and Multiply element by element to get the feature matrix Fusion_Weight i .
[0164]
[0165] Where ⊙ represents element-by-element multiplication.
[0166] Fusion_Weight is converted to i Divided into two parts: key tokens and redundant tokens.
[0167]
[0168] In the formula, the Divide operation uses the characteristics of the Boolean mask matrix to separate the values of 1 and data in the Boolean mask matrix. The original features will not be changed after multiplication, and the corresponding features will be hidden where the value is 0 in the Boolean mask matrix.
[0169] After the classification token extraction path and spatiotemporal feature token screening path are processed, TAMMQTS focuses on the changes of key tokens. and Fusion, get the output features This process fuses the information of classification tokens and spatiotemporal feature tokens to capture richer semantics and spatiotemporal dependencies, and reduces the calculation of redundant tokens by focusing on the changes of key tokens.
[0170]
[0171] Process through FFN to obtain the output features of TAMMQTS
[0172]
[0173] Among them, TAMMQTS i represents STPU in GSFEM i of TAMMQTS.
[0174] Furthermore, in step S24, the features after GSFEM are segmented, classification tokens are extracted and input into a classifier, and behavior classification and specification detection are performed in combination with the window service behavior specification detection method. The specific steps are as follows:
[0175] S241: Segment the input data and extract classification tokens, denoted as where N represents the total number of STPUs in GSFEM.
[0176] S242: Input the from step S41 into the classifier. The calculation formula of the classifier is as follows:
[0177]
[0178] In the formula, FCLayer represents a fully connected layer, whose input dimension is 768 and output dimension is K (i.e., the total number of predefined behavior categories).
[0179] Furthermore, the step S3 includes:
[0180] S31: Use the public dataset for training and the self-annotated dataset for testing. For the service personnel in the dataset.
[0181] S32: Set the parameters required for training: number of training epochs: epoch = 50, validation frequency: 1 epoch, number of frames per video: 8, input batch size: batch_size = 16, size of training images: train_img_size = 224, number of input channels: 3, global DropPath rate: 0.5, momentum of the optimizer method AdamW: 0.9, color perturbation intensity: 0.4, image interpolation method: bicubic, RandomErasing mode: pixel, base learning rate: 1×10 -5 , minimum learning rate: 1×10 -6, the cosine annealing learning rate adjustment strategy is adopted. The number of layers of the spatio-temporal feature attention module in the local feature aggregation module is set to 4, and the selection ratio of the key tokens to the redundant tokens of the token selection mechanism of the time attention module for multi-query tokens is 1:1.
[0182] Further, the step S4 includes:
[0183] After this classification result is cited into the system, the system will compare it with the predefined behavior specifications. If the detected behavior conforms to the specifications, the system records the behavior and continues to monitor; if an irregular behavior is detected, the system will trigger a prompt mechanism, such as reminding the staff by means of sound, pop-up window or log record, to ensure that the service behavior conforms to the specification standards.
[0184] The present invention provides a window service behavior specification detection device for use in the S1, including: a monitoring device, a computer.
[0185] The monitoring device can capture window service scene images in real time and transmit the construction site images to the computer, facilitating the detection of window service behavior specifications.
[0186] The computer can receive the construction site images transmitted by the monitoring device, input the computer program proposed by the present invention, and load the detection model for detection, so as to obtain the window service behavior specification detection result.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method and system for detecting window service behavior norms based on video spatio-temporal attention, comprising the following steps: S1: Specify the installation location and hardware conditions of the acquisition device, and construct a training and test data set; S2: Construct a behavior recognition method based on Transformer with multi-token query for behavior detection; S3: After preprocessing the video of the window service personnel, input it into the constructed behavior recognition method based on Transformer with multi-token query for training; S4: Use the trained behavior recognition model for actual window service behavior norm detection. Further, in step S1, it specifically includes: setting the installation location and hardware conditions of the acquisition device, for example, camera performance requirements, installation location and shooting angle, etc., to capture the video frames of the service personnel in the window service scenario to meet the requirements of behavior action detection and recognition. Further, step S3 includes the following steps: S31: Statistically screen the work videos and labels of the service personnel in the data set; S32: Normalize the input data to the same size and input it into the behavior recognition method based on Transformer with multi-token query for classification and prediction.
2. The method and system for detecting window service behavior specifications of video spatio-temporal attention according to claim 1, wherein: Step S2 includes the following steps: S21: Process the input video data through chunking and position encoding. S22: Capture the detailed information of the behavior through the constructed shallow local feature aggregation module, enhance the perception ability of local actions, and reduce the calculation of redundant information at the same time. S23: Capture the global context information of the behavior through the constructed global spatio-temporal feature extraction module, enhance the adaptability to complex scenarios, and model long-distance temporal dependencies at the same time. S24: Complete the classification of the behavior through the classification module. Furthermore, the specific process of step S21 is as follows: Assume that the input video data is 3×T×H×W, where T represents the sampling frequency of the video, and H and W represent the height and width of each frame. The input video is mapped into L tokens (i.e., D×K h ×K w by three-dimensional convolution. Here, D is the size of the convolutional kernel in the time direction, and K h and K w are the sizes of the convolutional kernel in the height and width directions of space, respectively). The input video is mapped into L tokens, which are represented as X i ∈R L×C , where L = (T×H×W) / (D×K h ×K w ), representing the number of tokens after mapping, C represents the dimension of the token, and D is the size of the convolutional kernel in the time direction. Position encoding is performed on the L tokens in X in , and a classification token is generated to obtain the encoded representation X in X pos ∈R K×C where K = L + 1 represents the total number of tokens after adding the classification token. Further, the specific process of step S24 is as follows: After GSFEM, the features are segmented, the classification tokens are extracted and input into the classifier, and the behavior classification and specification detection are carried out in combination with the window service behavior specification detection method. The specific steps are as follows: S241: Segment the input data and extract the classification tokens, denoted as where N represents the total number of STPUs in GSFEM. S242: Take the input classifier, and the calculation formula of the classifier is as follows: Wherein, FCLayer represents a fully connected layer, the input dimension of which is C (i.e., the feature dimension of each token), and the output dimension is K (i.e., the total number of predefined behavior categories).
3. The method and system for detecting window service behavior specifications of video spatio-temporal attention according to claim 2, wherein: The local feature aggregation module (Local Feature Aggregation Module, LFAM) composed of the local feature aggregation module block (Spatiotemporal Feature Attention Module, SFAM) constructed in step S22 aims to introduce local inductive bias, improve the robustness of the model, and solve the problem of resource waste caused by the inability of spatial attention in the initial stage of feature extraction of Vision Transformer to focus on local areas. Each SFAM includes the following specific steps: Normalize the input data X through batch normalization pos to stabilize the feature distribution and avoid gradient explosion or gradient disappearance during the training process. Use 3D convolution to compress the normalized features in the channel dimension to reduce the computational complexity and obtain where C d is the number of channels after convolution, and C d = C / dw_reduction, where dw_reduction is the reduction factor in depthwise separable convolution. Wherein, bn represents batch normalization. When i = 1, the input data is X pos , and the input of other layers is the output of the previous layer of SFAM The range of i is from 1 to M, and M is the number of layers of SFAM in LFAM. The 3D depthwise separable convolution DWConv3D is further used to extract features, obtaining the output features This convolution operation decomposes the standard convolution into a per-channel spatial convolution and a pointwise 1×1×1 convolution, effectively retaining the key information of the features while significantly reducing the number of model parameters and the computational complexity. Introduce the Channel Attention Module (CAM) to process and generate the channel attention map M T Multiply M T element-wise with to obtain CAM can adaptively adjust the channel weights, enabling the model to better focus on key channel information. In the formula, ⊙ C represents channel-wise multiplication, and Ch_attn represents passing through the attention module. Introduce the Spatial Attention Module (SAM) to process and generate the spatial attention weight map M S . Multiply M S element-wise with to obtain The Spatial Attention Module (SAM) can help the model effectively focus on the key region information within the frame. In the formula, ⊙ E represents element-wise multiplication, and Sp_attn represents passing through the attention module. The features obtained after SAM processing Enhance the channel dimension through 3D convolution and combine with the input features Perform residual connection to generate the final output In the formula, SFAM i represents the i-th layer SFAM in the local feature aggregation module LFAM.
4. A method and system for detecting window service behavior norms of video spatio-temporal attention according to claim 2, characterized in that: The step S23 constructs a Global Spatial Feature Extraction Module (GSFEM). The GSFEM is stacked by 12 identical Spatio-Temporal Processing Units (STPUs). Each STPU consists of three modules: a Hybrid Spatial Perception Module (HSPM), a Temporal Attention Module for Multi-Query Token Selection (TAMMQTS), and a Spatial-Temporal Feature Fusion Module (STFFM). The step S23 specifically includes the following: S231: Input the output features after LFAM into the HSPM of the STPU. The main task of the HSPM is to extract the global context information in the spatial dimension from the input features. S232: Input the output features of the HSPM after step S231 into the TAMMQTS. The main task of the TAMMQTS is to extract the global dependencies in the temporal dimension from the input features. S233: Input the output features of the HSPM in step S231 and the output features of the TAMMQTS in step S232 into the STFFM for fusion. The specific calculation formula is as follows: In the formula, represents the output feature of the HSPM of the STPU i , represents the output feature of the multi-TAMMQTS of the STPU i , represents the output feature of the STPU i .
5. A method and system for detecting window service behavior specifications of video spatio-temporal attention according to claim 4, characterized in that: The specific calculation process of the HSPM in step S231 includes: After normalizing the input data through batch normalization After that, the Local Spatiotemporal Multi-Head Relation Aggregator (LS_MHRA) and the enhanced features will be used to perform a residual connection with the input features to obtain locally enhanced features LS_MHRA describes the local spatio-temporal correlation between tokens by defining an affinity matrix. To enhance the attention to tokens within the local neighborhood, the affinity matrix of LS_MHRA is restricted to the local scope, including a learnable parameter matrix a ls ∈R t×3×3 . Wherein, represents the affinity matrix of the k-th token in the n-th head, where n = 1, …, N, and N represents the number of LS_MHRAs. k = 1, …, K, and K represents the total number of tokens. x k represents the input data the k-th token in, and x j represents x k a certain token within the range of a ls . By calculating and comparing all x j within this range, x k can learn the local spatio-temporal relationship with all x j within the affinity limit range. Each token in the input data has its own corresponding affinity matrix. Global affinity matrix It is formed by concatenating the affinity matrices of all tokens in, and its calculation formula is: In the formula, represents the global affinity matrix of the nth head in LS_MHRA. To further model the relationships between tokens, the relationship aggregator R(·) of the nth head is defined as: where V n (·) represents a linear projection of the input information of the function. The LS_MHRA concatenates the results of the relationship aggregators of N heads and performs a linear transformation. The calculation formula is as follows: where \(U\in R\) C×C is a learnable fusion matrix. The input data features are processed using layer normalization. A Global Spatial Multi-Head Relation Aggregator (GS_MHRA) that introduces global spatial associations is used to obtain the feature representation after global spatial associations. This process can enhance the global spatial perception ability of a single-frame image, enabling the model to more effectively model long-range spatial dependencies. Process through a feed-forward neural network FFN to obtain an output feature This process retains local details and global context information, further enhancing the network's feature expression ability. 。 6. A method and system for detecting window service behavior specifications of video spatio-temporal attention according to claim 4, characterized in that: In the step S232, TAMMQTS designs two processing paths: one path is the classification token extraction path, which extracts the classification tokens by splitting the input data and records the split classification tokens as The other path is the screening path for spatio-temporal feature tokens. The screening path for spatio-temporal feature tokens includes the following specific steps: Introduce temporal information into the input data through Dynamic Position Encoding (DPE), and the specific calculation formula is as follows: Filter out the motion features from the data and extract the spatio-temporal tokens from the input data using the method of segmented cutting, which are represented as Through (Multi-Query Token Selection, MQTS) for the screening of tokens, the key and redundant tokens are obtained and represented as This process introduces a context information vector called History_token to guide the cross-attention calculation. History_token does not come from the actual input data but serves as a fixed reference vector for the attention mechanism. The specific calculation formula is as follows: Cross-attention calculation is performed on and History_token to obtain enhanced features . This process integrates global context information into the input data to enhance the temporal feature representation. After performing a linear transformation on through a fully connected layer and introducing a non-linear transformation using an activation function, the importance score of each token is obtained. After performing a linear transformation on through a fully connected layer and introducing a non-linear transformation using an activation function, the importance score of each token is obtained. where Attn_Weight i represents the importance score of each token, and δ represents the Quick_Gelu activation function. Retrieve the top-k values of Attn_Weight through the Top-k screening algorithm and record the indices corresponding to these k values. A boolean mask matrix is constructed based on these indices to distinguish key tokens and redundant tokens. In the experiment, by adjusting the value of k, the ratio of key tokens to redundant tokens can be dynamically controlled to explore the impact of different ratios on the model performance. i The first k values are taken, and the indices corresponding to these k values are recorded. A boolean mask matrix is constructed based on these indices to distinguish key tokens and redundant tokens. In the experiment, by adjusting the value of k, the ratio of key tokens to redundant tokens can be dynamically controlled to explore the impact of different ratios on the model performance. Mask_Weight i = TOP-K(Attn_Weight i ) where Mask_Weight i represents a Boolean mask matrix constructed by k indices. Multiply Mask_Weight i element-wise with to obtain the feature matrix Fusion_Weight i . In the formula, ⊙ represents element-wise multiplication. Partition Fusion_Weight through a Boolean mask matrix i into two parts: key tokens and redundant tokens. In the formula, the Divide operation takes advantage of the characteristics of the Boolean mask matrix. When multiplying by the data where the value in the Boolean mask matrix is 1, the original feature will not be changed, and the corresponding feature will be hidden where the value in the Boolean mask matrix is 0. After multiplication, the original features are not changed, while the corresponding features are hidden where the value in the Boolean mask matrix is 0. After the processing of the classification token extraction path and the spatio-temporal feature token screening path, TAMMQTS focuses on the changes of key tokens. Through GC_MHRA for and fusion to obtain the output feature This process fuses the information of classification tokens and spatio-temporal feature tokens, captures richer semantic and spatio-temporal dependency relationships, and reduces the calculation of redundant tokens by focusing on the changes of key tokens. Process through FFN to obtain the output features of TAMMQTS Among them, TAMMQTS i represents the STPU in GSFEM i of TAMMQTS.
7. A method and system for detecting window service behavior specifications of video spatio-temporal attention according to claim 1, characterized in that: In step S1, first deploy the hardware system required for service behavior collection, including multi-angle camera devices and a video analysis server, to ensure that the server can process the service window video stream transmitted by the camera device in real time; then perform behavior annotation on the collected video data.
8. A method and system for detecting window service behavior specifications of video spatio-temporal attention according to claim 2, characterized in that: The monitoring device is used to capture the video of the window service scenario in real time and transmit it to the video analysis server; and transmit the video data to the video analysis server; the server is used to receive and process the real-time video data transmitted by the monitoring device, analyze it through a service behavior specification detection network based on spatio-temporal attention, identify the behavior categories of window service personnel, and compare them with the preset standard of standard behaviors, and finally output the service behavior compliance determination result.
9. A method and system for detecting window service behavior specifications of video spatio-temporal attention based on any one of the methods recited in claims 1-8, characterized in that: Including: A data input module, configured to transmit the video data captured by the monitoring device to a pre-trained Transformer behavior recognition model based on multi-token queries; The feature extraction and behavior classification module uses the behavior recognition network based on the multi-token query Transformer to extract spatio-temporal features and classify the behaviors in the video. The behavior determination and feedback module is used to determine the classification results, identify behaviors that do not meet the specifications, and remind the staff through sound, pop-up windows or log records to ensure that the service behaviors meet the specification standards.