Interaction action detection method and device based on multi-level features
Through the interactive action detection method of adaptive multi-level feature extraction and hierarchical relationship inference, the problem of insufficient feature extraction and fine-grained relationship modeling is solved, and efficient interactive behavior recognition is achieved in complex scenarios.
Patent Information
- Application Number
- CN202510472811.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-29
AI Technical Summary
The existing interactive action detection technology has the problems of insufficient feature extraction and utilization of interactive area and insufficient fine-grained relationship modeling, and it is difficult to accurately identify the interaction behavior between people and objects in complex scenarios.
Adaptive multi-level feature extraction module is used to optimize lightweight channel weights through DLA network and ECA module, and combine MobileViT encoder and multi-task decoder to perform character, object and interaction detection, and improve fine-grained interactive classification accuracy through mono-, paired and ternary relationship attention interaction and CLIP pre-trained text encoder.
It improves the accuracy and efficiency of interactive action detection in complex scenarios, can effectively identify the interaction behavior of multiple individuals and objects, solves the shortcomings of feature extraction and relationship modeling in traditional methods, and achieves a quick understanding of complex scenarios.
Smart Images

Figure CN120388227A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of interactive action detection, and in particular, to an interactive action detection method and device based on multi-level features. Background Art
[0002] The interactive action detection task is to detect the people and interactive objects with interactive relationships in an image or video and output their interactive action categories, and finally obtain an interactive triple <person, interactive action, interactive object>. The subject of the triple is a person, and the object includes people and objects. There are person-to-person interactions and person-to-object interactions. Since the person-object interactive action detection task needs to locate people and objects and identify the interactive relationships between them, a more intuitive idea for this task is to first perform object detection on the image, and then identify the existing interactive actions between the detected people and interactive objects. However, the interactive actions formed between people and objects are not simply the sum of the features of people and objects. Interactive action detection often requires reasoning, analysis, and judgment of fuzzy, complex, and difficult-to-identify interactive behaviors, which is a very challenging problem. The person-object interactive action detection task is not a simple combination of the visual content of people and objects existing in the image, but requires the interactive detection algorithm to understand these abstract concepts of the interactive relationships between people and objects and between people and people in the image.
[0003] Currently, there are two difficulties in interactive action detection technology. Difficulty one: insufficient extraction and utilization of interactive region features. There are often some small targets and occluded targets in the living scene, and these targets are more difficult to detect compared to normal targets. Some interactive actions often involve multiple individuals and objects, and the scenes of interactive actions are variable. Different lighting, perspectives, and scale changes, as well as complex target types and backgrounds, will affect the detection and judgment of the correct targets by the interactive detection network. Difficulty two: insufficient fine-grained relationship modeling. Traditional methods only model the simple association of person-object pairs (such as spatial co-occurrence), lacking hierarchical reasoning of individual attributes (human postures), pairwise relationships (person-object spatial distances), and overall semantics (scene context), resulting in difficulty in distinguishing similar interactive actions (such as "riding" a horse vs "leading" a horse).
[0004] Interactive action detection, as an emerging direction in the field of computer vision, combines the object localization task of object detection and the behavior classification task of behavior recognition, that is, locating the behavior subject (person) and behavior object (other people or objects) in an image or video and classifying the behavior exerted by the subject on the object. Its technical challenges lie in: complex scene understanding: it is necessary to model the three-way dependence relationship of people, objects, and interactions; multi-scale feature requirements: taking into account local details (such as hand postures) and global semantics (such as scene context), and the actual application needs to meet the low-latency requirements of edge devices.
[0005] With the development of neural network models and the increasing richness of datasets, computer vision has made rapid progress. The continuous improvement of real-world requirements has also put forward higher demands for understanding visual content, which has brought about a transformation in the feature extraction paradigm for artificial intelligence: from manually extracting features to automatically extracting features by neural networks. This transformation has also had a profound impact on interactive action detection. In this paradigm, the detection of human-object interactive actions can be divided into two subtasks: 1. Instance (human and interactive object) detection, 2. Relationship prediction (classification). According to whether the optimization of these two subtasks is synchronized, interactive action detection can generally be divided into two-stage detection and single-stage detection methods. The main difference between different detection methods lies in the different strategies adopted for instance objects in the interactive recognition stage.
[0006] The two-stage method (represented by iCAN and VSGNet) disassembles the human interaction detection task into two major steps according to the characteristics of the task: In the object detection stage, it focuses on the accurate detection of human and object targets in the image and generates candidate boxes; in the interactive reasoning stage, based on the candidate boxes, it extracts ROI features and predicts interactions through a graph network or an attention mechanism. The advantage of this method is that it can effectively separate the two tasks of object detection and interactive behavior classification, enabling them to be independently optimized and facilitating the integration of diverse information to improve the detection efficiency. However, the two-stage method also has its limitations, that is, the model structure is relatively complex, the recognition efficiency is low, and the staged design cannot process dense scenes in real time.
[0007] To overcome this defect, the single-stage method (represented by QPIC and IDN) emerged. It abandons the step-by-step execution mode of the two-stage method and instead reshapes the human interaction detection task into a parallel processing task, directly outputting triple information end-to-end from the image, thus significantly accelerating the recognition process. The technical process generally uses a ResNet or Transformer backbone to extract image features. Parallel task branches: Human / object detection branch: Generate bounding boxes and classes. Interactive classification branch: Directly predict the interaction classes of human-object pairs. The defect of the existing technology is flat interaction modeling (only human-object pairs), and most single-stage models adopt a framework of multi-task collaborative learning, sharing feature resources among tasks. However, there may be significant differences in feature focus and optimization goals between object detection and relationship prediction. This sharing mechanism sometimes leads to mutual interference between features, thereby restricting the overall optimization effect and making it difficult to reach the performance ceiling. Summary of the Invention
[0008] The present invention provides an interactive action detection method and device based on multi-level features to solve the technical problems existing in the above-mentioned prior art.
[0009] To achieve the above object, the present invention provides an interactive action detection method based on multi-level features, which includes:
[0010] S1: Extract low-level features, middle-level features, and high-level features from the image to be detected and represent them as low-level feature maps, middle-level feature maps, and high-level feature maps respectively. Perform lightweight channel weight optimization on the low-level feature maps, middle-level feature maps, and high-level feature maps respectively, and fuse the information of the low-level feature maps, middle-level feature maps, and high-level feature maps through a cascaded fusion method to obtain a fused feature map;
[0011] S2: Obtain local detail features from the fused feature map and perform global modeling to obtain global features. The global features are then fused with the local detail features to generate global context features. Perform person detection, object detection, and interaction detection in parallel through multiple task branches, and output person features, object features, and interaction features.
[0012] S3: Through the attention interaction of unary relationships, pairwise relationships, and ternary relationships, gradually fuse multi-level contexts, use a pre-trained text encoder to generate text embeddings for interaction categories, align the text embeddings with visual features, and improve the accuracy of fine-grained interaction classification;
[0013] S4: Integrate the context information of the features output by the hierarchical relationship reasoning module to achieve information complementarity between tasks, calculate the channel weights independently for each task branch, explicitly model complex interaction relationships, and finally use FFN to predict and output the final prediction result.
[0014] In an embodiment of the present invention, step S1 includes:
[0015] S11: Use a 3-layer DLA network to extract low-level features, middle-level features, and high-level features from the image i to be detected and represent them as low-level feature map f low , middle-level feature map f mid and high-level feature map f high , respectively represent extracting the low-level feature map, middle-level feature map, and high-level feature map from the image i to be detected,
[0016]
[0017] S12: For each feature map f ∈ {f low , f mid , f high}, use the ECA module to generate channel weights w through 1D convolution with a kernel size of k, and multiply them with the original feature map f channel by channel to obtain the feature map f eca after lightweight channel weight optimization. Then, adjust the feature resolution of the feature map f eca to obtain the low-level feature map to be stitched the middle-level feature map to be stitched and the high-level feature map to be stitched Wherein:
[0018] w = Sigmoid(Conv1D k (GAP(f))),
[0019]
[0020] Wherein, GAP represents channel compression, Conv1D k represents 1D convolution with a kernel size of k, sigmoid represents the sigmoid function, represents per-channel multiplication,
[0021] S13: Stitch the low-level feature map to be stitched the middle-level feature map to be stitched and the high-level feature map to be stitched in the channel dimension to obtain the stitched feature map f cat , concat represents the stitching operation,
[0022]
[0023] S14: Perform cascaded convolution fusion on the stitched feature map f cat to output the fused feature map F fused , F fuse1 、F fuse2 are intermediate process results, Conv represents the convolution operation,
[0024] F fuse1 = Conv3×3(f cat ),
[0025] F fuse2 = Conv3×3(F fuse1 ) + F fuse1 ,
[0026] F fused = Conv1×1(F fuse2 ).
[0027] In an embodiment of the present invention, step S2 includes:
[0028] S21: Enhance the fused feature map F fused ∈R H×W×C through position encoding to obtain the enhanced feature map X in , F fused +PE represents enhancing the fused feature map F fused through position encoding, H, W, and C respectively represent the fused feature map F fusedThe height, width, and number of channels,
[0029] X in = F fused + PE,
[0030] S22: Use a 3×3 depthwise separable convolution to perform a local convolution operation on the feature map X in to extract local detail features F local , and divide F local into P×P small blocks and flatten them into a sequence X patches ∈ R N×d , where, d = P 2 × C_outC_out is the number of output channels in the depthwise separable convolution;
[0031] S23: Perform self-attention operations on each block in the sequence X patches ∈ R N×d using the following formula:
[0032]
[0033] SelfAtt represents the Self-Attention self-attention mechanism, Q, K, and V represent query, key, and value respectively. Q, K, and V are matrices generated by linearly transforming the input sequence. T represents transpose, and W low ∈ R d×r (r << d) is a low-rank projection matrix, and the number of attention heads of SelfAtt is 4 heads.
[0034] S24: Reconstruct each block processed by S23 into a spatial feature F global ∈ R H×W×C ;
[0035] S25: Construct a multi-task decoder module. The multi-task decoder module includes a MobileViT encoder and a multi-task decoder. The MobileViT encoder performs the following functions:
[0036] Fuse the spatial feature F global and the local detail feature F local weightedly to output the global context feature X, where α is a learnable parameter, the initial value of α is 0.5 and is automatically optimized through training, and ⊙ represents channel-wise multiplication.
[0037] X = α ⊙ F local + (1 - α) ⊙ F global ;
[0038] S26: The multi-task decoder has three branches, which respectively correspond to three sub-tasks of person detection, object detection, and interaction classification. Each branch of the multi-task decoder consists of L decoder layers, and each decoder layer contains a self-attention layer, a cross-attention layer, and a Transformer decoder inside. The cross-attention layers in each branch share parameters, and the global context feature X and the independent learnable query tokens Q t are input into the corresponding branch of the multi-task decoder. In the L-th layer of the t-th branch, the multi-task decoder updates the output of the previous layer of the t-th branch by attending to the global context feature X to generate task-specific features, where t ∈ (H, O, I), l = 1, 2... L, and H, O, and I respectively correspond to the person detection branch, the object detection branch, and the interaction classification branch. The cross-attention layers share the Key-Value projection matrix and only retain the independent query projection:
[0039]
[0040] ShareCrossAtt is the multi-head attention mechanism, and W kv is the shared parameter, is the task-specific parameter, d k is the attention head dimension,
[0041] The query tokens are updated through shared attention between L decoder layers,
[0042]
[0043] LayerNorm represents the layer normalization operation, represents the learnable query tokens updated in the (l - 1)-th layer of the t-th branch, represents the output of the decoder layer in the l-th layer of the t-th branch;
[0044] S27: Output the person feature the object feature and the interaction feature
[0045] In an embodiment of the present invention, step S3 includes:
[0046] S31: Generate the ternary relation feature from the person feature the object feature and the interaction feature through the multi-layer perceptron MLP
[0047]
[0048] S32: Generate text embeddings for interaction categories using the CLIP pre-trained text encoder, align the text embeddings with visual features, and generate updated features that fuse cross-modal semantic information CrossAtt represents cross-attention, is the query projection weight for F HOI of, and are the key and value projection weights for the text embedding T respectively, d k is the dimension of the attention head,
[0049]
[0050] S33: Perform self-attention modeling on the internal features of single-person, object, and interaction instances to capture internal dependencies U l ,
[0051]
[0052] SelfAtt represents self-attention,
[0053] S34: Use cross-attention to embed U l into to obtain the embedding result F′ l , represents the l-th layer of the updated feature F text-HOI ,
[0054]
[0055] S35: Construct human-object features human-interaction features F l HI and object-interaction features F l OI respectively through a multi-layer perceptron MLP,
[0056]
[0057] S36: Use self-attention to extract the relationships P HO between human-object features F HI , human-interaction features F OI and object-interaction features F l :
[0058]
[0059] S37: Embed P l into F′ l through cross-attention to obtain the embedding result F″ l ,
[0060] F″ l = CrossAtt((F′ l , P l ))
[0061] S38: Fuse the global context feature X with the embedding result F″ l to obtain the relational context y l ,
[0062] y l = CrossAtt(F″ l , X).
[0063] In an embodiment of the present invention, step S4 includes:
[0064] S41: Concatenate the person feature F H , the object feature F O and the interaction feature F I with the relational context y l to generate the channel attention weight β through MLP, σ is the Sigmoid function, and the output weight β ∈ [0, 1] C , C is the number of channels,
[0065]
[0066] S42: Perform channel-level weighting on the concatenated features using the weight β to obtain the weighted result ⊙ is the element-wise multiplication,
[0067]
[0068] S43: Input the output of the L-th layer decoder of the person detection branch into the FFN to calculate the person bounding box coordinates
[0069] Input the output of the L-th layer decoder of the object detection branch into the FFN to calculate the object bounding box coordinates
[0070] FFN hbox denotes the FFN for predicting the bounding box coordinates of the person, and FFN obox denotes the FFN for predicting the bounding box coordinates of the object;
[0071] S44: Classify the object to obtain the classification result δ represents the softmax operation, and FFN oc denotes the FFN for object classification,
[0072]
[0073] S45: Calculate the result of the cross-classification. σ represents the Sigmoid operation. σ is a multi-label classification activation function used to independently determine the existence of each category, and the output value is between [0, 1].
[0074] FFN ic Denotes the FFN for cross-classification.
[0075]
[0076] In an embodiment of the present invention, C_out is 64.
[0077] The present invention also provides an interactive action detection device based on multi-level features for performing the above method, which includes: an adaptive multi-level feature extraction module, an encoder-multi-task decoder module, a hierarchical relationship reasoning module, and a task-aware channel gating module, where:
[0078] The adaptive multi-level feature extraction module is configured to execute step S1;
[0079] The encoder-multi-task decoder module is configured to execute step S2;
[0080] The hierarchical relationship reasoning module is configured to execute step S3;
[0081] The task-aware channel gating module is configured to execute step S4.
[0082] The interactive action detection method and device provided by the present invention solve the following problems:
[0083] Aiming at the problem of insufficient extraction and utilization of interactive region features, the present invention proposes an adaptive multi-level feature extraction network. Each DLA (Deep Layer Aggregation) network layer may learn different feature representations, and these feature representations may focus on different aspects of the interactive region (such as shape, color, texture, etc.). In addition, ECA (EfficientChannel Attention) acts inside a single feature map to optimize the lightweight channel weights of each layer of features and enhance the response of key channels. And through the cascade fusion method, the information of the low layer, middle layer, and high layer is gradually accumulated, allowing the features to transition from local to global. Thereby improving the accuracy of global semantics and the preservation of details.
[0084] To address the problem of insufficient fine-grained relationship modeling, the present invention proposes a hierarchical collaborative reasoning network, which adopts three decoder branches corresponding to three subtasks: person detection, object detection, and interaction classification. After the decoder branches, a hierarchical relationship reasoning module is designed to hierarchically reason (from local to global) the features of people (FH), objects (FO), and interactions (FI) output by the three decoder branches. Through the attention interaction of unary relationships, pairwise relationships, and ternary relationships, multi-level contexts are gradually fused, and the CLIP pre-trained text encoder is used to align text embeddings with visual features, improving the accuracy of fine-grained interaction classification. In the task-aware channel gating module, the features output by the hierarchical relationship reasoning module are integrated with the context information of the person, object, and interaction branches to achieve information complementarity between tasks and explicitly model complex interaction relationships. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0086] Figure 1 Schematic diagram of an interaction action detection method based on multi-level features according to an embodiment of the present invention;
[0087] Figure 2 Schematic diagram of a multi-level feature extraction network according to an embodiment of the present invention;
[0088] Figure 3 Schematic diagram of the DLA network framework according to an embodiment of the present invention;
[0089] Figure 4 ECA module framework;
[0090] Figure 5 Transformer encoder-multi-task decoder module framework in step S2;
[0091] Figure 6 Multi-task decoder module framework and complete three-branch structure;
[0092] Figure 7 Hierarchical relationship reasoning module framework in step S3. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0093] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0094] The present invention combines the advantages of two-stage and single-stage. First, in the feature extraction stage, an adaptive multi-level feature extraction module is proposed to strengthen feature extraction, and then a lightweight MobileViT encoder is adopted. Its core lies in combining local convolution and a lightweight Transformer structure, which can not only efficiently extract local details but also model global dependencies at low cost. And the decoder has three branches, independently corresponding to the three subtasks of person detection, object detection, and interaction classification respectively. In addition, a hierarchical relationship reasoning module and a task-aware channel gating module are introduced on the basis of the single-stage framework to perform hierarchical reasoning on the features output by the decoder branches and filter task-related channels to suppress interference from irrelevant features. And cross-modal semantic guidance is carried out, that is, CLIP text embedding is aligned with visual features to support open-vocabulary interaction recognition and solve the problem that pure visual models are difficult to handle semantic ambiguity.
[0095] The current interactive action detection task mainly faces the following two major problems:
[0096] (1) The problem of insufficient feature extraction and utilization of interactive regions in images. In real-life scenarios, there are often some small targets and occluded targets, and it is more difficult to detect key details of these targets compared to normal targets. And some interactive actions often involve multiple individuals and objects, and the scenes of interactive actions are variable. Different lighting, perspectives, and scale changes will all affect feature extraction. Therefore, how to fully extract and utilize the environmental information and fine-grained target information in the image is a problem faced by current research.
[0097] (2) Insufficient fine-grained relationship modeling. Traditional methods only model simple associations of person-object pairs (such as spatial co-occurrence), lacking hierarchical reasoning of individual attributes (human postures), pairwise relationships (human-object spatial distances), and overall semantics (scene context), resulting in difficulty in distinguishing similar interactive actions (such as "riding" vs "leading" a horse).
[0098] For the first problem, the present invention proposes an adaptive multi-level feature extraction network that updates the weights to connect to the DLA network. Each layer of the DLA network may learn different feature representations, and these feature representations may focus on different aspects of the interaction area (such as shape, color, texture, etc.). By fusing these different feature representations, it helps the model better understand and recognize the complex information in the interaction area. For the second problem, the present invention proposes a hierarchical collaborative inference network. It adopts three decoder branches and makes full use of the unary, pairwise, and ternary relationship features of people, objects, and interactions for context exchange and dynamically filters the task-related channels to suppress the interference of irrelevant features. This helps the model quickly and accurately recognize and understand interaction actions in complex interaction scenarios.
[0099] Figure 1 Schematic diagram of an interaction action detection method based on multi-level features according to an embodiment of the present invention, which consists of Figure 1As shown in the figure, the present invention first performs feature extraction by an adaptive multi-level feature extraction module. The input image gradually passes through the DLA network to extract low-level (local details: edges, textures), middle-level (local + global: object parts), and high-level (global semantics: scene context) features respectively. In addition, ECA acts inside a single feature map to optimize the lightweight channel weights of each layer of features and enhance the response of key channels. And through the cascading fusion method, the information of the low-level, middle-level, and high-level is gradually accumulated, allowing the features to gradually transition from local to global, thereby improving the accuracy of global semantics and the preservation of details. This step solves the problem of insufficient feature expression ability of traditional single-layer features, improves feature diversity through progressive cascading fusion, and retains details and global information at the same time. Then, the spatial position information is embedded into the fused features, and after position encoding, it is sent into the Transformer encoder. The lightweight encoder design uses the MobileViT encoder to reduce the computational complexity of self-attention. With the help of the lightweight encoder, image features with context awareness are generated. Then it is sent to the multi-task decoder module. There are three task branches: parallel processing of human (H), object (O), and interaction (I) detection tasks, and each branch contains L layers of decoders. Each branch initializes a set of query tokens, which are gradually updated by paying attention to the encoder output features to obtain learnable query vectors. The decoder layer dynamically refines features, integrates image features, and generates task-specific features (F_H, F_O, F_I). This module realizes multi-task collaborative learning and avoids information isolation between tasks (such as the position of a person and interaction actions need to be jointly inferred). The features of human (F_H), object (F_O), and interaction (F_I) output by the three branches of the decoder are sent to the hierarchical relationship reasoning module for hierarchical reasoning (from local to global). Through the attention interaction of unary relationship, pairwise relationship, and ternary relationship, multi-level context is gradually fused, and a pre-trained text encoder such as CLIP is used to generate a text embedding T for the interaction category, and the text embedding is aligned with the visual features to improve the accuracy of fine-grained interaction classification. In the task-aware channel gating module, the features output by the hierarchical relationship reasoning module are further integrated with the context information of the human, object, and interaction branches of the decoder to achieve information complementarity between tasks, and the channel weights are independently calculated for each task branch. Explicitly model complex interaction relationships (such as "human-knife-cut" needs to consider both the human's knife-holding action and the presence of the knife), and finally use the FFN to predict and output the final prediction results (person, interaction object, interaction category).
[0100] As Figure 1 shown, the present invention provides an interaction action detection method based on multi-level features, which includes:
[0101] S1: Extract low-level features, middle-level features, and high-level features from the image to be detected and represent them as low-level feature maps, middle-level feature maps, and high-level feature maps respectively. Perform lightweight channel weight optimization on the low-level feature maps, middle-level feature maps, and high-level feature maps respectively, and fuse the information of the low-level feature maps, middle-level feature maps, and high-level feature maps through a cascaded fusion method to obtain a fused feature map;
[0102] The goal of feature extraction is to extract useful information for tasks (such as classification, detection, and interaction recognition) from the original data (such as images). The essence of features is the high-level abstract representation of data. Deep learning features: Automatically learn the multi-level representation of data through neural networks (the convolutional layers of CNNs, the attention mechanisms of Transformers), gradually abstracting from low-level to high-level.
[0103] The multi-level feature extraction network designed in the present invention is as Figure 2 shown. The essence of the design is to utilize the multi-level outputs (low, middle, and high levels) of a single DLA-34 network, enhance channel sensitivity through the ECA module, and then fuse these hierarchical features in a cascaded manner.
[0104] S2: Obtain local detail features from the fused feature map and perform global modeling to obtain global features. The global features are then fused with the local detail features to generate global context features, and perform person detection, object detection, and interaction detection in parallel through multi-task branches, and output person features, object features, and interaction features.
[0105] S3: Through the attention interaction of unary relationships, pairwise relationships, and ternary relationships, gradually fuse multi-level contexts, use the pre-trained text encoder to generate text embeddings for interaction categories, align the text embeddings with visual features, and improve the accuracy of fine-grained interaction classification;
[0106] S4: Integrate the context information of the features output by the hierarchical relationship reasoning module to achieve information complementarity between tasks, calculate the channel weights independently for each task branch, explicitly model complex interaction relationships, and finally use the FFN to predict and output the final prediction result.
[0107] In an embodiment of the present invention, step S1 includes:
[0108] S11: Use a 3-layer DLA network to extract low-level features, middle-level features, and high-level features from the image I to be detected and represent them as low-level feature map f low 、middle-level feature map f mid and high-level feature map f high , respectively represent extracting the low-level feature map, middle-level feature map, and high-level feature map from the image I to be detected.
[0109]
[0110] The DLA network framework is as Figure 3 shown. The DLA network is selected because it enhances the feature reuse ability through dense connections and hierarchical aggregation, fuses feature maps at different stages according to rules, and can fuse cross-layer information more efficiently compared to ResNet or FPN. Selecting a three-layer structure (low, middle, and high layers) meets the dual requirements of the HOI task for details (human poses, object shapes) and semantics (interaction relationships). The low layer in the three-layer structure captures local detail information, the middle layer mixes local and global information, and the high layer captures global semantic information.
[0111] In addition, due to the large semantic gap between the low-layer and high-layer features, direct splicing may lead to information conflicts or easily cause the weakening of shallow features under the dominance of deep features. The present invention adopts the method of ECA (Efficient Channel Attention) + progressive cascading fusion. ECA can solve the problems of lack of semantics in shallow features and loss of details in deep features, and improve the feature discriminability. As Figure 4 shown in the ECA module framework, it acts inside a single feature map, adjusts weights in the channel dimension, highlights important channels, but it does not directly change the fusion method between different hierarchical features. Progressive cascading fusion acts between multi-layer features. It reduces the sudden change of low-layer and high-layer features and makes the information flow smoother by gradually accumulating information.
[0112] The ECA module performs lightweight channel attention modeling on the features of each layer through 1D convolution, optimizes the weights of each layer of features, and at the same time avoids information loss caused by channel dimensionality reduction, with less computational complexity. For the feature channels of different hierarchies, weights are adaptively assigned. For example, in low-layer features, channels related to edges and textures are enhanced; in high-layer features, background interference channels are suppressed to highlight the interaction subjects. Figure 4 In it, the convolution kernel size k can be determined adaptively, C represents the channel dimension, and γ and b are set to 2 and 1, for example.
[0113] S12: For each feature map f ∈ {f low , f mid , f high}, use the ECA module to generate channel weights w through 1D convolution with a kernel size of k, and multiply them with the original feature map f channel by channel to obtain the feature map f eca optimized with lightweight channel weights. Then, adjust the feature resolution of the feature map f eca to obtain the low-layer feature map to be spliced the middle-layer feature map to be spliced and the high-layer feature map to be spliced Wherein:
[0114] w = Sigmoid(Conv1D k (GAP(f))),
[0115]
[0116] wherein, GAP represents channel compression, Conv1D k represents 1D convolution with a kernel size of k, sigmoid represents the sigmoid function, represents per-channel multiplication,
[0117] Finally, first retain the original information of each layer through channel concatenation, and then perform hierarchical fusion (such as low layer → middle layer → high layer) to avoid detail blurring caused by direct addition or convolution. The cascaded fusion gradually integrates low-level details, middle-level structures, and high-level semantics. Compared with single fusion (such as FPN), it can more accurately maintain spatial details and global consistency.
[0118] S13: Concatenate the low-level feature map to be concatenated the middle-level feature map to be concatenated and the high-level feature map to be concatenated in the channel dimension to obtain the concatenated feature map f cat , concat represents the concatenation operation,
[0119]
[0120] S14: Perform cascaded convolution fusion on the concatenated feature map f cat to output the fused feature map F fused , F fuse1 、F fuse2 are intermediate process results, Conv represents the convolution operation (usually 3×3 convolution + BatchNorm + ReLU).
[0121] F fuse1 = Conv3×3(f eat ),
[0122] F fuse2 = Conv3×3(F fuse1 ) + F fuse1 ,
[0123] F fused = Conv1×1(F fuse2 ).
[0124] Although the proposed feature extraction network can already extract multi-level (low-level, middle-level, high-level) features, these features usually focus on local information and progressive abstraction. The convolutional network works within a local receptive field and can gradually obtain higher-level semantic information, but its global dependence modeling ability is limited. The MobileViT encoder, on the other hand, explicitly models global information through Transformer operations, capturing the correlations between distant pixels, which is difficult for traditional convolutions to directly extract. The design of MobileViT is to effectively fuse local details (captured by CNN) with global context (provided by Transformer), and the two complement each other. Even with multi-level features, adding MobileViT can enhance the global information of these features, making the final feature representation more discriminative and robust. Moreover, the normal Transformer encoder mainly captures long-range dependencies through multiple layers of self-attention, but the computational cost is relatively high. Now, the MobileViT encoder is adopted, the core of which is to combine local convolution and lightweight Transformer structure, which can not only efficiently extract local details but also model global dependencies at low cost.
[0125] In the commonly used Transformer-based HOI detection methods, most of the decoders are single-branch and double-branch. The single-branch method updates the token set through a single decoder and is responsible for multiple subtasks, namely human detection, object detection, and interaction classification, and then directly predicts the HOI instance using FFN; the double-branch method appears later to solve the problem of limited detection ability of the single-branch under multiple tasks, that is, an independent decoder branch is adopted, one branch is responsible for detecting the pair of people, and the other branch is responsible for classifying the interaction behavior. Based on the double-branch decoder, the present invention decouples the branch of pair-of-people detection, making the decoder have three branches, corresponding to the three subtasks of human detection, object detection, and interaction classification respectively.
[0126] Step S2 includes a Transformer encoder and a multi-task decoder branch. The long-range dependencies (such as the interaction between a person and a distant object) are captured through the encoder for global context modeling. The person, object, and interaction are detected in parallel through the multi-task branch, avoiding the error accumulation of step-by-step detection. The framework of the Transformer encoder-multi-task decoder module is as Figure 5 shown.
[0127] Step S2 replaces the traditional Transformer encoder with a MobileViT encoder. Such lightweight networks effectively reduce the cost of self-attention calculations. The MobileViT encoder can utilize a strategy that combines local convolution and lightweight Transformer modules to globally model local features through Unfold / Fold operations, and then fuse them with the original local features to generate stronger global context representations. This design does not simply repeat existing multi-level feature extraction but globally enhances local information. The two complement each other, helping to improve the performance of the overall model in complex scenarios. Finally, the global context feature X is output, which contains spatial-semantic joint information.
[0128] In an embodiment of the present invention, step S2 includes:
[0129] S21: Input processing:
[0130] The fused feature map F fused ∈R H×W×C is enhanced through position encoding to obtain the enhanced feature map X in , F fused +PE represents performing position encoding enhancement on the fused feature map F fused , where H, W, and C respectively represent the height, width, and number of channels of the fused feature map F fused ;
[0131] X in = F fused +PE,
[0132] S22: Local convolution and chunking:
[0133] Perform local convolution operations on the feature map x in using a 3×3 depthwise separable convolution to extract local detailed features F local , and divide F local into P×P small chunks and flatten them into a sequence X patches ∈R N×d , where d = P [[ID=X]] 2 ×C_out, where C_out is the number of output channels in the depthwise separable convolution (for example, C_out is 64);
[0134] S23: Lightweight Transformer layer:
[0135] Each chunk undergoes global context modeling through a lightweight Transformer layer. This layer uses a simplified self-attention mechanism and has parameter and computational complexity pruning in its implementation (such as reducing the number of attention heads and using low-rank approximation) to reduce the computational complexity:
[0136] For each block in sequence X patches ∈R N×d perform the self-attention operation of the following formula respectively:
[0137]
[0138] Self Att represents the Self-Attention self-attention mechanism. Q, K, and V represent query, key, and value respectively. Q, K, and V are matrices generated by linearly transforming the input sequence. Q: represents the position (or token) that needs to be focused on currently, used to "query" the relevance with other positions. K: represents the "keys" of all positions, used to be queried to calculate similarity. V: represents the actual feature values, used for weighted aggregation according to the attention weights. T represents transpose, and W low ∈R d×r (r << d) is a low-rank projection matrix. The number of attention heads of Self Att is 4, and W low ∈R d×r (r << d) is a low-rank projection matrix, and its specific value is determined in the following way:
[0139] 1. Random initialization: At the beginning of training, the value of W low is usually randomly sampled from a normal distribution.
[0140] 2. End-to-end learning: During training, optimize the value of W low through backpropagation and gradient descent to minimize the task loss,
[0141] 3. Low-rank constraint: Since r << d, the rank of matrix W low is explicitly restricted, reducing the number of parameters from d 2 to d × r, thereby reducing the computational complexity and the amount of calculation.
[0142] In addition, the reduction in the number of attention heads will be reflected in the value of d k and the number of projection matrices. Increase in single-head dimension: Assuming the total dimension d is fixed, the single-head dimension d k will increase from d / 8 to d / 4 (for example, if d = 512, then d k changes from 64 to 128). Parameter pruning corresponds to a reduction in the number of projection matrices (from 8 groups to 4 groups), and the total number of parameters is reduced from 8 × 3 × d × d k to 4 × 3 × d × d k .
[0143] S24: Feature reconstruction and global fusion: Each block after being processed by Transformer is restored to the original spatial dimension through a reconstruction operation (Fold / Unfold). Unfolding:
[0144] Reconstruct each block after S23 processing into the spatial feature F global ∈R H×W×C ;
[0145] S25: Construct a multi-task decoder module. The multi-task decoder module includes a MobileViT encoder and a multi-task decoder. The MobileViT encoder performs the following functions:
[0146] Fuse the spatial feature F global with the local detail feature F local weightedly to output the global context feature X. Here, α is a learnable parameter, the initial value of α is 0.5 and is automatically optimized through training. The value range of α is constrained between [0,1] by the Sigmoid function to ensure the stability of the fusion. ⊙ represents element-wise multiplication (multiplying the corresponding elements of two tensors) and is used to achieve the adaptive weighted fusion of the local feature F local and the global feature F global to finally output the feature X that retains both details and global context.
[0147] X = α ⊙ F local + (1 - α) ⊙ F global ;
[0148] The global context feature X contains spatial-semantic joint information and retains both local details and global dependencies.
[0149] S26: Figure 6 For the multi-task decoder module framework and the complete three-branch architecture, as Figure 6 shown, the multi-task decoder has 3 branches. The 3 branches correspond to three subtasks of person detection, object detection, and interaction classification respectively. Each branch of the multi-task decoder consists of L decoder layers. Each decoder layer contains a self-attention layer, a cross-attention layer, and a Transformer decoder inside. The cross-attention layers in each branch share parameters. Input the global context feature X and the independent learnable query token Q t of each branch into the corresponding branch of the multi-task decoder. In the L-th layer of the t-th branch, the multi-task decoder updates the output of the previous layer of the t-th branch by focusing on the global context feature X to generate task-specific features, t ∈ (H, O, I), l = 1, 2... L, where H, O, and I correspond to the person detection branch, the object detection branch, and the interaction classification branch respectively. The cross-attention layers share the Key-Value projection matrix and only retain the independent query projection:
[0150]
[0151] ShareCrossAtt is the multi-head attention mechanism, and W kv is the shared parameter, is the task-specific parameter, d k is the attention head dimension,
[0152] The sharing of parameters in the cross-attention layer reduces redundant calculations while ensuring information complementarity among the three tasks of person, object, and interaction, achieving information interaction and efficient feature fusion.
[0153] The query tokens are updated through shared attention among L decoder layers,
[0154]
[0155] LayerNorm represents the layer normalization operation, represents the learnable query token updated in the (l - 1)-th layer of the t-th branch, represents the output of the decoder layer in the l-th layer of the t-th branch;
[0156] Each task branch (H, O, I) receives independent learnable query tokens Q t (such as Q H , Q O , Q I ) at the initial layer of the decoder. These tokens are obtained through random initialization or task-related pre-training. In each layer of the decoder, Q t gradually updates itself through cross-attention interaction with the global features X output by the encoder to generate task-specific feature representations. represents the output of the decoder in the l-th layer (1 < l < L) of branch t (t ∈ (H, O, I)).
[0157] S27: Output the person features Object features And interaction features
[0158] represents the output of the decoder in the l-th layer (1 < l < L) of branch t (t ∈ (H, O, I)). represents the features output in the l-th layer of the person detection branch of the multi-task decoder, represents the features output in the l-th layer of the object detection branch of the multi-task decoder, represents the features output in the l-th layer of the interaction detection branch of the multi-task decoder.
[0159] To address the problem of difficult to capture complex interaction logic, step S3 explicitly models the complex dependencies of human-object-interaction by progressively embedding unary, pairwise, and ternary relationships. Additionally, traditional visual models rely solely on pure visual features and are susceptible to background interference or detail ambiguity. Introducing text information can compensate for this shortcoming, enabling the model to more clearly capture semantic information in interactions, improving classification accuracy and robustness. After initially obtaining the overall "human-object-interaction" features, step S3 introduces text semantic information, providing cross-modal guidance for subsequent self-attention and further pairwise relationship modeling, thereby better eliminating ambiguity (such as distinguishing "riding" from "leading" a horse) and enhancing the accuracy of fine-grained interaction classification. The framework of the hierarchical relationship reasoning module is as Figure 7 shown.
[0160] In an embodiment of the present invention, step S3 includes:
[0161] S31: Generate ternary relationship features for the human feature the object feature and the interaction feature
[0162]
[0163] through a multi-layer perceptron MLP CrossAtt represents cross-attention, is the query projection weight for F HOI and and are the key and value projection weights for the text embedding T respectively, d k is the dimension of the attention head,
[0164]
[0165] S33: Perform self-attention modeling on the internal features of single human, object, and interaction instances to capture the internal dependencies U l ,
[0166]
[0167] SelfAtt represents self-attention,
[0168] S34: Use cross-attention to embed U l into to obtain the embedding result F' l , represents the updated feature F text-HOIat the l-th layer,
[0169]
[0170] S35: Construct the human-object features human-interaction feature F l HI and the object-interaction feature F l OI ,
[0171]
[0172] S36: Use self-attention to extract the relationships P HO between the human-object feature F HI , the human-interaction feature F OI and the object-interaction feature F l :
[0173]
[0174] S37: Embed P l into F' l through cross-attention to obtain the embedded result F'' l ,
[0175] F'' l = CrossAtt(F' l , P l )
[0176] S38: Fuse the global context feature X with the embedded result F'' l to obtain the relationship context y l ,
[0177] y l = CrossAtt(F'' l , X).
[0178] Step S3 designs a hierarchical reasoning process that conforms to cognitive logic (from local to global), enhancing the fine-grained classification ability. Traditional methods compress multiple relationships into a single feature, resulting in information confusion (e.g., it is difficult to distinguish between "a person riding a horse" and "a person feeding a horse"). The design of Step S3 can strengthen the adaptability to long-tail data. For example, in rare interaction categories (such as "a person playing the harp"), explicitly modeling the relationship can alleviate overfitting caused by insufficient data. And to improve semantic understanding ability, by introducing a CLIP pre-trained text encoder, text embeddings describing interaction categories can be obtained, thus providing semantic priors for visual features. This cross-modal alignment can help the model distinguish those interaction situations that are prone to ambiguity only based on visual features (e.g., "a person holding a cup" may mean "drinking" or "cleaning"). Moreover, the text embeddings have good generalization ability, so even if certain interaction categories are not seen in the training data, text semantics can be used for reasoning.
[0179] Step S4 makes a theoretical innovation by combining the channel attention mechanism with multi-task learning, proposing a task-aware gating strategy, which is different from traditional global channel compression (such as SENet). Since each sub-task requires different context information for relationship reasoning, an MLP and each task-specific token are used to transmit the relationship context. The MLP is used for non-linear mapping and can better capture the complexity of the relationship context. And to select the context information required for each sub-task, the channel attention mechanism is further used, which can assign different attentions to the weights of different channels or features in order to select the information most relevant to the task.
[0180] Dynamically select the feature channels most relevant to a specific task (H / O / I) through channel attention, and use the attention weight α activated by sigmoid to achieve soft selection of feature channels (similar to a gating mechanism). Concatenate the task features (such as the person branch feature F H ) with the relationship context (y l ) and generate channel attention weights through an MLP to solve the lack of global information in isolated task reasoning and perform task-context joint modeling.
[0181] In an embodiment of the present invention, Step S4 includes:
[0182] S41: Concatenate the person feature F H , the object feature F O and the interaction feature F I with the relationship context y l , generate channel attention weights β through an MLP, where σ is the Sigmoid function, and the output weight β ∈ [0, 1] C , C is the number of channels,
[0183]
[0184] S42: Channel - level weighting is performed on the concatenated features using the weight β to obtain a weighted result ⊙ represents element - wise multiplication
[0185]
[0186] S43: The output of the L - th layer decoder of the person detection branch is input into the FFN to calculate the person bounding box coordinates
[0187] The output of the L - th layer decoder of the object detection branch is input into the FFN to calculate the object bounding box coordinates
[0188] FFN hbox The FFN representing the bounding box coordinates of the predicted person. The FFN obox The FFN representing the bounding box coordinates of the predicted object; FFN is Feed - Forward Networks. The macroscopic structures of these four FFNs are roughly similar, but their input processing, output layer design, and intermediate layer parameters (such as the number of layers, width, activation function) will be adjusted according to the task requirements
[0189] S44: Classify the object to obtain a classification result δ represents the softmax operation. The FFN oc represents the FFN used for object classification
[0190]
[0191] S45: Calculate the result of interaction classification. σ represents the Sigmoid operation. σ is a multi - label classification activation function used to independently judge whether each category exists, and the output value is between [0,1]
[0192] FFN ic represents the FFN used for interaction classification
[0193]
[0194] The design of step S4 enables different tasks to focus on different feature channels (e.g., person detection focuses on the human body pose channel, and interaction classification focuses on the body - object contact channel). Additionally, irrelevant channels (such as interference from background regions) are suppressed by the α weight, and the channel attention weight α is learnable. It can automatically strengthen key channels (such as focusing on the "person - ball" contact area to judge the "kick" action) through gradient update, and can adapt to different scenario requirements
[0195] The present invention also provides an interactive action detection device based on multi-level features for performing the above method, which includes: an adaptive multi-level feature extraction module, an encoder-multi-task decoder module, a hierarchical relationship reasoning module, and a task-aware channel gating module, where:
[0196] The adaptive multi-level feature extraction module is configured to perform step S1;
[0197] The encoder-multi-task decoder module is configured to perform step S2;
[0198] The hierarchical relationship reasoning module is configured to perform step S3;
[0199] The task-aware channel gating module is configured to perform step S4.
[0200] Summarize the present invention as described above:
[0201] 1. The present invention proposes an adaptive multi-level feature extraction module. This module utilizes the DLA network to enhance the feature reuse ability through dense connections and hierarchical aggregation, and fuses the feature maps at different stages according to rules. Low-level, middle-level, and high-level features are extracted respectively. Low-level: Capture local detail information. Middle-level: Mix of local and global. High-level: Capture global semantic information. Enhance channel sensitivity through the ECA module, which acts inside a single feature map, performs lightweight channel weight optimization on each layer of features, and enhances the response of key channels. And through the cascade fusion method, gradually accumulate the information of the low-level, middle-level, and high-level, allowing the features to gradually transition from local to global, and reducing the mutation of the low-level and high-level features.
[0202] 2. The present invention proposes a multi-task decoder module, which adopts three task branches to parallelly process the tasks of human, object, and interaction detection, avoiding the error accumulation of step-by-step detection.
[0203] 3. The present invention proposes a hierarchical relationship reasoning module. This module realizes information complementarity between tasks by progressively embedding unary, pairwise, and ternary relationships, explicitly models the complex dependencies of human-object-interaction, and breaks through the limitations of traditional flat interaction modeling. And utilizes pre-trained text encoders such as CLIP to align the text embedding with the visual features. This module can improve the accuracy of fine-grained interaction classification and enhance the discrimination of similar interaction actions.
[0204] 4. The present invention proposes a task-aware channel gating module, which combines the channel attention mechanism with multi-task learning and proposes a task-aware gating strategy. The features output by the hierarchical relationship reasoning module are further integrated with the context information of the human, object, and interaction branches of the decoder to achieve information complementarity between tasks. In addition, different tasks are made to focus on different feature channels, and irrelevant channels are suppressed by the α weight. Moreover, the channel attention weight α is learnable and can automatically reinforce key channels through gradient update, enabling it to adapt to the requirements of different scenarios.
[0205] With the diverse evolution of production and living scenarios (such as smart home collaboration, industrial robot interaction, etc.), human-object interaction detection still faces challenges: complex background interference, such as dense occlusion (e.g., multiple people overlapping), dynamic light changes (e.g., reflection / shadow), small targets (e.g., handheld micro-tools); insufficient fine-grained interaction discrimination, such as the difference between "riding" and "leading" a horse depends on subtle differences (hand gestures, object spatial positions). Existing human-object interaction detection methods have limitations. Therefore, the present invention proposes an interaction action detection technology based on multi-level features, which can effectively improve the above-mentioned problems and enhance the detection accuracy.
[0206] The present invention addresses the problem of complex background interference. Models based on single - level features (such as the last layer of ResNet) are difficult to distinguish targets from background noise, and relying on manually designed attention mechanisms has poor generalization performance. The present invention proposes an adaptive multi - level feature extraction network module. This module uses the DLA network for feature extraction: low - level (edge / texture), middle - level (component structure), and high - level (scene semantics) features are extracted independently. The ECA channel attention dynamically enhances the response of key regions. Through progressive cascading fusion, the information of the low - level, middle - level, and high - level is gradually accumulated, enabling the features to transition from local to global. Aiming at the problem of insufficient fine - grained interaction discrimination resulting in a high misclassification rate for similar interaction actions, the present invention proposes a hierarchical collaborative network, which includes three modules: a Transformer encoder - multi - task decoder module, a hierarchical relationship reasoning module, and a task - aware channel gating module. First, the Transformer encoder uses the MobileViT encoder, which combines local convolution and a lightweight Transformer structure, capable of efficiently extracting local details and low - cost modeling of global dependencies. The decoder adopts three - task branches: parallelly processing the detection tasks of humans (H), objects (O), and interactions (I), and then delivering the features of the output humans (F_H), objects (F_O), and interactions (F_I) to the hierarchical relationship reasoning module, progressively embedding unary, pairwise, and ternary relationships, explicitly modeling the complex dependencies of human - object - interaction, and performing a hierarchical reasoning process (from local to global). Moreover, after initially obtaining the overall "human - object - interaction" features, text semantic information is introduced, providing cross - modal guidance for subsequent self - attention and further pairwise relationship modeling, thereby better eliminating ambiguity and improving the accuracy of fine - grained interaction classification. Then, the task - aware channel gating module generates channel attention weights to solve the lack of global information in isolated task reasoning and perform task - context joint modeling.
[0207] The interaction action detection method and device based on multi - level features provided by the present invention solve the following problems:
[0208] Aiming at the problem of insufficient extraction and utilization of interaction region features, the present invention proposes an adaptive multi - level feature extraction network. Each DLA (Deep Layer Aggregation) network layer may learn different feature representations, which may focus on different aspects of the interaction region (such as shape, color, texture, etc.). In addition, ECA (EfficientChannel Attention) acts inside a single feature map, performing lightweight channel weight optimization on each layer of features to enhance the response of key channels. And through the cascading fusion method, the information of the low - level, middle - level, and high - level is gradually accumulated, enabling the features to transition from local to global, thereby improving the accuracy of global semantics and the preservation of details.
[0209] To address the problem of insufficient fine-grained relationship modeling, the present invention proposes a hierarchical collaborative inference network that uses three decoder branches corresponding to three subtasks: person detection, object detection, and interaction classification. After the decoder branches, a hierarchical relationship inference module is designed to hierarchically infer (from local to global) the features of people (FH), objects (FO), and interactions (FI) output by the three decoder branches. Through attention interactions of unary, pairwise, and ternary relationships, multi-level contexts are gradually fused, and the CLIP pre-trained text encoder is used to align text embeddings with visual features, improving the accuracy of fine-grained interaction classification. In the task-aware channel gating module, the features output by the hierarchical relationship inference module are integrated with the context information of the person, object, and interaction branches to achieve information complementarity between tasks and explicitly model complex interaction relationships.
[0210] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0211] Those of ordinary skill in the art can understand that the modules in the device in the embodiment can be distributed in the device of the embodiment as described in the embodiment, or can be correspondingly changed and located in one or more devices different from this embodiment. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An interactive action detection method based on multi-level features, characterized in that Including: S1: Extract low-level features, middle-level features, and high-level features from the image to be detected and represent them as low-level feature maps, middle-level feature maps, and high-level feature maps respectively. Perform lightweight channel weight optimization on the low-level feature maps, middle-level feature maps, and high-level feature maps respectively, and fuse the information of the low-level feature maps, middle-level feature maps, and high-level feature maps through a cascaded fusion method to obtain a fused feature map; S2: Obtain local detail features from the fused feature map and perform global modeling to obtain global features. The global features are then fused with the local detail features to generate global context features. Perform person detection, object detection, and interaction detection in parallel through multi-task branches, and output person features, object features, and interaction features. S3: Through the attention interaction of unary relationships, pairwise relationships, and ternary relationships, gradually fuse multi-level contexts, use a pre-trained text encoder to generate text embeddings for interaction categories, align the text embeddings with visual features, and improve the accuracy of fine-grained interaction classification; S4: Integrate the context information of the features output by the hierarchical relationship reasoning module to achieve information complementarity between tasks, calculate the channel weights independently for each task branch, explicitly model complex interaction relationships, and finally use FFN to predict and output the final prediction result.
2. The interactive action detection method based on multi-level features according to claim 1, wherein Step S1 includes: S11: Extract low-level features, middle-level features, and high-level features from the image I to be detected using a 3-layer DLA network, and represent them as low-level feature map f low , middle-level feature map f mid , and high-level feature map f high , respectively represent the extraction of the low-level feature map, middle-level feature map, and high-level feature map from the image I to be detected S12: For each feature map f ∈ {f low , f mid , f high}, use the ECA module to generate channel weights w through 1D convolution with a kernel size of k, and multiply them with the original feature map f channel by channel to obtain the feature map f eca with optimized lightweight channel weights. Then, adjust the feature resolution of the feature map f eca to obtain the low-level feature map to be concatenated the middle-level feature map to be concatenated and the high-level feature map to be concatenated where: w = Sigmoid(Conv1D k (GAP(f))) Among them, GAP represents channel compression, and Conv1D k represents 1D convolution with a kernel size of k, and sigmoid represents the sigmoid function. represents per-channel multiplication. S13: Concatenate the low-level feature map to be stitched with the middle-level feature map to be stitched and the high-level feature map to be stitched in the channel dimension to obtain the concatenated feature map f cat , where concat represents the concatenation operation S14: Cascade convolution is performed on the stitched feature map f cat to fuse features and output the fused feature map F fused , F fuse1 , F fuse2 are intermediate process results, and Conv represents the convolution operation F fuse1 = Conv3×3(f cat ) F fuse2 = Conv3×3(F fuse1 ) + F fuse1 , F fused = Conv1×1(F fuse2 )。 3. The interactive action detection method based on multi-level features according to claim 1, wherein Step S2 includes: S21: Enhance the fused feature map F fused ∈R H×W×C through positional encoding to obtain the enhanced feature map X in , where F fused + PE represents enhancing the fused feature map F fused through positional encoding. H, W, and C respectively represent the height, width, and number of channels of the fused feature map F fused . X in = F fused + PE, S22: Perform local convolution operation on the feature map X using a 3×3 depthwise separable convolution in to extract local detailed features F local , and divide F local into P×P small blocks and flatten them into a sequence X patches ∈R N×d , where d = P 2 × C_out, where C_out is the number of output channels in the depthwise separable convolution; S23: Perform the self-attention operation of the following formula on each block in the sequence X patches ∈R N×d respectively: SelfAtt represents the Self-Attention mechanism. Q, K, and V represent query, key, and value respectively. Q, K, and V are matrices generated by linearly transforming the input sequence respectively. T represents transpose, and W low ∈R d×r (r << d) is a low-rank projection matrix. The number of attention heads of SelfAtt is 4 heads. S24: Reconstruct each block after S23 processing into the spatial feature F global ∈R H×W×C ; S25: Construct a multi-task decoder module, which includes a MobileViT encoder and a multi-task decoder. The MobileViT encoder performs the following functions: Fuse the spatial feature F global with the local detail feature F local through weighted fusion to output the global context feature X, where α is a learnable parameter with an initial value of 0.5 and is automatically optimized through training, and ⊙ represents channel-wise multiplication X = α ⊙ F local + (1 - α) ⊙ F global ; S26: The multi-task decoder has three branches, which respectively correspond to three subtasks of person detection, object detection, and interaction classification. Each branch of the multi-task decoder consists of L decoder layers. Each decoder layer contains a self-attention layer, a cross-attention layer, and a Transformer decoder. The cross-attention layers in each branch share parameters. The global context feature X and the independent learnable query tokens Q of each branch are input into the corresponding branches of the multi-task decoder. In the L-th layer of the t-th branch, the multi-task decoder updates the output of the previous layer of the t-th branch by attending to the global context feature X. t To generate task-specific features, where t ∈ (H, O, I), l = 1, 2... L, and H, O, I respectively correspond to the person detection branch, the object detection branch, and the interaction classification branch. The cross-attention layers share the Key-Value projection matrix and only retain the independent query projection: The global context feature X and the independent learnable query tokens Q of each branch are input into the corresponding branches of the multi-task decoder. In the L-th layer of the t-th branch, the multi-task decoder updates the output of the previous layer of the t-th branch by attending to the global context feature X. ShareCrossAtt is the multi-head attention mechanism, and W kv is the shared parameter, is the task-specific parameter, and d k is the attention head dimension, The query tokens are updated through shared attention among L decoder layers. LayerNorm represents the layer normalization operation, represents the learnable query token updated at the (l-1)-th layer of the t-th branch, represents the output of the decoder layer at the l-th layer in the t-th branch; S27: Output person characteristics Object characteristics and interaction characteristics 4. The interactive action detection method based on multi-level features according to claim 1, characterized in that Step S3 includes: S31: Generate the triple relationship features from the person characteristics object characteristics and interaction characteristics through a multi-layer perceptron MLP S32: Use the CLIP pre-trained text encoder to generate text embeddings for the interaction category, align the text embeddings with the visual features, and generate updated features that fuse cross-modal semantic information CrossAtt represents cross-attention, is the query projection weight for F HOI and and are the key and value projection weights for the text embedding T respectively, and d k is the dimension of the attention head, S33: Perform self-attention modeling on the internal features of a single person, object, or interaction instance to capture the internal dependencies U of the individual l , SelfAtt represents self-attention. S34: Embed U using cross-attention l into to obtain the embedding result F l ′, representing the l-th layer of the updated feature F text-HOI S35: Construct human-object features, human-interaction features, and object-interaction features respectively through a multi-layer perceptron (MLP). human-interaction features and object-interaction features S36: Extract the human-object feature F, HO human-interaction feature F, HI and object-interaction feature F OI and the relationship P l between each pair of them: S37: Embed P into F' through cross-attention to obtain the embedding result F'' l l l F l ″ = CrossAtt(F l ′, P l ) S38: Fuse the global context feature X with the embedding result F l to obtain the relational context y l , y l = CrossAtt(F l ″, X).
5. The interactive action detection method based on multi-level features according to claim 1, wherein Step S4 includes: S41: Concatenate the human feature F H , the object feature F O and the interaction feature F I with the relationship context y l , generate the channel attention weight β through MLP, σ is the Sigmoid function, and the output weight β ∈ [0, 1] C , C is the number of channels S42: Channel-level weighting is performed on the spliced features using the weight β to obtain a weighted result ⊙ represents element-wise multiplication S43: Input the output of the decoder of the L-th layer of the person detection branch into the FFN to calculate the person boundary coordinates The output of the decoder of the L-th layer of the object detection branch is input into the FFN to calculate the object boundary coordinates FFN hbox The FFN representing the bounding box coordinates of the predicted person, FFN obox The FFN representing the bounding box coordinates of the predicted object; S44: Classify the object to obtain the classification result δ represents the softmax operation, FFN oc represents the FFN for object classification S45: Calculate the result of interaction classification. σ represents the Sigmoid operation. σ is a multi-label classification activation function used to independently judge whether each category exists, and the output value is between [0,1]. FFN ic Indicates the FFN for interactive classification 6. The interactive action detection method based on multi-level features according to claim 3, characterized in that C_out is 64.
7. An interactive action detection device based on multi-level features, for performing the method according to any one of claims 1-6, characterized in that, Including: An adaptive multi-level feature extraction module, an encoder-multi-task decoder module, a hierarchical relationship reasoning module, and a task-aware channel gating module, where: The adaptive multi-level feature extraction module is configured to perform Step S1; The encoder-multi-task decoder module is configured to perform Step S2; The hierarchical relationship reasoning module is configured to perform Step S3; The task-aware channel gating module is configured to perform Step S4.
Citation Information
Cited By
Space-time decoupling sentiment analysis method and system based on multi-modal data
CN120930072A
Demand analysis and dynamic inquiry method and system based on cognitive reasoning
CN121094137A
Text embedding method and system based on large model intermediate layer fusion
CN121168473A
A text embedding method and system based on large model intermediate layer fusion
CN121168473B
Visual defect detection method based on multi-level Transform
CN121329967A