Character interaction detection method based on Mama architecture
By adopting a character interaction detection method based on Mamba architecture in character interaction detection, combining CEM units and detection context communication modules, a comprehensive progressive learning strategy is designed, and the existing methods are solved inadequate performance when identifying rare and complex interaction categories, and efficient and accurate character interaction detection is achieved.
Patent Information
- Application Number
- CN202510261091.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing character interaction detection method based on Transformer architecture has insufficient performance problems when identifying rare and complex interaction categories, and the model training time is long and the training cost is high, making it difficult to meet the practical application needs.
Using a character interaction detection method based on Mamba architecture, a novel comprehensive progressive learning strategy is designed to extract more expressive feature representations through global feature encoding module, multi-view feature aggregation module, human body branch, object branch and interaction branch, combined with CEM unit and detection context propagation module.
With small parameters, efficient end-to-end detection tasks are achieved, with better recognition performance, especially on HICO-DET and V-COCO datasets, which outperform existing methods, and can effectively identify rare and complex interaction categories.
Smart Images

Figure CN120107897A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the fields of computer vision, target detection and human interaction detection, and specifically is a human interaction detection method based on Mamba architecture. Background Art
[0002] Human interaction detection aims to detect human-object pairs that interact in an image and identify the types of interactions between them. Due to its importance in scene understanding and relational reasoning, human interaction detection plays a key role in a wide range of applications such as image retrieval, robotics, and visual question answering. Therefore, building an efficient and accurate human interaction detection method will play a positive role in the development of computer vision technology and promoting computer perception of complex scenes in the real world.
[0003] Previous methods for detecting human interaction can be divided into two paradigms: convolutional neural network-based and Transformer-based. Limited by the local receptive field of the convolution kernel, methods based on convolutional neural networks have difficulty identifying complex interaction categories. However, due to the advantage of the Transformer architecture in modeling long-distance interactions, methods based on the Transformer architecture have achieved significant performance improvements, thus becoming the mainstream method for current human interaction detection tasks.
[0004] Existing methods based on the Transformer architecture can be divided into three types according to the number of decoder branches: single-branch, dual-branch, and triple-branch methods. However, these methods all have some problems: the single-branch method uses a decoder to simultaneously handle the three subtasks of human detection, object detection, and interaction classification. Although the structure is simple, it is difficult for a single decoder to achieve a good performance trade-off between the three subtasks; the dual-branch method further decomposes human interaction detection into two subtasks: instance detection (human detection and object detection) and interaction classification, and uses two decoders to process these two subtasks respectively. Although this method achieves the decoupling between target detection and interaction classification, there is still a coupling relationship between the two subtasks of human detection and object detection, which restricts the possibility of further performance improvement; the three-branch method completely decouples the three branches of human, object, and interaction, so that each branch explicitly uses different parameters for the three subtasks. However, due to the lack of effective prior knowledge, the interaction branch requires a long training time and converges slowly. In addition, the additional decoder branch parameters based on the Transformer architecture bring huge training costs, which restricts the further development of human interaction detection technology.
[0005] At the same time, current methods ignore the need to capture more advanced visual expressions to identify rare and complex interaction categories, resulting in the captured interaction semantics being not comprehensive and rich enough, limiting the model's ability to further discover and identify more human interaction instances, making it difficult to meet the needs of actual applications. Summary of the invention
[0006] The present invention aims to address the deficiencies of the above-mentioned prior art and proposes a human interaction detection method based on the Mamba architecture, in order to capture high-level and comprehensive visual semantics to drive the detection of interacting human-object pairs, alleviate the recognition problem of rare and complex interaction categories, thereby helping to discover more human interaction detection instances and realizing human interaction detection more accurately and efficiently.
[0007] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme:
[0008] The character interaction detection method based on the Mamba architecture of the present invention is characterized in that it is performed according to the following steps:
[0009] Step 1: Obtain a dataset of human interaction detection images and preprocess them to obtain a preprocessed image dataset ,in, Indicates preprocessed human interaction detection images; N represents the total number of human interaction detection images; Indicates the length of the human interaction detection image, Indicates the width of the human interaction detection image. Indicates the number of channels of the human interaction detection image;
[0010] Step 2: Construct a human interaction detection network based on the Mamba architecture, including: a global feature encoding module, a multi-view feature aggregation module, a human branch, an object branch, and an interaction branch; wherein the global feature encoding module includes: a ResNet-50 model, a feature mapping unit, and a Transformer encoder; the multi-view feature aggregation module includes: two CEM units; the human branch and the object branch both include: two cascaded LoRA units, a Router unit for dynamic weighting, a detection context propagation module, and a fully connected layer; the interaction branch includes: two cascaded LoRA units, two Router units for dynamic weighting, and a fully connected layer;
[0011] Step 3: The global feature encoding module Process it and get A global visual representation ;
[0012] Step 4: The multi-view feature aggregation module Process it and get Shallow shared semantics and Deep shared semantics , and input them into the human body branch, object branch, and interaction branch for processing, and the human body branch outputs The human bounding box prediction results and its confidence , object branch output The object category prediction results , object bounding box prediction results and its confidence And the interactive branch output The interactive category prediction results and its confidence ;
[0013] Step 5: Based on , , And the real result of human bounding box , object bounding box real results , Object category real results , Interaction category real results , construct the total loss function of the human interaction detection network , and used to train the human interaction detection network until the total loss function Until convergence, the optimal human interaction detection model is obtained;
[0014] Step 6: In the reasoning process, The image to be tested Input into the optimal human interaction detection model for processing, and obtain The Individual human bounding box prediction results and its detection confidence , object bounding box prediction results and its detection confidence , Object category prediction results , Interaction category prediction results and its confidence ; , express The total number of prediction results in ;
[0015] Will , , After adding, as the The confidence score of the prediction result of the interaction between the characters is obtained. The first few prediction results with high confidence scores are taken as The final prediction result.
[0016] The character interaction detection method based on Mamba architecture described in the present invention is also characterized in that step 3 comprises:
[0017] Step 3.1: Input into the ResNet-50 model for processing and obtain the Initial visual features , represents the length of the visual feature, represents the width of the visual feature, The number of channels representing visual features;
[0018] Step 3.2: The feature mapping unit uses a convolution operation to The channel dimension is reduced to , and then use the flattening operator to expand the visual features after dimensionality reduction, and get the first Intermediate visual features ; The number of channels after the dimension is expressed;
[0019] Step 3.3: Encode the position and After superposition, it is input into the Transformer encoder for processing to obtain the first A global visual representation .
[0020] Further, the step 4 comprises:
[0021] Step 4.1: Random Generation query vectors to be learned ; and and Input into the first CEM unit for processing, and obtain the Shallow shared semantics ;in, Indicates the number of query vectors; Indicates the number of channels of the query vector;
[0022] Step 4.2: Input to the first LoRA unit of the human body branch for processing, and get the Individual low-level detection features ;
[0023] Will Input to the first LoRA unit of the object branch for processing, and get the Low-level detection features of objects ;
[0024] Step 4.3: and After superposition, and input them into the second CEM unit for processing to obtain Deep shared semantics ;
[0025] Step 4.4: Input the second LoRA unit of the human body branch for feature enhancement processing, and obtain the first Intermediate detection features of the individual body ;
[0026] Will The second LoRA unit of the input object branch performs feature enhancement processing to obtain the first Intermediate detection features of objects ;
[0027] Step 4.5: and Input the Router unit of the human body branch and perform adaptive dynamic weighting to obtain the first Advanced human detection features ;
[0028] Will and The Router unit in the input object branch performs adaptive dynamic weighting to obtain the first High-level object detection features ;
[0029] Step 4.6: The first Router unit in the interactive branch , , Perform adaptive dynamic weighting to generate the Low-level interaction features ;
[0030] Step 4.7: The first LoRA unit pair in the interactive branch Perform feature enhancement to obtain Intermediate Interaction Features ;
[0031] Step 4.8: The second Router unit in the interactive branch is , , , Perform adaptive dynamic weighting to obtain Advanced Interaction Features ;
[0032] Step 4.9: The second LoRA unit pair in the interactive branch Perform feature enhancement to generate Comprehensive interaction features ;
[0033] Step 4.10: The detection context propagation module in the human body branch and Process it and get Comprehensive detection characteristics of the individual ;
[0034] The detection context propagation module of the object branch is and Process it and get Comprehensive detection features of objects ;
[0035] Step 4.11: The fully connected layer in the human body branch Process and obtain The human bounding box prediction results And the detection confidence of the human bounding box , express Number of people;
[0036] The fully connected layer of the object branch is Process and obtain Object bounding box prediction results , detection confidence of the object bounding box And the object category prediction results , Indicates the number of detected objects. Indicates the number of categories of detected objects;
[0037] The fully connected layers in the interaction branch are Process and obtain The interactive category prediction results and interaction category confidence ; Indicates the number of detected interactions; Represents the number of categories of detected interactions.
[0038] Furthermore, each CEM unit includes: two self-Mamba blocks, two feature integration blocks and two cross-Mamba blocks; the first CEM unit in step 4.1 is obtained by the following process Shallow shared semantics :
[0039] Step 4.1.1. The input is fed into a self-Mamba block in the first CEM unit for feature enhancement, and the second The global visual state of the first stage and The global visual output of the first stage , thus obtaining The global visual output of the first stage As the One-stage global enhanced visual features :
[0040] (1)
[0041] (2)
[0042] In formula (1)-formula (2), Indicates The global visual state of the first stage is k=1. = , Indicates the total number of steps; express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix;
[0043] Step 4.1.2: Another self-Mamba block input to the first CEM unit is used for feature enhancement, and the second One-stage query state quantity and The output of the first-stage query , thus obtaining The output of the first-stage query As a first-stage enhancement query vector :
[0044] (3)
[0045] (4)
[0046] In formula (3)-formula (4), Indicates The first-stage query state quantity of the step, when k=1, let = , express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix;
[0047] Step 4.1.3: The first feature integration block uses equations (5) and (6) to obtain Two-stage enhancement of global visual features and two-stage enhanced query vector :
[0048] (5)
[0049] (6)
[0050] In formula (5)-(6), Norm represents a LayNorm layer;
[0051] Step 4.1.4: and Input to a cross Mamba block in the first CEM unit for Perform feature enhancement and use formula (7)-formula (8) to obtain the first The three-stage global visual state of the step and The three-stage global visual output of the step , thus obtaining The three-stage global visual output of the step As the Three-stage enhanced visual features :
[0052] (7)
[0053] (8)
[0054] In formula (7)-formula (8), Indicates The three-stage global visual state of the step, when k=1, let = , express right The state transition matrix affected, express for The influence matrix of is obtained by formula (9): express for The influence matrix of is obtained by formula (10): express for The impact matrix, represents a hyperparameter;
[0055] (9)
[0056] (10)
[0057] In formula (9)-formula (10), represents the fully connected layer, represents a hyperparameter;
[0058] Step 4.1.5: and Enter another cross Mamba block in the first CEM unit for Perform feature enhancement and use formula (11)-formula (12) to obtain the first The three-stage query state quantity of the step and The output of the three-stage query , thus obtaining The output of the three-stage query As a three-stage enhanced query vector :
[0059] (11)
[0060] (12)
[0061] In formula (11)-formula (12), Indicates The three-stage query state quantity of the step, when k=1, let = , express right The state transition matrix affected, express for The influence matrix of is obtained by formula (13): express for The influence matrix of is obtained by formula (14): express for The impact matrix;
[0062] (13)
[0063] (14)
[0064] Step 4.1.6: The second feature integration block uses equations (15) and (16) to obtain Four-stage enhancement of global visual features and four-stage enhanced query vector :
[0065] (15)
[0066] (16)
[0067] Step 4.1.7: Use formula (17) to get Shallow shared semantics :
[0068] (17)
[0069] In formula (17), Indicates channel fusion.
[0070] Further, each detection context propagation module includes: two self-Mamba blocks and two cross-Mamba blocks. The detection context propagation module in the human body branch in step 4.10 is obtained by the following process: Comprehensive detection characteristics of the individual :
[0071] Step 4.10.1: A self-Mamba block in the detection context propagation module of the human body branch uses formula (1)-formula (2) to Process it and get One-stage enhanced human detection features ;
[0072] Step 4.10.2: Another self-Mamba block in the detection context propagation module of the human body branch is based on formula (3)-formula (4) Process it and get One-stage reinforcement interaction feature ;
[0073] Step 4.10.3: A cross Mamba block in the detection context propagation module of the human body branch performs and to process Perform adaptive feature enhancement to obtain the Two-stage enhanced human detection features ;
[0074] Step 4.10.4: Another cross-Mamba block in the detection context propagation module of the human body branch is based on equation (11)-(12). as well as and The multiplication of the Interactive Perception Human Detection Features to process Perform adaptive feature enhancement to obtain the Three-stage enhanced human detection feature ;
[0075] Step 4.10.5: and Add together to get the comprehensive detection features of the human body .
[0076] Furthermore, in step 5, the total loss function is constructed using formula (16): :
[0077] (16)
[0078] In formula (16), , , and There are 4 hyperparameters. represents the GIOU loss of the human body bounding box, Represents the L1 loss of the human body bounding box; represents the object bounding box GIOU loss, Represents the L1 loss of the object bounding box; represents the cross entropy loss of object category, Represents the interaction category Focal loss.
[0079] An electronic device of the present invention includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute any of the character interaction detection methods, and the processor is configured to execute the program stored in the memory.
[0080] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the character interaction detection method when the computer program is executed by a processor.
[0081] Compared with the prior art, the present invention has the following beneficial effects:
[0082] 1. The present invention uses the Mamba architecture to drive the character interaction detection task. With a small number of parameters, it can achieve a good end-to-end detection task and has better recognition performance on different data sets than existing methods. Experimental results show that the proposed method outperforms the most advanced methods on both HICO-DET and V-COCO datasets.
[0083] 2. The present invention designs a decoder based on the Mamba architecture, which effectively absorbs the advantages of existing single-branch, dual-branch and triple-branch methods, while alleviating the common problems of previous methods. In terms of parameter quantity and inference speed, each branch is designed through cascaded low-rank adaptation, maintaining the simplicity of the single-branch method. The human and object features provide a good prior for the interaction branch, inheriting the advantage of fast convergence of the dual-branch method. The human interaction detection task is divided into three subtasks: human body detection, object detection and interaction classification, to absorb the advantage of the thorough decoupling of the three-branch method. In addition, the method of the present invention sets an independent branch for each subtask to explicitly isolate different parameters, and trains a specific Router unit for each subtask to dynamically allocate weight combinations, thereby realizing implicit gradient separation, and effectively alleviating conflicts between subtasks from both explicit and implicit perspectives.
[0084] 3. This paper designs a novel comprehensive progressive learning strategy to promote the recognition of rare and complex interaction categories. By designing the CEM unit and the detection context propagation module to extract more expressive feature representations, the detection performance of the model is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 It is the overall flow chart of the present invention;
[0086] Figure 2 It is a structural diagram of the self-Mamba block, cross-Mamba block, CEM unit and detection context propagation module of the present invention. DETAILED DESCRIPTION
[0087] In this embodiment, the overall process of a method for detecting human interaction based on the Mamba architecture is as follows: Figure 1 Specifically, the following steps are followed:
[0088] Step 1: Obtain a dataset of human interaction detection images and preprocess them to obtain a preprocessed image dataset ,in, Indicates preprocessed human interaction detection images; N represents the total number of human interaction detection images; Indicates the length of the human interaction detection image, Indicates the width of the human interaction detection image. Indicates the number of channels of the human interaction detection image;
[0089] Step 2: Figure 1 As shown in the figure, a human interaction detection network based on the Mamba architecture is constructed, including: a global feature encoding module, a multi-view feature aggregation module, a human branch, an object branch, and an interaction branch; wherein the global feature encoding module includes: a ResNet-50 model, a feature mapping unit, and a Transformer encoder; the multi-view feature aggregation module includes: two CEM units; the human branch and the object branch both include: two cascaded LoRA units, a Router unit for dynamic weighting, a detection context propagation module, and a fully connected layer; the interaction branch includes: two cascaded LoRA units, two Router units for dynamic weighting, and a fully connected layer.
[0090] Step 3: Global feature encoding module Process it and get A global visual representation ;
[0091] Step 3.1: Input into the ResNet-50 model for processing and obtain the Initial visual features , represents the length of the visual feature, represents the width of the visual feature, The number of channels representing visual features;
[0092] Step 3.2: The feature mapping unit uses convolution operation to The channel dimension is reduced to , and then use the flattening operator to expand the visual features after dimensionality reduction, and get the first Intermediate visual features ; Indicates the number of channels after dimension. In this embodiment, the global feature encoding module uses the DETR model (Nicolas Carion, et.al. End-to-end object detection with transformers. InEuropean Confer- ence on Computer Vision, pages 213–229. Springer, 2020) pre-trained on the MS-COCO dataset to initialize the model parameters. Take 256.
[0093] Step 3.3: Encode the position and After superposition, it is input into the Transformer encoder for processing to obtain the first A global visual representation .
[0094] Step 4: Multi-view feature aggregation module Process it and get Shallow shared semantics and Deep shared semantics , and input them into the human body branch, object branch, and interaction branch for processing, and the human body branch outputs The human bounding box prediction results and its confidence , object branch output The object category prediction results , object bounding box prediction results and its confidence And the interactive branch output The interactive category prediction results and its confidence .
[0095] Step 4.1: Random Generation query vectors to be learned ; and and Input into the first CEM unit for processing, and obtain the Shallow shared semantics ;in, Indicates the number of query vectors; represents the number of channels of the query vector. In this embodiment, and Take 64 and 256 respectively. The first CEM unit includes: two self-Mamba blocks, two feature integration blocks and two cross-Mamba blocks. The first CEM unit is obtained by the following process. Shallow shared semantics ,like Figure 2 As shown:
[0096] Step 4.1.1. The input is fed into a self-Mamba block in the first CEM unit for feature enhancement, and the second The global visual state of the first stage and The global visual output of the first stage , thus obtaining The global visual output of the first stage As the One-stage global enhanced visual features :
[0097] (1)
[0098] (2)
[0099] In formula (1)-formula (2), Indicates The global visual state of the first stage is k=1. = , Indicates the total number of steps; express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix.
[0100] Step 4.1.2: Another self-Mamba block input to the first CEM unit is used for feature enhancement, and the second One-stage query state quantity and The output of the first-stage query , thus obtaining The output of the first-stage query As a first-stage enhancement query vector :
[0101] (3)
[0102] (4)
[0103] In formula (3)-formula (4), Indicates The first-stage query state quantity of the step, when k=1, let = , express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix.
[0104] Step 4.1.3: The first feature integration block of the first CEM unit uses equations (5)-(6) to obtain the first Two-stage enhancement of global visual features and two-stage enhanced query vector :
[0105] (5)
[0106] (6)
[0107] In formula (5)-formula (6), Norm represents a LayNorm layer.
[0108] Step 4.1.4: and Input to a cross Mamba block in the first CEM unit for Perform feature enhancement and use formula (7)-formula (8) to obtain the first The three-stage global visual state of the step and The three-stage global visual output of the step , thus obtaining The three-stage global visual output of the step As the Three-stage enhanced visual features :
[0109] (7)
[0110] (8)
[0111] In formula (7)-formula (8), Indicates The three-stage global visual state of the step, when k=1, let = , express right The state transition matrix affected, express for The influence matrix of is obtained by formula (9): express for The influence matrix of is obtained by formula (10): express for The impact matrix, represents a hyperparameter. In this embodiment, Take 0.75.
[0112] (9)
[0113] (10)
[0114] In formula (9)-formula (10), represents the fully connected layer, represents a hyperparameter. In this embodiment, Take 0.50.
[0115] Step 4.1.5: and Enter another cross Mamba block in the first CEM unit for Perform feature enhancement and use formula (11)-formula (12) to obtain the first The three-stage query state quantity of the step and The output of the three-stage query , thus obtaining The output of the three-stage query As a three-stage enhanced query vector :
[0116] (11)
[0117] (12)
[0118] In formula (11)-formula (12), Indicates The three-stage query state quantity of the step, when k=1, let = , express right The state transition matrix affected, express for The influence matrix of is obtained by formula (13): express for The influence matrix of is obtained by formula (14): express for The impact matrix.
[0119] (13)
[0120] (14)
[0121] Step 4.1.6: The second feature integration block of the first CEM unit uses equations (15)-(16) to obtain Four-stage enhancement of global visual features and four-stage enhanced query vector :
[0122] (15)
[0123] (16)
[0124] Step 4.1.7: Use formula (17) to get Shallow shared semantics :
[0125] (17)
[0126] In formula (17), Indicates channel fusion.
[0127] Step 4.2: Input to the first LoRA unit of the human body branch for processing, and get the Individual low-level detection features ;
[0128] Will Input to the first LoRA unit of the object branch for processing, and get the Low-level detection features of objects ;
[0129] Step 4.3: and After superposition, The two are input into the second CEM unit for processing, realizing the context information exchange between the human and object branches, as well as the shallow shared semantics. The enhancement of Deep shared semantics The second CEM unit includes: two self-Mamba blocks, two feature integration blocks and two cross-Mamba blocks. The second CEM unit is obtained by the following process. Deep shared semantics .
[0130] Step 4.3.1. The input is sent to a self-Mamba block in the second CEM unit for feature enhancement, and the first The shared semantic state of the first step and The shared semantic output of the first stage , thus obtaining The shared semantic output of the first stage As the One-phase shared semantics :
[0131] (18)
[0132] (19)
[0133] In formula (18)-formula (19), Indicates The shared semantic state of the first stage of the step, when k=1, let = express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix.
[0134] Step 4.3.2: and Superposition to obtain Interactive person pair detection features , which is input into another self-Mamba block in the second CEM unit for feature enhancement, and the first The state quantity of the interactive character pair in one stage and The output of the interactive character pair in the first stage , thus obtaining The output of the interactive character pair in the first stage As the One-stage interactive person pair detection features :
[0135] (20)
[0136] (twenty one)
[0137] In formula (20)-formula (21), Indicates The state quantity of the interactive character pair in one stage of the step, when k=1, let = , express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix.
[0138] Step 4.3.3: The first feature integration block of the second CEM unit uses equations (22) and (23) to obtain Two-stage interactive person pair detection features and Two-phase shared semantics :
[0139] (twenty two)
[0140] (twenty three)
[0141] Step 4.3.4: and Input to a cross Mamba block in the second CEM unit for Perform feature enhancement and use equations (24)-(25) to obtain The three-stage shared semantic state and The three-stage shared semantic output , thus obtaining The three-stage shared semantic output As the Three-phase shared semantics :
[0142] (twenty four)
[0143] (25)
[0144] In formula (24)-formula (25), Indicates The three-stage shared semantic state quantity of the step, when k=1, let = , express right The state transition matrix affected, express for The influence matrix of is obtained by formula (26): express for The influence matrix of is obtained by formula (27): express for The impact matrix, Represents a hyperparameter.
[0145] (26)
[0146] (27)
[0147] Step 4.3.5: and Enter another cross Mamba block in the second CEM unit for Perform feature enhancement and use equations (28) and (29) to obtain The three-stage interactive character pair and The output of the three-stage interactive character pair , thus obtaining The output of the three-stage interactive character pair As a three-stage interactive person pair detection feature :
[0148] (28)
[0149] (29)
[0150] In formula (28)-formula (29), Indicates The state quantity of the three-stage interactive character pair of steps, when k=1, let = , express right The state transition matrix affected, express for The influence matrix of is obtained by formula (30): express for The influence matrix of is obtained by formula (31): express for The impact matrix.
[0151] (30)
[0152] (31)
[0153] Step 4.3.6: The second feature integration block of the second CEM unit uses equations (32) and (33) to obtain Four-phase shared semantics and Four-stage interactive person pair detection features :
[0154] (32)
[0155] (33)
[0156] Step 4.3.7: Use formula (34) to get Deep shared semantics :
[0157] (34)
[0158] Step 4.4: Input the second LoRA unit of the human body branch for feature enhancement processing, and obtain the first Intermediate detection features of the individual body ;
[0159] Will The second LoRA unit of the input object branch performs feature enhancement processing to obtain the first Intermediate detection features of objects .
[0160] Step 4.5: and Input the Router unit of the human body branch and perform adaptive dynamic weighting to obtain the first Advanced human detection features ;
[0161] Will and The Router unit in the input object branch performs adaptive dynamic weighting to obtain the first High-level object detection features .
[0162] Step 4.6: Improving the performance of the human interaction detection algorithm on rare and complex interaction categories is crucial for practical applications. Inspired by the fact that the human brain needs a gradual process to learn complex things, a comprehensive gradual learning strategy for interaction classification is constructed. It gradually obtains low-level, medium-level, high-level and comprehensive interaction features in four stages to obtain a more advanced interaction semantic understanding. , , Perform adaptive dynamic weighting to generate the Low-level interaction features .
[0163] Step 4.7: The first LoRA unit pair in the interactive branch Perform feature enhancement to obtain Intermediate Interaction Features ;
[0164] Step 4.8: The second Router unit pair in the interactive branch , , , Perform adaptive dynamic weighting to obtain Advanced Interaction Features ;
[0165] Step 4.9: The second LoRA unit pair in the interactive branch Perform feature enhancement to generate Comprehensive interaction features .
[0166] Step 4.10: In order to pass the information in the interaction branch that is beneficial to human detection to the human branch, the detection context propagation module in the human branch and Process it and get Comprehensive detection characteristics of the individual ;
[0167] In order to transfer the information in the interaction branch that is beneficial to object detection to the object branch, the detection context propagation module of the object branch and Process it and get Comprehensive detection features of objects ;
[0168] Each detection context propagation module includes two self-Mamba blocks and two cross-Mamba blocks. The detection context propagation module in the human body branch is obtained as follows: Comprehensive detection characteristics of the individual :
[0169] Step 4.10.1. The feature enhancement is performed in a self-Mamba block in the detection context propagation module of the human body branch, and the first The state quantity of human body detection in one stage and The output of human body detection in one stage , thus obtaining The output of human body detection in one stage As the One-stage enhanced human detection features :
[0170] (35)
[0171] (36)
[0172] In formula (35)-formula (36), Indicates The state quantity of human body detection in one stage of step, when k=1, let = . express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix.
[0173] Step 4.10.2: Another self-Mamba block in the detection context propagation module of the human body branch is input for feature enhancement, and the first The human body interaction state quantity of one stage and The output of human interaction in one stage , thus obtaining The output of human interaction in one stage As the One-stage enhancement of human interaction features :
[0174] (37)
[0175] (38)
[0176] In formula (37)-formula (38), Indicates The human body interaction state quantity of one stage of step, when k=1, let = , express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix.
[0177] Step 4.10.3: and A cross-Mamba block in the detection context propagation module of the input human branch is Perform adaptive feature enhancement and use equations (39)-(40) to obtain Two-stage human body detection state quantity and Output of the two-stage human detection , thus obtaining Output of the two-stage human detection As the Two-stage enhanced human detection features :
[0178] (39)
[0179] (40)
[0180] In formula (39)-formula (40), Indicates The state quantity of the two-stage human body detection is: = express right The state transition matrix affected, express for The influence matrix of is obtained by formula (41): express for The influence matrix of is obtained by formula (42): express for The impact matrix.
[0181] (41)
[0182] (42)
[0183] Step 4.10.4: and Multiply them together to get the interactive perception human detection feature ,Will and Another cross-Mamba block in the detection context propagation module of the input human branch is Perform guided feature enhancement and use equations (43) and (44) to obtain The three-stage human body detection state quantity and Output of the three-stage human body detection , thus obtaining Output of the three-stage human body detection As the Three-stage enhanced human detection feature :
[0184] (43)
[0185] (44)
[0186] In formula (43)-(44), Indicates The three-stage human body detection state quantity of the step, when k=1, let = express right The state transition matrix affected, express for The influence matrix of is obtained by formula (45): express for The influence matrix of is obtained by formula (46): express for The impact matrix.
[0187] (45)
[0188] (46)
[0189] Step 4.10.5: and Add together and get Comprehensive detection characteristics of the individual ;
[0190] The detection context propagation module in the object branch is obtained as follows: Comprehensive detection features of objects :
[0191] Step 4.10.6: A self-Mamba block in the detection context propagation module input to the object branch performs feature enhancement, and the first The state quantity of object detection in one stage and The output of one-stage object detection , thus obtaining The output of one-stage object detection As the One-stage enhanced object detection features :
[0192] (47)
[0193] (48)
[0194] In formula (47)-formula (48), Indicates The state quantity of the object detection in the first stage of the step, when k=1, let = . express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix.
[0195] Step 4.10.7. Another self-Mamba block in the detection context propagation module of the object branch is input for feature enhancement, and the first The state quantity of the object interaction in one stage of the step and The output of the object interaction in one stage , thus obtaining The output of the object interaction in one stage As the One-stage enhanced object interaction features :
[0196] (49)
[0197] (50)
[0198] In formula (49)-formula (50), Indicates The state quantity of the object interaction in one stage of the step, when k=1, let = , express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix.
[0199] Step 4.10.8: and A cross-Mamba block in the detection context propagation module of the input object branch is Perform adaptive feature enhancement and use equations (51) and (52) to get The state of the two-stage object detection and The output of the second stage object detection , thus obtaining The output of the second stage object detection As the Two-stage enhanced object detection features :
[0200] (51)
[0201] (52)
[0202] In formula (51)-formula (52), Indicates The state quantity of the second-stage object detection of the step, when k=1, let = express right The state transition matrix affected, express for The influence matrix of is obtained by formula (19): express for The influence matrix of is obtained by formula (20): express for The impact matrix;
[0203] (53)
[0204] (54)
[0205] Step 4.10.9: and Multiply them together to get the interactive perception object detection feature ,Will and Another cross-Mamba block in the detection context propagation module of the input object branch is Perform guided feature enhancement and use equations (55)-(56) to obtain The three-stage object detection state quantity and The output of the three-stage object detection , thus obtaining The output of the three-stage object detection As the Three-stage enhanced object detection features :
[0206] (55)
[0207] (56)
[0208] In formula (55)-(56), Indicates The three-stage object detection state quantity of the step, when k=1, let = express right The state transition matrix affected, express for The influence matrix of is obtained by formula (57): express for The influence matrix of is obtained by formula (58): express for The impact matrix;
[0209] (57)
[0210] (58)
[0211] Step 4.10.10: and Add together and get Comprehensive detection features of objects .
[0212] Step 4.11: Fully connected layer pairs in the human body branch Process and obtain The human bounding box prediction results And the detection confidence of the human bounding box , express Number of people;
[0213] The fully connected layer pair of the object branch Process and obtain Object bounding box prediction results , detection confidence of the object bounding box And the object category prediction results , Indicates the number of detected objects. Indicates the number of categories of detected objects;
[0214] Fully connected layer pairs in the interaction branch Process and obtain The interactive category prediction results and interaction category confidence ; Indicates the number of detected interactions; Represents the number of categories of detected interactions.
[0215] Step 5: Based on , , And the real result of human bounding box , object bounding box real results , Object category real results , Interaction category real results , construct the total loss function of the human interaction detection network , and used to train the human interaction detection network until the total loss function Until convergence, the optimal person interaction detection model is obtained, and the total loss function It is built according to the following steps:
[0216] Step 5.1: Based on and , using formula (59) to construct the GIOU loss of the human detection frame :
[0217] (59)
[0218] In formula (59), Represents the intersection area of the predicted result of the human bounding box and the true result of the human bounding box, Represents the union area of the predicted result of the human bounding box and the true result of the human bounding box, Represents the area of the smallest rectangle that contains the predicted result of the human bounding box and the true result of the human bounding box;
[0219] Step 5.2: Based on and , using formula (60) to construct the L1 loss of the human detection frame :
[0220] (60)
[0221] In formula (60), , , , Respectively The horizontal coordinate of the center point, the vertical coordinate of the center point, the width of the human body bounding box prediction result, and the height of the human body bounding box prediction result; , , , Respectively The horizontal coordinate of the center point, the vertical coordinate of the center point, the width of the true result of the human body bounding box, and the height of the true result of the human body bounding box;
[0222] Step 5.3: Based on and , using formula (61) to construct the GIOU loss of the object detection box :
[0223] (61)
[0224] In formula (61), Represents the intersection area of the object bounding box prediction result and the object bounding box true result, Represents the union area of the object bounding box prediction result and the object bounding box true result, Represents the area of the smallest rectangle that contains the predicted result of the object bounding box and the actual result of the object bounding box;
[0225] Step 5.4: Based on and , use formula (62) to construct the L1 loss of the object detection box :
[0226] (62)
[0227] In formula (62), , , , Respectively The horizontal coordinate of the center point, the vertical coordinate of the center point, the width of the object bounding box prediction result, and the height of the object bounding box prediction result; , , , Respectively The horizontal coordinate of the center point, the vertical coordinate of the center point, the width of the object bounding box true result, and the height of the object bounding box true result.
[0228] Step 5.5: Based on and , using formula (63) to construct the object category cross entropy loss :
[0229] (1- (63)
[0230] In formula (63), Indicates The logarithmic function with base .
[0231] Step 5.6: Based on and , using formula (64) to construct the interactive category Focal loss :
[0232] (64)
[0233] Step 5.7: Use formula (65) to construct the total loss function :
[0234] (65)
[0235] In formula (65), , , and There are 4 hyperparameters. In this embodiment, , , and Take 3, 1, 1.25 and 1 respectively.
[0236] Step 6: In the reasoning process, The image to be tested Input into the optimal human interaction detection model for processing, and obtain The Individual human bounding box prediction results and its detection confidence , object bounding box prediction results and its detection confidence , Object category prediction results , Interaction category prediction results and its confidence ; , express The total number of prediction results in ;
[0237] Will , , After adding, as the The confidence score of the prediction result of the interaction between the characters is obtained. The first few prediction results with high confidence scores are taken as The final prediction result.
[0238] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0239] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium, and the computer program executes the steps of the above method when executed by a processor.
[0240] Example
[0241] In order to verify the effectiveness of the method HOIMamba of the present invention, the common HICO-DET dataset (Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng. Learningto detect human-object interactions. workshop on applications of computervision, 2017.) and V-COCO dataset (Saurabh Gupta and Jitendra Malik. Visualsemantic role labeling. arXiv: Computer Vision and Pattern Recognition, 2015.) were selected for training and testing in this embodiment, and compared with advanced methods based on Transformer architecture, namely QPIC (Tamura, M.; Ohashi, H.; and Yoshinaga, T. 2021. QPIC: Query-basedpairwise human-object interaction detection with image-wide contextualinformation. In Proceedings of the IEEE / CVF Conference on Computer Vision andPattern Recognition, 10410–10419.), UPT (Zhang, FZ; Campbell, D.; andGould, S. 2022. Efficient two-stage detection of human-object interactions with a novel unary-pairwise transformer. In Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition, 20104–20112.), MUREN(Kim, S.; Jung, D.; and Cho, M. 2023. Relational context learning for human-object interaction detection.In Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition, 2925–2934.), DPAD (Gao, J.; Liang,K.; Wei, T.; Chen, W.; Ma, Z.; and Guo, J. 2024. Dual-Prior AugmentedDecoding Network for Long Tail Distribution in HOI Detection. In Proceedingsof the AAAI Conference on Artificial Intelligence, volume 38, 1806–1814.),CDN (Zhang, A.; Liao, Y.; Liu, S.; Lu, M.; Wang, Y.; Gao, C.; and Li, X.2021. Mining the benefits of two-stage and one-stage hoi detection. Advancesin Neural Information Processing Systems, 34: 17209–17220.), GEN-VLKT (Liao,Y.; Zhang, A.; Lu, M.; Wang, Y.; Li, X.; and Liu, S. 2022. Gen-vlkt: Simplifyassociation and enhance interaction understanding for hoi detection. InProceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition, 20123–20132.), RmLR (Cao, Y.; Tang, Q.; Yang, F.; Su, X.; You,S.; Lu, X.; and Xu, C. 2023. Re-mine, learn and reason: Exploring thecrossmodal semantic correlations for language-guided hoi detection.InProceedings of the IEEE / CVF International Conference on Computer Vision,23492–23503.), SCTC (Jiang, W.; Ren, W.; Tian, J.; Qu, L.; Wang, Z.; and Liu,H. 2024. Exploring Self-and Cross-Triplet Correlations for Human-ObjectInteraction Detection. In Proceedings of the AAAI Conference on ArtificialIntelligence, volume 38, 2543–2551.), FGAHOI (Ma, S.; Wang, Y.; Wang, S.; andWei, Y. 2023. Fgahoi: Fine-grained anchors for human-object interactiondetection. IEEE Transactions on Pattern Analysis and Machine Intelligence.),PViC (Zhang, F. Z.; Yuan, Y.; Campbell, D.; Zhong, Z.; and Gould, S. 2023.Exploring predicate visual context in detecting of human-object interactions.In Proceedings of the IEEE / CVF International Conference on Computer Vision,10411–10421.), MP-HOI (Yang, J.; Li, B.; Zeng, A.; Zhang, L.; and Zhang, R.2024a. Open-World Human-Object Interaction Detection via Multimodal Prompts.In Proceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition, 1695416964.), HybHOI (Wu, EZ; Li, Y.; Wang, Y.; and Wang, S. 2024. Exploring Pose-Aware Human-Object Interaction via Hybrid Learning. InProceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition, 17815–17825.). The present invention uses the mean Average Precision (mAP) as the evaluation index, and replaces the ResNet-50 model in the global feature encoding module with the ResNet-101 model and the Swin-Large (Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin,S.; and Guo, B. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE / CVF International Conference on Computer Vision, 10012–10022.) model to verify the scalability of the present invention's method HOIMamba. .
[0242] Table 1
[0243]
[0244] The experimental results show that the proposed method performs better than other methods in various settings of the two datasets, which proves the feasibility of the proposed method. In particular, it significantly surpasses the previous method in the rare category setting of the HICO-DET dataset, proving the rationality of the designed decoupled progressive learning. The experiment shows that the proposed method can effectively cope with the recognition of rare and complex interaction categories, thereby extracting advanced interaction feature information and completing the task of human interaction detection.
[0245] In order to verify the efficiency of the method of the present invention, the proposed method HOIMamba is compared with QPIC (Tamura, M.; Ohashi, H.; and Yoshinaga, T. 2021. QPIC: Query-based pairwise human-object interaction detection with image-wide contextual information. InProceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition, 10410–10419.), AS-Net (Chen, M.; Liao, Y.; Liu, S.; Chen, Z.;Wang, F.; and Qian, C. 2021. Reformulating hoi detection as adaptive setprediction. In Proceedings of the IEEE / CVF Conference on Computer Vision andPattern Recognition, 9004–9013.), MUREN (Kim, S.; Jung, D.; and Cho, M. 2023.Relational context learning for human-object interaction detection. InProceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2925–2934.) to compare the number of parameters and computational complexity. The experimental results are shown in Table 2:
[0246] Table 2
[0247]
[0248] The experimental results show that the proposed method is superior to the previous methods in terms of parameter quantity and computational complexity, thus proving the feasibility of the proposed method. The experimental results show that the proposed method can capture high-level and comprehensive visual semantics and complete the task of human interaction detection efficiently and accurately.
Claims
1. A method for detecting human interaction based on Mamba architecture, characterized in that: The steps are as follows: Step 1: Obtain the human interaction detection image dataset and preprocess it to obtain the preprocessed image dataset ,in, Indicates preprocessed human interaction detection images; N represents the total number of human interaction detection images; Indicates the length of the human interaction detection image, Indicates the width of the human interaction detection image. Indicates the number of channels of the human interaction detection image; Step 2: Construct a human interaction detection network based on the Mamba architecture, including: a global feature encoding module, a multi-view feature aggregation module, a human branch, an object branch, and an interaction branch; wherein the global feature encoding module includes: a ResNet-50 model, a feature mapping unit, and a Transformer encoder; the multi-view feature aggregation module includes: two CEM units; the human branch and the object branch both include: two cascaded LoRA units, a Router unit for dynamic weighting, a detection context propagation module, and a fully connected layer; the interaction branch includes: two cascaded LoRA units, two Router units for dynamic weighting, and a fully connected layer; Step 3: The global feature encoding module Process it and get A global visual representation ; Step 4: The multi-view feature aggregation module Process it and get Shallow shared semantics and Deep shared semantics , and input them into the human body branch, object branch, and interaction branch for processing, and the human body branch outputs The human bounding box prediction results and its confidence , object branch output The object category prediction results , object bounding box prediction results and its confidence And the interactive branch output Interaction category prediction results and its confidence ; Step 5: Based on , , And the real result of human bounding box , object bounding box real results , Object category real results , Interaction category real results , construct the total loss function of the human interaction detection network , and used to train the human interaction detection network until the total loss function Until convergence, the optimal human interaction detection model is obtained; Step 6: In the reasoning process, The image to be tested Input into the optimal human interaction detection model for processing, and obtain The Individual human bounding box prediction results and its detection confidence , object bounding box prediction results and its detection confidence , Object category prediction results , Interaction category prediction results and its confidence ; , express The total number of prediction results in ; Will , , After adding, as the The confidence score of the prediction result of the interaction between the characters is obtained. The first few prediction results with high confidence scores are taken as The final prediction result.
2. A method for detecting human interaction based on Mamba architecture according to claim 1, characterized in that: The step 3 comprises: Step 3.1: Input into the ResNet-50 model for processing and obtain the Initial visual features , represents the length of the visual feature, represents the width of the visual feature, The number of channels representing visual features; Step 3.2: The feature mapping unit uses a convolution operation to The channel dimension is reduced to , and then use the flattening operator to expand the visual features after dimensionality reduction, and get the first Intermediate visual features ; The number of channels after the dimension is expressed; Step 3.3: Encode the position and After superposition, it is input into the Transformer encoder for processing to obtain the first A global visual representation .
3. A method for detecting human interaction based on Mamba architecture according to claim 2, characterized in that: The step 4 comprises: Step 4.1: Random Generation query vectors to be learned ; and and Input into the first CEM unit for processing, and obtain the Shallow shared semantics ;in, Indicates the number of query vectors; Indicates the number of channels of the query vector; Step 4.2: Input to the first LoRA unit of the human body branch for processing, and get the Individual low-level detection features ; Will Input to the first LoRA unit of the object branch for processing, and get the Low-level detection features of objects ; Step 4.3: and After superposition, and input them into the second CEM unit for processing to obtain Deep shared semantics ; Step 4.4: Input the second LoRA unit of the human body branch for feature enhancement processing, and obtain the first Intermediate detection features of the individual body ; Will The second LoRA unit of the input object branch performs feature enhancement processing to obtain the first Intermediate detection features of objects ; Step 4.5: and Input the Router unit of the human body branch and perform adaptive dynamic weighting to obtain the first Advanced human detection features ; Will and The Router unit in the input object branch performs adaptive dynamic weighting to obtain the first Advanced object detection features ; Step 4.6: The first Router unit in the interactive branch , , Perform adaptive dynamic weighting to generate the Low-level interaction features ; Step 4.7: The first LoRA unit pair in the interactive branch Perform feature enhancement to obtain Intermediate Interaction Features ; Step 4.8: The second Router unit in the interactive branch is , , , Perform adaptive dynamic weighting to obtain Advanced Interaction Features ; Step 4.9: The second LoRA unit pair in the interactive branch Perform feature enhancement to generate Comprehensive interaction features ; Step 4.10: The detection context propagation module in the human body branch and Process it and get Comprehensive detection characteristics of the individual ; The detection context propagation module of the object branch is and Process it and get Comprehensive detection features of objects ; Step 4.11: The fully connected layer in the human body branch Process and obtain The human bounding box prediction results And the detection confidence of the human bounding box , express Number of people; The fully connected layer of the object branch is Process and obtain Object bounding box prediction results , detection confidence of the object bounding box And the object category prediction results , Indicates the number of detected objects. Indicates the number of categories of detected objects; The fully connected layers in the interaction branch are Process and obtain Interaction category prediction results and interaction category confidence ; Indicates the number of detected interactions; Represents the number of categories of detected interactions.
4. The method for detecting human interaction based on Mamba architecture according to claim 3, characterized in that: Each CEM unit includes: two self-Mamba blocks, two feature integration blocks and two cross-Mamba blocks; the first CEM unit in step 4.1 is obtained by the following process Shallow shared semantics : Step 4.1.
1. The input is fed into a self-Mamba block in the first CEM unit for feature enhancement, and the second The global visual state of the first stage and The global visual output of the first stage , thus obtaining The global visual output of the first stage As the One-stage global enhanced visual features : (1) (2) In formula (1)-formula (2), Indicates The global visual state of the first stage is k=1. = , Indicates the total number of steps; express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix; Step 4.1.2: Another self-Mamba block input to the first CEM unit is used for feature enhancement, and the second One-stage query state quantity and The output of the first-stage query , thus obtaining The output of the first-stage query As a first-stage enhancement query vector : (3) (4) In formula (3)-formula (4), Indicates The first-stage query state quantity of the step, when k=1, let = , express right The state transition matrix affected, express for The impact matrix, express for The impact matrix, express for The impact matrix; Step 4.1.3: The first feature integration block uses equations (5) and (6) to obtain Two-stage enhancement of global visual features and two-stage enhanced query vector : (5) (6) In formula (5)-(6), Norm represents a LayNorm layer; Step 4.1.4: and Input to a cross Mamba block in the first CEM unit for Perform feature enhancement and use formula (7)-formula (8) to obtain the first The three-stage global visual state of the step and The three-stage global visual output of the step , thus obtaining The three-stage global visual output of the step As the Three-stage enhanced visual features : (7) (8) In formula (7)-formula (8), Indicates The three-stage global visual state of the step, when k=1, let = , express right The state transition matrix affected, express for The influence matrix of is obtained by formula (9): express for The influence matrix of is obtained by formula (10): express for The impact matrix, represents a hyperparameter; (9) (10) In formula (9)-formula (10), represents the fully connected layer, represents a hyperparameter; Step 4.1.5: and Enter another cross Mamba block in the first CEM unit for Perform feature enhancement and use formula (11)-formula (12) to obtain the first The three-stage query state quantity of the step and The output of the three-stage query , thus obtaining The output of the three-stage query As a three-stage enhanced query vector : (11) (12) In formula (11)-formula (12), Indicates The three-stage query state quantity of the step, when k=1, let = , express right The state transition matrix affected, express for The influence matrix of is obtained by formula (13): express for The influence matrix of is obtained by formula (14): express for The impact matrix; (13) (14) Step 4.1.6: The second feature integration block uses equations (15) and (16) to obtain Four-stage enhancement of global visual features and four-stage enhanced query vector : (15) (16) Step 4.1.7: Use formula (17) to get Shallow shared semantics : (17) In formula (17), Indicates channel fusion.
5. The method for detecting human interaction based on Mamba architecture according to claim 4, characterized in that: Each detection context propagation module includes: two self-Mamba blocks and two cross-Mamba blocks. The detection context propagation module in the human body branch in step 4.10 is obtained by the following process. Comprehensive detection characteristics of the individual : Step 4.10.1: A self-Mamba block in the detection context propagation module of the human body branch uses formula (1)-formula (2) to Process it and get One-stage enhanced human detection features ; Step 4.10.2: Another self-Mamba block in the detection context propagation module of the human body branch is based on formula (3)-formula (4) Process it and get One-stage reinforcement interaction feature ; Step 4.10.3: A cross Mamba block in the detection context propagation module of the human body branch performs and to process Perform adaptive feature enhancement to obtain the Two-stage enhanced human detection features ; Step 4.10.4: Another cross-Mamba block in the detection context propagation module of the human body branch is based on equation (11)-(12). as well as and The multiplication of the Interactive Perception Human Detection Features to process Perform adaptive feature enhancement to obtain the Three-stage enhanced human detection feature ; Step 4.10.5: and Add together to get the comprehensive detection features of the human body .
6. The method for detecting human interaction based on Mamba architecture according to claim 5, characterized in that: In step 5, the total loss function is constructed using formula (16): : (16) In formula (16), , , and There are 4 hyperparameters. represents the GIOU loss of the human body bounding box, Represents the L1 loss of the human body bounding box; represents the object bounding box GIOU loss, Represents the L1 loss of the object bounding box; represents the cross entropy loss of object category, Represents the interaction category Focal loss.
7. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute any one of the character interaction detection methods described in claims 1-6, and the processor is configured to execute the program stored in the memory.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the character interaction detection method described in any one of claims 1 to 6 are executed.
Citation Information
Patent Citations
Character interaction detection method and device, equipment and storage medium
CN115097941A
TED-Net-based non-contact human-object interaction detection method
CN116563605A
Target detection method and device based on human interactive perception and storage medium
CN116883804A
Human-object interaction detection method based on virtual enhancement
CN117115695A
A zero-sample robot control method, device, terminal and storage medium
CN119748461A