Method for Detecting Human-Object Interaction Based on DETR's Pairwise Decoding Interaction of Persons

Through the DETR-based character pair decoding interaction detection method, combining semantic modality and loss function, the problems of training difficulties and feature connection neglect are solved, and the accuracy and network performance of person-to-object interaction detection are improved.

CN115147931BActive Publication Date: 2025-07-18ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210864552.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-21
Publication Date
2025-07-18
Estimated Expiration
2042-07-21

AI Technical Summary

Technical Problem

The existing DETR-based character interaction detection algorithm has problems with training difficulties, neglect of multiple query feature connections and lack of reliability in triple prediction, resulting in insufficient detection accuracy.

Method used

A DETR-based character pairwise decoding interaction detection method is adopted, combined with semantic modality, and a variety of feature information is fused by querying vector classifiers, semantic networks and pairwise fusion detection networks, and supervised training is used using verb cross entropy and semantic relative entropy loss functions.

Benefits of technology

It improves the accuracy of human-object interaction detection, reduces training resource consumption, expands the network's receptive field, and enhances network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147931B_ABST
    Figure CN115147931B_ABST
Patent Text Reader

Abstract

The present invention discloses a human-object interaction detection method based on DETR for paired decoding and interaction of characters. By passing an image through a trained DETR model, object bounding boxes, object categories, and query vectors of characters are obtained, thereby reducing the model training time. Then, the query vectors and object categories are input into a query vector classifier to obtain query vectors of humans, query vectors of objects, and object categories. The object categories are input into a semantic network to obtain semantic query vectors of objects. The query vectors of objects and the semantic query vectors of objects are fused to obtain fused query vectors of objects. The fused query vectors of objects and the query vectors of humans are combined to obtain object query vectors. Finally, the object query vectors are input into a paired fusion detection network to achieve human-object interaction detection. The present invention improves the accuracy of human-object interaction detection, expands the receptive field of the network, and improves the performance of the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of human-object interaction detection, and particularly relates to a human-object interaction detection method based on DETR for pairwise decoding interaction of characters. Background Art

[0002] Human-object interaction detection is a downstream task of object detection and a currently popular computer vision task. Compared with object detection that detects object bounding boxes and object categories, human-object interaction detection locates the interacting human-object pairs in an image and classifies the actions.

[0003] Before the Transformer model was applied to vision tasks, human-object interaction detection algorithms often used convolutional networks to extract visual features, such as HO-RCNN, which is a typical human-object interaction detection algorithm based on convolutional neural networks. The algorithm uses R-CNN to locate relevant regions, and after the backbone network crops the features, the features are fused through multiple branches; STIGPN uses graph convolution to iterate feature information. However, these methods still have limitations, that is, traditional convolutional networks cannot introduce global information and will also cause feature pollution.

[0004] Recently, the end-to-end Detection Transformer network DETR (Detection Transformer) algorithm has become popular. It uses depth self-attention to replace convolution and can introduce global information to complete set prediction. The DETR algorithm is used to handle object detection problems. Therefore, it is a very natural idea to introduce DETR into the downstream task of object detection, human-object relationship detection, and even various fields of vision. The QPIC algorithm is to introduce DETR into the field of human-object relationship interaction detection, use it as a basic detector, and extract corresponding context information to predict the final triple set.

[0005] However, the related human-object interaction detection algorithms that directly complete triple set prediction based on DETR still have some problems. One is the difficulty in training, which is a disadvantage brought by the Transformer model; the other is that a single query is used as a whole to predict features, ignoring the more intuitive feature connections between multiple queries. Therefore, a special structure needs to be designed to fuse the corresponding feature connections. At the same time, the finally predicted <human, object, interaction> triple lacks corresponding reliability judgment and requires a semantic model to make constraints. Summary of the Invention

[0006] This application proposes a human-object interaction detection method based on DETR for pairwise decoding interaction of characters to reduce training resources and improve the accuracy of human-object interaction detection by combining semantic modalities.

[0007] To achieve the above object, the technical solution of the present application is as follows:

[0008] A method for human-object interaction detection based on DETR for pairwise decoding interaction of people, including:

[0009] Inject the feature map obtained from the original image through the backbone network into the trained DETR network, where the DETR network includes an encoder, a decoder, and an MLP layer, to obtain the query vector output by the decoder, as well as the target box and target category finally output by the DETR network;

[0010] Input the query vector and the target category into the query vector classifier to obtain the query vector of people, the query vector of objects, and the category of objects;

[0011] Input the category of objects into the semantic network to obtain the semantic query vector of objects;

[0012] Fuse the query vector of objects and the semantic query vector of objects to obtain the fused query vector of objects, and merge the fused query vector of objects and the query vector of people to obtain the object query vector;

[0013] Input the object query vector into the pairwise fusion detection network to achieve human interaction detection.

[0014] Further, the semantic network includes a spatial attention module and a semantic aggregation module. The input feature of the semantic spatial attention module is the verb embedding vector of the dataset, and the output is the semantic spatial attention feature;

[0015] The input feature of the semantic aggregation module is the semantic spatial attention feature output by the semantic spatial attention module and the category of objects output by the query vector classifier. The semantic spatial attention feature passes through a linear layer, a ReLU activation function, a linear layer, and a sigmoid activation function to obtain the attention feature, which is multiplied by the feature obtained by another linear layer of the category of objects. The result passes through a linear layer, a normalization layer, a ReLU activation function, and a linear layer in sequence and then adds the category of objects, and then is input into the Transformer layer to obtain the semantic query vector of objects.

[0016] Further, the fusion of the query vector of objects and the semantic query vector of objects to obtain the fused query vector of objects includes:

[0017] Add the query vector of objects and the semantic query vector of objects and then pass through the ReLU activation function, and subtract the square of the difference between the query vector of objects and the semantic query vector of objects.

[0018] Further, the pairwise fusion detection network sequentially includes an improved Transformer encoder, a pairwise fusion module, a Transformer decoder, and an MLP layer;

[0019] The improved Transformer encoder has input features that are an object query vector and a pairwise box position encoding respectively. In the improved Transformer encoder, the object query vectors are paired and combined with the pairwise box position encoding. Through a linear layer and a sigmoid activation function, the output of the first branch is obtained; the object query vectors are copied and multiplied by the elements of the pairwise box position encoding to obtain the output of the second branch; the output elements of the two branches are multiplied, passed through a linear layer, added to the input object query vector, and then passed through a normalization layer, a forward propagation layer, and a normalization layer to output the pairwise query vector;

[0020] In the pairwise fusion module, the pairwise query vectors are respectively combined with the pairwise box position encoding and the global visual features after adaptive average pooling, then multiplied after passing through a linear layer, and then successively passed through a ReLU activation function, a linear layer, and a ReLU activation function to obtain the final pairwise query vector that fuses multiple features;

[0021] The pairwise query vector that fuses multiple features is decoded by the Transformer decoder and output to the MLP to obtain the probability score of the human-object interaction action, thus completing the detection of the human-object interaction action.

[0022] Furthermore, the human-object interaction detection method based on DETR pairwise decoding interaction for human and object also includes:

[0023] Calculate the overall loss function of the network, perform backpropagation, and update the network parameters;

[0024] Among them, the overall loss function of the network is:

[0025] L total = L a + L SKL

[0026] Among them, L total represents the overall loss function, L a and L SKL respectively represent the verb cross-entropy loss function and the semantic relative entropy loss function;

[0027] The verb cross-entropy loss function L a is:

[0028]

[0029] Among them, N q represents the number of verb categories, represents the number of predicted verb categories corresponding to the object in statistics, and Φ represents the set of all true values, Indicates in the prediction set, l f is the focal loss, l f (p t ) = -α t (1 - p t ) γ log(p t ), α t is the parameter to suppress the imbalance between positive and negative samples, γ is the parameter to control the imbalance between the number of easy / hard samples, p t is the sample, where represents the true verb category;

[0030] The semantic relative entropy loss function L SKL is:

[0031]

[0032] where is the symmetric conditional distribution of verbs in the dataset, A is the adjacency matrix of verbs processed by the semantic space attention module, is the KL divergence loss function;

[0033] can be obtained through the following calculation:

[0034]

[0035] where N p is the number of verbs in the dataset, c ij is:

[0036]

[0037] A can be obtained through the following calculation:

[0038]

[0039] where τ is the temperature parameter for scaling the normalized semantic inner product softmax distribution, is the verb embedding vector processed by the semantic space attention module, and T is the transpose symbol.

[0040] A method for human-object interaction detection based on DETR's pairwise decoding interaction for characters proposed in this application uses a trained DETR model to alleviate the problem of long training time. To enhance the representation of features in the semantic modality, the semantic modality is added to improve the accuracy of human-object interaction detection. Adding the Transformer module improves the network's ability to extract global information, expands the network's receptive field, and improves the network's performance. Finally, the semantic relative entropy loss function is proposed to strengthen the network's supervision of semantics. Description of the Drawings

[0041] Figure 1 This is a flowchart of the method for detecting human-object interaction based on DETR's paired decoding interaction of humans;

[0042] Figure 2 This is a schematic diagram of the overall network structure of this application;

[0043] Figure 3 This is a schematic diagram of the DETR network structure of this application;

[0044] Figure 4 This is a schematic diagram of the multi-modal fusion network structure of an embodiment of this application;

[0045] Figure 5 This is a schematic diagram of the semantic network structure of an embodiment of this application;

[0046] Figure 6 This is a schematic diagram of the paired fusion detection network structure;

[0047] Figure 7 This is a schematic diagram of the improved Transformer encoder structure of an embodiment of this application;

[0048] Figure 8 This is a schematic diagram of the paired fusion module structure of an embodiment of this application. Detailed implementation manner

[0049] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0050] As Figure 1 shown, a method for detecting human-object interaction based on DETR's paired decoding interaction of humans includes:

[0051] S1. Inject the feature map obtained from the original image through the backbone network into the trained DETR network, where the DETR network includes an encoder, a decoder, and an MLP layer, to obtain the query vector output by the decoder, as well as the target box and target category finally output by the DETR network.

[0052] As Figure 2 shown, the entire network of this application includes a backbone network, a DETR network, a multi-modal fusion network, a semantic network, and a paired fusion detection network. The following elaborates in detail the process of detecting human interaction through the above networks for the original image.

[0053] First, extract the global visual features of the original image through the backbone network. In this embodiment, the backbone network can adopt ResNet50.

[0054] After extracting the global visual features of the original image, input them into the trained DETR network. As Figure 3 shown, the DETR network includes an encoder, a decoder, and an MLP layer. The input features pass through N Transformer encoding layers, N Transformer decoding layers, and an MLP layer in sequence from the input side to the output side. In this model, N is 6.

[0055] It should be noted that the DETR network of this application includes two parts. One is the output features of the decoder at the end, that is, the query vectors (Queries); the other is the output of the last MLP layer of the DETR network, that is, the target boxes and target classes.

[0056] S2. Input the query vectors and target classes into the query vector classifier to obtain the query vectors of people, the query vectors of objects, and the classes of objects.

[0057] As Figure 4 shown, the multimodal fusion network includes a query vector (Queries) classifier and a modality fusion module.

[0058] The query vectors output by the decoder of the DETR network and the target classes output by the MLP layer of the DETR network are input into the query vector classifier one by one. The query vectors of people (Human Queries), the query vectors of objects (Object Queries), and the classes of objects (Object Classes) are obtained. The object class information is obtained by removing the objects classified as people from the object classes obtained by the DETR network.

[0059] Specifically, the query vector classifier first divides the corresponding indices of people (the index of people is 1), the corresponding indices of objects, and the corresponding indices of the background in the target classes output by the MLP layer according to the corresponding indices in the dataset object labels (a total of 80 classes) into clusters of people, objects, and the background. The number of elements in the clusters of people and objects should be greater than or equal to K1 and less than or equal to K2 (K1, K2 are hyperparameters, set manually). If the number of elements in the cluster of people or objects is K and less than K1, then sort the confidence scores of the corresponding target classes in the background cluster from largest to smallest, and retain the top K1 - K target classes to make the number of cluster elements meet the conditions; if the number of elements in the cluster of people or objects is K and greater than K2, then sort the confidence scores of the corresponding target classes in the cluster from largest to smallest, and retain the top K2 target classes to make the number of cluster elements meet the conditions; after obtaining the clusters of people and objects, obtain the corresponding query vectors of people and objects according to the target classes in the clusters; finally, output the classes of objects in the cluster of objects, the corresponding query vectors of objects, and the query vectors of people corresponding to the classes of people in the cluster of people.

[0060] Step S3: Input the category of the object into the semantic network to obtain the semantic query vector of the object.

[0061] The semantic network in this embodiment is as Figure 5 shown, including a semantic space attention module and a semantic aggregation module.

[0062] Among them, the semantic space attention module is used to learn the verb embedding vectors (embeddings) of the dataset and learn the distribution of the relationship between objects and actions in the dataset.

[0063] The semantic aggregation module, according to the input category of the object, combines the distribution of the relationship between objects and actions in the dataset to obtain the semantic query vector of the object.

[0064] Specifically, the datasets used by the semantic space attention module are V-COCO and HICO-DET. These two datasets are used to detect human interaction actions. The datasets include images and corresponding labels. The labels include the target boxes of the people and objects of the interaction objects, the category labels of the objects (the objects include people and objects), and the interaction action category labels. The semantic space attention module counts the corresponding objects and action categories in the dataset to obtain the relationship distribution for subsequent processing.

[0065] The input feature of the semantic space attention module is the verb embedding vector of the dataset. The features are accumulated through a recurrent network including an attention layer and a ReLU activation function. The attention layer obtains query, key, and value features by passing the input feature through three linear layers respectively. Then, the query feature is multiplied by the transpose matrix of the key feature and divided by the square root of the hidden layer dimension to obtain the attention map feature. Then, the attention map feature is multiplied by the value feature after passing through the softmax. Finally, the feature generated by the module and the input feature are added as the output feature of the module, which is called the semantic space attention feature. The spatial attention mechanism is a relatively mature technology in this field and will not be elaborated here.

[0066] The input features of the semantic aggregation module are the semantic space attention features output by the above semantic space attention module and the category of the object output by the query vector classifier, and they pass through a cross-attention layer and a Transformer layer in sequence. In the cross-attention layer, the semantic space attention features pass through a linear layer, a ReLU activation function, a linear layer, and a sigmoid activation function to obtain attention features, which are multiplied by the features obtained by passing the category of the object through another linear layer. The result passes through a linear layer, a normalization layer (layerNorm), a ReLU activation function, and a linear layer in sequence and then adds the category of the object, and then is input into the Transformer layer to obtain the semantic query vector of the object.

[0067] The Transformer layer processes the input features through three linear layers respectively to obtain query, key, and value features. Then, the attention map features are obtained by multiplying the query features with the transposed matrix of the key features. After that, the attention map features are multiplied with the value features after passing through the softmax function. Finally, the output features of the semantic aggregation module are obtained by passing through a linear layer, a layer normalization layer, a ReLU activation function, and a linear layer and adding the input features. The Transformer is a relatively mature technology in this field and will not be elaborated here.

[0068] In this embodiment, the semantic network filters the category information of the object as the input according to the query vector classifier in the multi-modal fusion network to obtain the corresponding semantic query.

[0069] Step S4: Fuse the query vector of the object and the semantic query vector of the object to obtain the fused object query vector, and merge the fused object query vector and the query vector of the person to obtain the object query vector.

[0070] The multi-modal fusion network fuses the semantic query vector of the object returned by the semantic network with the query vector of the object output by the query vector classifier. The fusion is performed in the modal fusion module to obtain the fused object query vector. In the modal fusion module, the two input features are added and then passed through the ReLU activation function, and then the square of the subtraction of the two features is subtracted.

[0071] Then, the fused object query vector and the query vector of the person are merged to obtain the object query vector, that is, the query vectors of M persons and the query vectors of N modally fused objects are concatenated to obtain M + N query vectors.

[0072] Step S5: Input the object query vector into the pairwise fusion detection network to implement person-object interaction detection.

[0073] As Figure 6 shown, the pairwise fusion detection network includes multiple stages, successively including an improved Transformer encoder, a pairwise fusion module, a Transformer decoder, and an MLP layer.

[0074] Among them, for the improved Transformer encoder, the input features are the object query vector and the pairwise box position encoding respectively. The pairwise box position encoding is a vector composed of the coordinates, length, width, and intersection over union (IoU) of the corresponding pair of target boxes of the person and the object as the input, and is obtained through a linear layer and a ReLU activation function. The pairwise box position encoding is a relatively mature technology in this field and will not be elaborated here.

[0075] As Figure 7As shown, in the improved Transformer encoder, the object query vectors are paired and combined with the paired box position encodings. Through a linear layer and a sigmoid activation function, the output of the first branch is obtained; the object query vectors are copied and multiplied by the elements of the paired box position encodings to obtain the output of the second branch; the elements of the outputs of the two branches are multiplied, passed through a linear layer, added to the input object query vectors, and then passed through a normalization layer (layerNorm layer), a feed-forward network (FFN), and a normalization layer to output the paired query vectors.

[0076] Among them, the pairing operation pairs the query vectors of people and objects pairwise, changing the query vector dimension from 256 to 512, while the copying operation copies the query vectors of people and objects respectively, also changing the query vector dimension from 256 to 512. Thus, while retaining the features of a single query vector, both encode pairwise feature information.

[0077] As Figure 8 shown, in the pairwise fusion module, the paired query vectors are respectively combined (concatenated) with the paired box position encodings and the global visual features after adaptive average pooling, multiplied after passing through a linear layer, and then successively passed through a ReLU activation function, a linear layer, and a ReLU activation function to obtain the final paired query vectors that fuse multiple features.

[0078] Finally, the paired query vectors that fuse multiple features are decoded by the Transformer decoder and output to the MLP to obtain the probability scores of the human-object interaction actions, thereby completing the detection of human-object interaction actions.

[0079] In a specific embodiment, this application also calculates the overall loss function of the network during training, performs backpropagation, and updates the network parameters. Among them, the overall loss function of the network is linearly fused by the verb cross-entropy loss function L a and the semantic relative entropy loss function L SKL where:

[0080] The verb cross-entropy loss function L a is:

[0081]

[0082] Among them, N q represents the number of verb categories, represents the number of predicted verb categories corresponding to the objects, Φ represents the set of all ground-truth (true values), represents in the prediction set. l fIt is the Focal loss, and the specific manifestation of the Focal loss is l f (p t ) = -α t (1 - p t ) γ log(p t ), where α t is the parameter to suppress the imbalance between positive and negative samples, γ is the parameter to control the imbalance in the number of easy / hard samples, and p t is the sample. Therefore, in the verb loss, l f is used to calculate the loss between the true verb category and the predicted verb category, where represents the true verb category.

[0083] The semantic relative entropy loss function L SKL is as follows:

[0084]

[0085] where is the symmetric conditional distribution of verbs in the dataset, A is the adjacency matrix of the verb embeddings processed by the semantic space attention module, is the KL divergence loss function.

[0086] can be obtained through the following calculation

[0087]

[0088] where N p is the number of verbs in the dataset, and c ij is

[0089]

[0090] A can be obtained through the following calculation:

[0091]

[0092] where τ is the temperature parameter for scaling the normalized semantic inner product softmax distribution, is a verb embedding processed by the semantic space attention module, and T is the transpose symbol.

[0093] The overall loss function of the network is:

[0094] L total = L a + L SKL

[0095] where Ltotal Represents the overall loss function, L a and L SKL respectively represent the verb cross-entropy loss function and the semantic relative entropy loss function.

[0096] The above-described embodiments merely represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A person-object interaction detection method based on DETR for paired decoding interaction of people, characterized in that, The person-object interaction detection method based on DETR's paired decoding interaction includes: Inject the feature map obtained from the original image through the backbone network into the trained DETR network. The DETR network includes an encoder, a decoder, and an MLP layer to obtain the query vector output by the decoder, as well as the target box and target category finally output by the DETR network. Input the query vector and target category into the query vector classifier to obtain the query vector of the person, the query vector of the object, and the category of the object. Input the category of the object into the semantic network to obtain the semantic query vector of the object. Fuse the query vector of the object and the semantic query vector of the object to obtain the fused object query vector, and merge the fused object query vector and the query vector of the person to obtain the object query vector. Input the object query vector into the paired fusion detection network to achieve person-object interaction detection. Among them, the paired fusion detection network sequentially includes an improved Transformer encoder, a paired fusion module, a Transformer decoder, and an MLP layer. For the improved Transformer encoder, the input features are the object query vector and the paired box position encoding. In the improved Transformer encoder, the object query vectors are paired and combined with the paired box position encoding, and through a linear layer and a sigmoid activation function, the output of the first branch is obtained; the object query vector is copied and multiplied by the elements of the paired box position encoding to obtain the output of the second branch; the output elements of the two branches are multiplied and then passed through a linear layer, added to the input object query vector, and then passed through a normalization layer, a forward propagation layer, and a normalization layer to output the paired query vector. In the paired fusion module, the paired query vectors are respectively combined with the paired box position encoding and the global visual features obtained through adaptive average pooling, and then multiplied after passing through a linear layer, and then sequentially passed through a ReLU activation function, a linear layer, and a ReLU activation function to obtain the final paired query vector that fuses multiple features. The paired query vector that fuses multiple features is decoded by the Transformer decoder and then output to the MLP to obtain the probability score of the person-object interaction action, thereby completing the detection of the person-object interaction action.

2. The method for detecting human-object interaction based on DETR-based paired decoding interaction of people according to claim 1, characterized in that The semantic network includes a semantic space attention module and a semantic aggregation module. The input feature of the semantic space attention module is the verb embedding vector of the dataset, and the semantic space attention feature is output. The input feature of the semantic aggregation module is the semantic space attention feature output by the semantic space attention module and the category of the object output by the query vector classifier. The semantic space attention feature passes through a linear layer, a ReLU activation function, a linear layer, and a sigmoid activation function to obtain the attention feature, which is multiplied by the feature obtained by the category of the object through another linear layer. The result is sequentially passed through a linear layer, a normalization layer, a ReLU activation function, and a linear layer, and then added to the category of the object, and then input to the Transformer layer to obtain the semantic query vector of the object.

3. The method for human-object interaction detection based on DETR's paired decoding interaction for humans as claimed in claim 1, wherein The fusion of the query vector of the object and the semantic query vector of the object to obtain the fused object query vector includes: After adding the query vector of the object and the semantic query vector of the object, pass it through the ReLU activation function, and subtract the square of the difference between the query vector of the object and the semantic query vector of the object.

4. The method for detecting human-object interaction based on DETR-based paired decoding interaction of humans as claimed in claim 1, wherein, The person-object interaction detection method based on DETR for paired decoding and interaction of people further includes: Calculate the overall loss function of the network, perform backpropagation, and update the network parameters; Among them, the overall loss function of the network is: L total = L a + L SKL Among them, L total represents the overall loss function, and L a and L SKL represent the verb cross-entropy loss function and the semantic relative entropy loss function respectively; The verb cross-entropy loss function L a is as follows: Among them, N q represents the number of types of verbs, represents the number of predicted verb categories corresponding to the statistics and the object, and Φ represents the set of all true values, represents that in the prediction set, l f is the focal loss, and l f (p t ) = -α t (1 - p t ) γ log(p t ), where α t is the parameter for suppressing the imbalance of positive and negative sample parameters, γ is the parameter for controlling the imbalance of the number of easy / difficult samples, and p t is the sample, where represents the true verb category; The semantic relative entropy loss function L SKL is as follows: Among them is the conditional distribution of verb symmetry in the dataset, and A is the adjacency matrix of the verbs processed by the semantic space attention module. is the KL divergence loss function. It can be obtained through the following calculation: where N p is the number of verbs in the dataset, and c ij is defined as: A can be obtained through the following calculation: where τ is the temperature parameter for scaling the normalized semantic inner product softmax distribution, is the verb embedding vector processed by the semantic space attention module, and T is the transpose symbol.