A Transformer-based Social Relationship Recognition Method

Social relationship features are extracted through the fully connected network and the Transformer encoder network, and social relationship recognition is combined with residual embedding blocks and graph neural networks, which solves the problems of insufficient feature extraction and rough fusion methods, and achieves more efficient social relationship recognition.

CN115858943BActive Publication Date: 2025-07-25SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111116796.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-23
Publication Date
2025-07-25
Estimated Expiration
2041-09-23

AI Technical Summary

Technical Problem

The existing social relationship recognition methods have insufficient feature extraction, too rough feature fusion method, and unreasonable inference recognition classification, resulting in low accuracy of social relationship recognition.

Method used

The fully connected network, Vision Transformer network and ResNet-50 network are used to extract character pairs and scene characteristics, and the intrinsic connections between the features through the Transformer encoder network are used to infer information, and the residual embedding block is introduced to transmit information, construct a social relationship-scene graph for graph reasoning, and finally feature fusion and classification are performed through residual connections.

Benefits of technology

It improves the accuracy and rationality of social relationship recognition, fully explores the inner connection between characters and features, saves training time, and improves the recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115858943B_ABST
    Figure CN115858943B_ABST
Patent Text Reader

Abstract

The present invention proposes a social relationship recognition method based on Transformer, which mainly involves the problem of extracting and reasoning related features and performing social relationship recognition through Transformer in deep learning. First, relevant features are extracted through the feature extraction module; then, the feature reasoning module is used to infer the connection between features and form relationships and scene nodes, and the residual block is introduced to pass the character pair information to the classification layer; finally, a fully connected social relationship-scene graph is constructed, and the residual block is fused after graph reasoning using GGNN to classify social relationships. The present invention fully considers the feature extraction related to social relationship recognition, uses Transformer to explore the intrinsic connection between character pair features and introduces the logical relationship between graph neural network reasoning relationships, which solves the problems of insufficient feature extraction, overly rough feature fusion method, and unreasonable reasoning recognition and classification in social relationship recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the problem of social relationship recognition in the field of deep learning, and particularly to a social relationship recognition method based on Transformer. Background Art

[0002] In the field of computer vision, social relationship recognition is an important task for studying the social relationships between people, providing important clues for understanding human interaction behaviors. Most of the existing studies at present are based on relevant features such as faces, body regions, scenes, etc. to conduct social relationship recognition, and certain achievements have been made. In recent years, graph neural networks designed specifically for graph-structured data have developed rapidly and have also promoted other fields. Therefore, some researchers have introduced them into the field of social relationship recognition to simulate human thinking to reason about the relationships between people and objects in a scene to improve the accuracy of social relationship recognition. In addition, relying on its own powerful reasoning ability for information, the network based on the Transformer structure has made great breakthroughs in both the field of natural language processing and the field of computer vision, and thus has also shown certain potential in the field of social relationship recognition in computer vision. At present, social relationship recognition plays an important role in fields such as photo classification, group division, and crowd activity analysis.

[0003] Social relationship recognition, as an important research task in the field of computer vision, has received extensive attention from relevant researchers at home and abroad. Existing methods often use traditional convolutional neural networks as the backbone network for feature extraction, thus paying more attention to local information and being unable to effectively extract the interactive information of person pairs hidden in global information. In addition, most methods only adopt a simple splicing and fusion method for the extracted feature vectors, or directly splice them as the nodes of the graph neural network, unable to fully explore the internal connections between person pair features, and the complex network reduces the attention to person pair-related features. Therefore, in this patent, the relative position features of person pairs, the features of each person in the person pair, the features of the common area of the person pair, and the scene features of the entire image are extracted at one time through a fully connected layer, two weight-sharing Vision Transformer networks, a parameter-independent Transformer network, and a ResNet-50 network; then, the Transformer decoder network is used to reason about the internal connections between person pair-related features, explore the internal connections between features, and introduce a residual embedding block to transfer person pair-related feature information to the classification layer; subsequently, all features except scene features are fused to form social relationship nodes, and the extracted scene features are used as scene nodes. Then, these social relationship nodes and scene nodes are connected in a fully connected manner to form a social relationship graph with scenes introduced and sent into the graph neural network for graph reasoning; finally, the scene nodes are removed, and the person pair-related information is fused with the social relationship nodes after graph reasoning through residual connection, and the fused information is classified to improve the rationality and recognition accuracy of social relationships. Summary of the Invention

[0004] The purpose of the present invention is to provide a social relationship recognition method based on Transformer. First, the Vision Transformer network and the fully connected network are used to fully extract the person pair features related to social relationship recognition, and the ResNet-50 network is used to extract scene features. Then, the Transformer decoder is used to reason about the internal connections between person pair-related features and introduce a residual embedding block to transfer person pair feature information. Next, a social relationship-scene graph form is adopted to introduce the graph neural network to fuse and reason about the processed features. Finally, a residual connection is introduced to fuse the person pair feature information with the social relationship node information after graph reasoning for reasonable classification, effectively solving the problems of insufficient feature extraction, overly rough feature fusion method, and unreasonable reasoning and recognition classification in social relationship recognition.

[0005] For the convenience of description, the following concepts are introduced first:

[0006] Pre-trained model: The training of a neural network requires a large amount of data, time, and sufficient computing resources. To avoid repeated training of the network, the model parameters of a model with good performance trained by other researchers are transferred to the model for a specific task and fine-tuned to meet the requirements of this task.

[0007] Transformer: A deep learning network that extracts input data in a parallel manner by adopting the self-attention mechanism and generally consists of two parts: a decoder and an encoder.

[0008] Vision Transformer (ViT): A Transformer network application extended from the field of natural language processing to the field of computer vision. This network divides an image into equal-sized patches, compresses them into one dimension as the input of the Transformer encoder, and uses the self-attention mechanism of the Transformer itself to globally extract features.

[0009] Graph: Refers to the graph in graph theory, which is a graph in a non-Euclidean space and consists of nodes (Nodes) and edges (Edges) connecting the nodes.

[0010] Graph Neural Network (GNN): A neural network structure that directly computes on a graph. It learns the representation of nodes through message passing, updates the information of the current node with adjacent nodes until the entire graph converges to a stable state.

[0011] Gated Graph Neural Network (GGNN): To address the application limitations brought by the fixed-point theory (Banach's Fixed Point Theorem) in traditional graph neural networks, a new graph neural network is formed by introducing the update method of the Gated Recurrent Unit (GRU).

[0012] Deep Residual Network (ResNets): A deep learning network that addresses the side effects caused by network depth by introducing residual blocks; it is divided into ResNet-18, ResNet-34, ResNet-50, ResNet-101, and ResNet-152 according to the number of network layers.

[0013] The present invention specifically adopts the following technical solutions:

[0014] A social relationship recognition method based on Transformer, characterized in that:

[0015] a. Extracting character features and scene features associated with social relationship recognition through a fully connected network, Vision Transformer network, and convolutional neural network;

[0016] b. Introduce the Transformer encoder network to infer the intrinsic relationship between each character feature, and pass the character feature information by introducing a residual embedding block;

[0017] c. Construct a social relationship-scene graph in a non-Euclidean space, and use graph neural networks to infer the connections between social relationships and between scenes and social relationships;

[0018] d. Integrate the global character feature information with the social relationship information after graph reasoning in a fully connected manner, increase the attention to character feature information, and enhance the rationality of social relationship recognition;

[0019] The method mainly includes the following steps:

[0020] (1) Data processing and enhancement: The bounding box regions of the two characters and the joint region of a character pair as input are uniformly cropped to a size of 224×224, and the entire image is cropped to a size of 448×448. The cropped images are normalized and randomly flipped horizontally. In addition, the position information and area information of the bounding boxes of the two characters are normalized and used as one input.

[0021] (2) Feature extraction: The relative position features of the person pairs, the features of each person in the person pairs, the features of the common areas of the person pairs, and the scene features of the entire image are extracted in sequence through a fully connected layer in the node generation model, two weight-shared and pre-trained Vision Transformer networks, a parameter-independent and pre-trained Vision Transformer network, and a pre-trained ResNet-50 network. It should be noted that since scene feature extraction is relatively simple and there is a scene recognition model for the ResNet-50 network, ResNet-50 is used as the network for scene information extraction.

[0022] (3) Character feature reasoning: The four character features in the feature extraction module are inferred through the Transformer encoder network to explore the intrinsic relationship between the four character features. In addition, a learnable residual embedding block is introduced to extract and transmit the global character feature information inferred by the Transformer encoder to the classification layer using a multi-head self-attention mechanism. It should be noted that the output of the network is a residual embedding block and four new feature blocks.

[0023] (4) Node generation: For the relative position features of the person pairs inferred by the person feature inference module, the features of each person in the person pair, and the features of the common area of the person pair, a fully connected layer is used for fusion to form social relationship nodes, and the scene features extracted by the ResNet-50 network are used as scene nodes; it should be noted that for each two people in an RGB image, a social relationship node is formed, but there is only one scene node, and the social relationship node contains not only the information related to the person pair, but also the internal connection between the relevant features of the person pair;

[0024] (5) Graph construction and inference: The social relationship nodes and scene nodes generated in step (4) are connected in a fully connected manner to construct a social relationship graph introducing the scene, and the graph is inferred through a gated graph neural network, while mining the connections between social relationships and between social relationships and the scene;

[0025] (6) Social relationship classification: Remove the scene node, and fuse the residual embedding block in step (3) and the social relationship nodes after graph construction and inference in step (5) in a fully connected manner to improve the attention to the feature information of the person pair and increase the rationality of social relationship recognition; in addition, the fused information is classified through an additional fully connected layer.

[0026] The beneficial effects of the present invention are as follows:

[0027] (1) Make full use of the Vision Transformer pre-trained model for feature extraction, saving a large amount of training time and improving the social relationship recognition effect.

[0028] (2) Extract the internal connection between the relevant feature information of the person pair through the Transformer encoder network, that is, the new middle-level information.

[0029] (3) Infer the connections between social relationships and between social relationships and the scene in the form of a graph, effectively mining the information between relationships.

[0030] (4) Introduce an additional residual embedding block into the input of the Transformer encoder network to transfer the global person pair-related feature information to the classification layer, improving the rationality of social relationship recognition. Description of the Drawings

[0031] Figure 1 It is a schematic diagram of the person feature inference module.

[0032] Figure 2 It is a schematic diagram of the graph construction and inference module.

[0033] Figure 3 It is a schematic diagram of the overall model framework. Detailed Implementation Manner

[0034] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and should not be construed as limiting the protection scope of the present invention. Those skilled in the art can make some non-essential improvements and adjustments to the present invention according to the above-mentioned inventive content and still fall within the protection scope of the present invention.

[0035] A social relationship recognition method based on Transformer and graph neural network specifically includes the following steps:

[0036] (1) Data Processing and Enhancement

[0037] Uniformly crop the two input person bounding box regions and a person pair joint region to a size of 224×224, crop the entire image to a size of 448×448, and perform normalization processing and random horizontal flipping on the cropped image. In addition, the position information and area information of the two person bounding boxes are normalized and used as one input path, which is specifically expressed as:

[0038]

[0039] Among them, x min_A , y min_A , x max_A , y max_A respectively represent Figure 1 the coordinates of the upper left and lower right corners of the bounding box of person A in A , and area min_B , y min_B , x max_B , y max_B , area B correspond to the coordinates and area of the bounding box of person B. When this vector is used as the network input, its respective coordinates and area are normalized to [-1, 1].[[]END]

[0040] (2) Feature Extraction

[0041] Through Figure 3In the feature extraction module, a fully connected layer, two Vision Transformer networks with shared weights and pre-trained on the ImageNet dataset, a Vision Transformer network with independent parameters and pre-trained on ImageNet, and a ResNet-50 network pre-trained on the Places365-Standard dataset are used to extract the relative position features of the person pair, the features of each person in the person pair, the features of the common area of the person pair, and the scene features of the entire image in sequence. It should be noted that the last classification layers in the Vision Transformer network and the ResNet-50 network are removed because they are only used as feature extraction networks.

[0042] (3) Inference of person features

[0043] As Figure 1 shown, the extracted relative position features of the person pair, the features of each person in the person pair, the features of the common area of the person pair, and the scene features of the entire image are used as four feature vectors, which together with a learnable additional residual embedding block are used as the input of the Transformer encoder network. The self-attention mechanism of the kernel of the Transformer encoder enables the network to explore the internal relationships between the person pair-related features. At the same time, the residual embedding block also obtains the global person pair-related information as the content transmitted through the residual connection.

[0044] In addition, the four feature vectors and the learnable residual embedding block are regarded as the five inputs of a sentence sequence, and all are artificially assigned position information embeddings (Position Embedding, PE). Among them, the residual embedding block is assigned position 0, and the first person region feature, the second person region feature, the person pair joint region feature, and the person pair relative position feature are assigned positions 1, 2, 3, and 4 respectively.

[0045] (4) Node generation

[0046] As Figure 2 shown, a fully connected layer is used to fuse the relative position features of the person pair, the features of each person in the person pair, and the features of the common area of the person pair to form a social relationship node, and the scene features extracted by the ResNet-50 network are used as the scene node. It should be noted that for every two people in an RGB image, a social relationship node is formed, but there is only one scene node. As Figure 2 shown, if there are four people in the input picture, they are paired up in six pairs to form six social relationship nodes, and the scene features of the entire picture are extracted as one relationship node.

[0047] In addition, both the social relationship nodes and the scenario nodes are uniformly transformed to 512 dimensions in the last layer of the network generated by their respective nodes, that is, the node representations of all inputs to the graph neural network are 512 dimensions.

[0048] (5) Graph construction and reasoning

[0049] Such as Figure 2 Therefore, the social relationship nodes and scenario nodes generated in step (4) are connected in a fully connected manner to construct a social relationship graph incorporating the scenario. When there are four characters in the input image, its social relationship graph is as Figure 2 shown. After forming the social relationship graph incorporating the scenario, it is fed into a gated graph neural network for graph reasoning to fully exploit the logical relationships between social relationships and transfer the information contained in the scenario nodes to the relationship nodes.

[0050] The number of layers of the gated graph neural network in this part is 3, and the input and output of each layer are 512 dimensions.

[0051] (6) Social relationship classification

[0052] After fully reasoning about the graph, the scenario nodes are removed, and the residual embedding block in step (3) and the social relationship nodes after graph construction and reasoning in step (5) are fused in a fully connected manner to enhance the attention to the feature information of the person pairs and increase the rationality of social relationship recognition; in addition, the fused information is classified through an additional fully connected layer.

Claims

1. A social relationship recognition method based on Transformer, characterized by: a. Extracting character features and scene features associated with social relationship recognition through a fully connected network, Vision Transformer network, and convolutional neural network; b. Introduce the Transformer encoder network to infer the intrinsic relationship between each character feature, and pass the character feature information by introducing a residual embedding block; c. Construct a social relationship-scene graph in a non-Euclidean space, and use graph neural networks to infer the connections between social relationships and between scenes and social relationships; d. Integrate the global character feature information with the social relationship information after graph reasoning in a fully connected manner to increase the focus on character feature information; The method comprises the following steps: (1) Data processing and enhancement: The bounding box regions of the two characters and the joint region of a character pair as input are uniformly cropped to a size of 224×224, and the entire image is cropped to a size of 448×448. The cropped images are normalized and randomly flipped horizontally. In addition, the position information and area information of the bounding boxes of the two characters are normalized and used as one input. (2) Feature extraction: The relative position features of the person pairs, the features of each person in the person pairs, the features of the common areas of the person pairs, and the scene features of the entire image are extracted in turn through a fully connected layer in the node generation model, two weight-shared and pre-trained Vision Transformer networks, a parameter-independent and pre-trained Vision Transformer network, and a pre-trained ResNet-50 network; ResNet-50 is used as the network for scene information extraction; (3) Character feature reasoning: The four character features in the feature extraction module are inferred through the Transformer encoder network to explore the intrinsic relationship between the four character features. In addition, a learnable residual embedding block is introduced to extract and pass the global character feature information inferred by the Transformer encoder to the classification layer using the multi-head self-attention mechanism. The output of the network is a residual embedding block and four new feature blocks. (4) Node generation: The relative position features of the person pairs inferred by the person feature inference module, the features of each person in the person pairs, and the features of the common areas of the person pairs are fused using a fully connected layer to form social relationship nodes. The scene features extracted by the ResNet-50 network are used as scene nodes. In an RGB image, one social relationship node is formed for every two people, but there is only one scene node. (5) Graph construction and reasoning: The social relationship nodes and scene nodes generated in step (4) are connected in a fully connected manner to construct a social relationship graph that introduces the scene, and the graph is reasoned through a gated graph neural network to mine the connections between social relationships and between social relationships and scenes. (6) Social relationship classification: Remove the scene node, and fuse the residual embedding block in step (3) and the social relationship node after graph construction and reasoning in step (5) in a fully connected manner to improve the attention to the feature information of the person pair and increase the rationality of social relationship recognition; in addition, classify the fused information through an additional fully connected layer.

2. The Transformer-based social relationship recognition method according to claim 1, wherein In step (2), extract the features within two single-person bounding boxes through two Vision Transformer networks with shared parameters; extract the features within a person pair joint region through a separate Vision Transformer network; extract the scene features through a ResNet-50 network applicable to the scene classification task.

3. The social relationship recognition method based on Transformer according to claim 1, characterized in that In step (3), for the person pair related features extracted in step (2), regard different features as different words in a sentence sequence, and then infer the interaction between the person pair related features in the sequence through a Transformer encoder network to explore the connections between different features.

4. The method for identifying social relationships based on Transformer according to claim 1, wherein In step (4), fuse the output of step (3) passing through the Transformer encoder network except for the residual embedding block in a fully connected manner, so that the generated social relationship node contains the internal connection between the person pair related features.

5. The social relationship recognition method based on Transformer according to claim 1, characterized in that In step (6), fuse the residual embedding block described in step (3) and the output after reasoning through the gated graph neural network described in step (5) in a fully connected manner, and transfer the person pair related information to the final classification layer to improve the attention to the person pair related feature information.

Citation Information

Patent Citations

  • Multi-relation collaborative filtering algorithm based on graph neural network

    CN111523047A

  • Document analyzing apparatus and method thereof

    US20100049499A1