Visual relation extraction method based on attention intersection perception

CN118674974BActive Publication Date: 2026-09-29NAT UNIV OF DEFENSE TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410684305.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2026-09-29
Estimated Expiration
2044-05-30

AI Technical Summary

Technical Problem

然而,对于物体多的图像而言,想要抽取出其中两个物体的视觉关系,不一定需要利用上整张图片的信息,而只需要利用上这两个物体相关的信息,其余的不管信息反而会对模型的判断产生干扰

Benefits of technology

[0051]与现有方法相比,本发明方法的优点在于:本技术提供了基于注意力交集感知的视觉关系抽取方法,本方法利用视觉注意力感知机制,创新性地提出了注意力交集感知模块,通过建模目标间的注意力热图交集,提取出每对视觉目标的共同关注特征,使模型能够关注到图片的关键信息,降低无效信息的干扰,提高视觉关系抽取的准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118674974B_ABST
    Figure CN118674974B_ABST
Patent Text Reader

Abstract

The application discloses a visual relation extraction method based on attention intersection perception, and the method comprises the following steps: constructing a visual relation extraction model, including a pre-training visual model, an attention intersection perception module and a relation prediction layer; extracting picture features by using the pre-training visual model; obtaining feature representations of all targets according to picture features and target region labels; calculating common attention features of each pair of targets by using the attention intersection perception module; inputting the feature representations of each pair of targets and the common attention features into the relation prediction layer to perform relation prediction; calculating a relation prediction loss function, optimizing the visual relation extraction model by using the relation prediction loss function, and performing visual relation extraction by using the optimized visual relation extraction model. The application uses a visual attention perception mechanism, innovatively proposes an attention intersection perception module, extracts common attention features of each pair of visual targets, and enables the model to pay attention to key information of a picture, thereby improving the visual relation extraction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of deep learning and image processing, and in particular to a method for visual relation extraction based on attention intersection perception. Background Technology

[0002] Visual relation extraction is a task in computer vision that aims to identify and understand the relationships between different objects in an image. It involves modeling and inferring spatial, functional, and semantic relationships between objects. Typically, the input is an image and the locations of objects within it, and the output is labels or categories describing the relationships between objects. For example, in an image containing a person and a ball, the visual relation extraction task can classify the relationship between the person and the ball as "holding," "kicking," or "contacting."

[0003] The value of visual relation extraction lies in improving computers' understanding and reasoning abilities regarding image scenes, supporting visual search and other computer vision tasks, and playing a crucial role in practical applications such as scene understanding, visual reasoning, visual search, and computer-aided vision tasks. However, for images with many objects, extracting the visual relationship between two objects doesn't necessarily require utilizing the information from the entire image; only the information related to those two objects is needed. Ignoring other information can interfere with the model's judgment. Guiding the model to focus on local information related to objects is a feasible technical path to improve the performance of visual relation extraction. Summary of the Invention

[0004] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention discloses a visual relationship extraction method based on attention intersection perception. This method can identify the relationships between targets in an image. Compared with existing methods, this method innovatively proposes an attention intersection perception module by utilizing a visual attention perception mechanism to extract the common attention features of each pair of visual targets, enabling the model to focus on key information in the image and improving the accuracy of visual relationship extraction.

[0005] The objective of this invention is achieved through the following technical solution: a visual relationship extraction method based on attention intersection perception, the method comprising:

[0006] Step 1: Construct a visual relationship extraction model, including a pre-trained visual model, an attention intersection perception module, and a relationship prediction layer;

[0007] Step 2: Extract image features using a pre-trained visual model;

[0008] Step 3: Obtain feature representations of all targets based on image features and target region annotations;

[0009] Step 4: Use the attention intersection perception module to calculate the common attention features for each pair of targets;

[0010] Step 5: Input the feature representations and common interest features of each pair of targets into the relationship prediction layer to perform relationship prediction;

[0011] Step 6: Calculate the relationship prediction loss function, optimize the visual relationship extraction model using the relationship prediction loss function, and use the optimized visual relationship extraction model to extract visual relationships.

[0012] Specifically, the extraction of image features using a pre-trained visual model includes the following steps:

[0013] The input image is represented as The visual features of the input image are extracted using a pre-trained visual model, expressed as follows:

[0014] F img =ResNet(x)

[0015] in, The image is represented by its visual features, W and H represent the width and height of the image, d represents the hidden layer dimension of ResNet, and s represents the downsampling factor of ResNet.

[0016] Specifically, obtaining feature representations of all targets based on image features and target region annotations includes the following steps:

[0017] Based on the target area annotation, the target's position is represented as the coordinates (x, y) of a bounding box. i ,y i ,w i ,h i ), where i represents the index of the target, x i and y i w represents the coordinates of the top-left corner of the rectangle. i and h i This indicates the width and height of the rectangle.

[0018] The ROIAlign algorithm is used to map the target region annotations into the image feature space, resulting in a visual feature map of the target, expressed as:

[0019]

[0020] in, Let d represent the visual feature map of the i-th target, d represent the hidden layer dimension of the ResNet, and s represent the hidden layer dimension of the target. o Indicates the output size of ROIAlign;

[0021] Adaptive pooling is performed on the visual feature map of the i-th target to obtain the feature representation of the i-th target, expressed as:

[0022]

[0023] in, The feature representation of the i-th target;

[0024] Visual features are extracted for each target to obtain feature representations for all targets:

[0025]

[0026] in, Let n represent the characteristic representation of all targets, and n represent the number of targets.

[0027] Furthermore, the aforementioned use of the attention intersection perception module to calculate common attention features for each pair of targets includes the following steps:

[0028] Step 401, calculate the attention heatmap for each target; the feature representation of the i-th target is Will After mapping, the visual features F of the image are... img The attention heatmap is calculated using the following expression:

[0029]

[0030] in, W represents the attention heatmap for the i-th target. a These are learnable parameters, and softmax is the activation function used to normalize a vector into a probability distribution.

[0031] Step 402, calculate the attention intersection of each pair of targets; multiply the attention of each pair of targets bitwise, the expression is:

[0032]

[0033] in, Let A represent the intersection of attention between the i-th and j-th targets. i A represents the attention heatmap of the i-th target. j Attention heatmap of the j-th target;

[0034] Step 403: Calculate the common interest features for each pair of targets; aggregate the common interest features of each pair of targets using the attention intersection, expressed as:

[0035]

[0036] Among them, h i,j C represents the common feature of interest for the i-th and j-th objectives. i,j,klC represents the intersection of attention between the i-th and j-th targets. i,j The element in the k-th row and l-th column, F img,kl The element in the k-th row and l-th column represents the visual features of the image;

[0037] For each target, perform step 401; for each pair of targets, perform steps 402 and 403 to obtain the common features of interest for each pair of targets.

[0038] Furthermore, the step of inputting the feature representations and common interest features of each pair of targets into the relationship prediction layer for relationship prediction includes the following steps:

[0039] It is the feature representation of the i-th target. h is the feature representation of the j-th target. i,j This indicates the common features of interest for the i-th and j-th objectives. and h i,j After concatenation, the data is input into a fully connected layer to calculate the probability vector of the relationship between the i-th target and the j-th target. The expression is as follows:

[0040]

[0041] in, Let W represent the predicted probability vector of the relationship between the i-th target and the j-th target, where l represents the number of relationship categories. o and b o These are the learnable parameters of the fully connected layer, and softmax is the activation function used to normalize a vector into a probability vector.

[0042] Preferably, the calculation of the relationship prediction loss function, the optimization of the visual relationship extraction model using the relationship prediction loss function, and the performance of visual relationship extraction using the optimized visual relationship extraction model include the following steps:

[0043] Step 601, construct the relationship label q between the i-th target and the j-th target. ij q ij q is a one-hot vector of length l, which is the number of relation categories. When the i-th target and the j-th target have a k-th relation, q... ijk =1.

[0044] Step 602: Calculate the relationship prediction loss function L between the i-th target and the j-th target. ij The expression is:

[0045]

[0046] Where, p ijkL represents the k-th element of the predicted probability vector of the relationship between the i-th target and the j-th target. ij The prediction loss function represents the relationship between the i-th objective and the j-th objective;

[0047] Step 603, calculate the relationship prediction loss function, the expression of which is:

[0048]

[0049] Where L is the relationship prediction loss function;

[0050] Step 604: Based on the relationship prediction loss function L, use the Adam optimizer to optimize the visual relationship extraction model until L converges, and obtain the final model parameters; use the optimized visual relationship extraction model to perform visual relationship extraction.

[0051] Compared with existing methods, the advantages of the present invention are as follows: This technology provides a visual relationship extraction method based on attention intersection perception. This method innovatively proposes an attention intersection perception module by utilizing the visual attention perception mechanism. By modeling the intersection of attention heatmaps between targets, it extracts the common attention features of each pair of visual targets, enabling the model to focus on the key information of the image, reducing the interference of invalid information, and improving the accuracy of visual relationship extraction. Attached Figure Description

[0052] Figure 1 A flowchart illustrating an embodiment of the present invention is shown. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0054] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0055] In this embodiment, we assume we are building an intelligent security system. This system aims to automatically detect and identify relationships between objects by analyzing real-time video streams captured by surveillance cameras, thereby providing more intelligent and efficient security monitoring functions. The system can identify relationships between people and objects in images, such as "carrying," "placing," and "pushing," using visual relationship extraction technology. When the system detects abnormal behavioral relationships, such as "carrying a large object into a prohibited area" or "placing an object in a dangerous location," it can automatically trigger an alarm and notify security personnel. For the task of visual relationship extraction, a visual relationship extraction method based on attention intersection perception is used.

[0056] Therefore, as Figure 1 As shown, a visual relation extraction method based on attention intersection perception includes:

[0057] Step 1: Construct a visual relationship extraction model, including a pre-trained visual model, an attention intersection perception module, and a relationship prediction layer;

[0058] Step 2: Extract image features using a pre-trained visual model;

[0059] Step 3: Obtain feature representations of all targets based on image features and target region annotations;

[0060] Step 4: Use the attention intersection perception module to calculate the common attention features for each pair of targets;

[0061] Step 5: Input the feature representations and common interest features of each pair of targets into the relationship prediction layer to perform relationship prediction;

[0062] Step 6: Calculate the relationship prediction loss function, optimize the visual relationship extraction model using the relationship prediction loss function, and use the optimized visual relationship extraction model to extract visual relationships.

[0063] The method of extracting image features using a pre-trained visual model includes the following steps:

[0064] The input image is represented as The visual features of the input image are extracted using a pre-trained visual model, expressed as follows:

[0065] F img =ResNet(x)

[0066] in, The image is represented by its visual features, W and H represent the width and height of the image, d represents the hidden layer dimension of ResNet, and s represents the downsampling factor of ResNet.

[0067] ResNet is a deep convolutional neural network architecture proposed by Kaiming He et al. from Microsoft Research. Its main innovation is the introduction of residual connections, also known as skip connections, to solve the vanishing and exploding gradient problems in deep network training.

[0068] The introduction of ResNet is of great significance in deep learning. Through the design of residual connections, it allows for the construction of very deep networks, resulting in significant performance improvements in large-scale image classification and other computer vision tasks. ResNet's flexibility and powerful feature representation capabilities have made it an important foundational model in many fields.

[0069] By pre-training on large-scale image datasets (such as ImageNet), ResNet can learn a set of general image feature representations. These features can capture low-level visual features (such as edges and textures) and high-level semantic features (such as object shapes and parts) in images. This pre-training gives the network a better initial feature extraction capability, providing a stronger foundation for subsequent tasks.

[0070] The method of obtaining feature representations of all targets based on image features and target region annotations includes the following steps:

[0071] Based on the target area annotation, the target's position is represented as the coordinates (x, y) of a bounding box. i ,y i ,w i ,h i ), where i represents the index of the target, x i and y i w represents the coordinates of the top-left corner of the rectangle. i and h i This indicates the width and height of the rectangle.

[0072] The ROIAlign algorithm is used to map the target region annotations into the image feature space, resulting in a visual feature map of the target, expressed as:

[0073]

[0074] in, Let d represent the visual feature map of the i-th target, d represent the hidden layer dimension of the ResNet, and s represent the hidden layer dimension of the target. o Indicates the output size of ROIAlign;

[0075] RoIAlign is a feature alignment method for object detection and recognition tasks. It addresses the issues of information loss and accuracy degradation that can occur when using RoIPooling.

[0076] In object detection tasks, a common approach is to extract regions of interest (ROIs) and project them onto a feature map, then use the RoIPooling operation to map each ROI onto a fixed-size feature map region. However, RoIPooling is based on rounding operations, which can lead to subtle feature location shifts, thus reducing localization accuracy.

[0077] RoIAlign addresses this issue by using bilinear interpolation to extract features of RoI regions from the feature map in a more accurate manner.

[0078] Adaptive pooling is performed on the visual feature map of the i-th target to obtain the feature representation of the i-th target, expressed as:

[0079]

[0080] in, The feature representation of the i-th target;

[0081] Visual features are extracted for each target to obtain feature representations for all targets:

[0082]

[0083] in, Let n represent the characteristic representation of all targets, and n represent the number of targets.

[0084] The aforementioned use of the attention intersection perception module to calculate common attention features for each pair of targets includes the following steps:

[0085] Step 401, calculate the attention heatmap for each target; the feature representation of the i-th target is Will After mapping, the visual features F of the image are... img The attention heatmap is calculated using the following expression:

[0086]

[0087] in, W represents the attention heatmap for the i-th target. a These are learnable parameters, and softmax is the activation function used to normalize a vector into a probability distribution.

[0088] Step 402, calculate the attention intersection of each pair of targets; multiply the attention of each pair of targets bitwise, the expression is:

[0089]

[0090] in, Let A represent the intersection of attention between the i-th and j-th targets. i A represents the attention heatmap of the i-th target. j Attention heatmap of the j-th target;

[0091] Step 403: Calculate the common interest features for each pair of targets; aggregate the common interest features of each pair of targets using the attention intersection, expressed as:

[0092]

[0093] Among them, h i,j C represents the common feature of interest for the i-th and j-th objectives. i,j,kl C represents the intersection of attention between the i-th and j-th targets. i,j The element in the k-th row and l-th column, F img,kl The element in the k-th row and l-th column represents the visual features of the image;

[0094] For each target, perform step 401; for each pair of targets, perform steps 402 and 403 to obtain the common features of interest for each pair of targets.

[0095] The process of inputting the feature representations and common interest features of each pair of targets into the relationship prediction layer for relationship prediction includes the following steps:

[0096] It is the feature representation of the i-th target. h is the feature representation of the j-th target. i,j This indicates the common features of interest for the i-th and j-th objectives. and h i,j After concatenation, the data is input into a fully connected layer to calculate the probability vector of the relationship between the i-th target and the j-th target. The expression is as follows:

[0097]

[0098] in, Let W represent the predicted probability vector of the relationship between the i-th target and the j-th target, where l represents the number of relationship categories. o and b o These are the learnable parameters of the fully connected layer, and softmax is the activation function used to normalize a vector into a probability vector.

[0099] The calculation of the relationship prediction loss function, the optimization of the visual relationship extraction model using the relationship prediction loss function, and the extraction of visual relationships using the optimized visual relationship extraction model include the following steps:

[0100] Step 601, construct the relationship label q between the i-th target and the j-th target. ij qij It is a one-hot vector of length l, which is the number of relation categories, q ijk It is a relation labeling; when the i-th target and the j-th target have a k-th type of relation, q ijk =1, otherwise q ijk =0;

[0101] Step 602: Calculate the relationship prediction loss function L between the i-th target and the j-th target. ij The expression is:

[0102]

[0103] Where, p ijk L represents the k-th element of the predicted probability vector of the relationship between the i-th target and the j-th target. ij The prediction loss function represents the relationship between the i-th objective and the j-th objective;

[0104] Step 603, calculate the relationship prediction loss function, the expression of which is:

[0105]

[0106] Where L is the relationship prediction loss function;

[0107] Step 604: Based on the relationship prediction loss function L, use the Adam optimizer to optimize the visual relationship extraction model until L converges, and obtain the final model parameters; use the optimized visual relationship extraction model to perform visual relationship extraction.

[0108] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A visual relation extraction method based on attention intersection perception, characterized in that, The method includes: Step 1, constructing a visual relationship extraction model, including a pre-trained visual model, an attention intersection perception module, and a relationship prediction layer; Step 2, using the pre-trained visual model to extract image features; Step 3, obtaining feature representations of all targets based on image features and target region annotations; Step 4, using the attention intersection perception module to calculate common attention features for each pair of targets; Step 5, inputting the feature representations and common attention features of each pair of targets into the relationship prediction layer to perform relationship prediction; Step 6, calculating the relationship prediction loss function, optimizing the visual relationship extraction model using the relationship prediction loss function, and performing visual relationship extraction using the optimized visual relationship extraction model. The aforementioned use of the attention intersection perception module to calculate common attention features for each pair of targets includes the following steps: Step 401, calculate the attention heatmap for each target; The feature representation of each target is: ,Will The visual features of the mapped image The attention heatmap is calculated using the following expression: in, Indicates the first Attention heatmap of each target These are learnable parameters, and softmax is the activation function used to normalize a vector into a probability distribution. Step 402, calculate the attention intersection of each pair of targets; multiply the attention of each pair of targets bitwise, the expression is: in, Indicates the first The first goal and the first The intersection of attention to each target Indicates the first Attention heatmap of each target Indicates the first Attention heatmap of each target; Step 403: Calculate the common interest feature for each pair of targets; use the attention intersection of each pair of targets to derive the common interest feature for each pair of targets, expressed as: in, Indicates the first The first goal and the first The common focus characteristics of each goal Indicates the first The first goal and the first Intersection of attention of each target No. Line 1 Column elements, The visual features of the image Line 1 The elements of the column; for each target, perform step 401, and for each pair of targets, perform steps 402 and 403 to obtain the common features of interest for each pair of targets; The method of extracting image features using a pre-trained visual model includes the following steps: The input image is represented as... The visual features of the input image are extracted using a pre-trained visual model, expressed as follows: in, Indicates the visual features of an image. and This indicates the width and height of the image. This represents the hidden layer dimension of ResNet. This indicates the downsampling factor of ResNet.

2. The visual relation extraction method based on attention intersection perception according to claim 1, characterized in that, The method of obtaining feature representations of all targets based on image features and target region annotations includes the following steps: Representing the position of the target as the coordinates of a bounding box based on the target region annotations. ,in Indicates the index of the target. and This represents the coordinates of the top-left corner of the rectangle. and This represents the width and height of the bounding box. Using the ROIAlign algorithm, the target region is labeled and mapped into the image feature space to obtain the visual feature map of the target, expressed as: in, Indicates the first Visual feature map of each target, This represents the hidden layer dimension of ResNet. Indicates the output size of ROIAlign; for the ... Adaptive pooling is performed on the visual feature map of the nth target to obtain the nth target. The feature representation of each target is expressed as: in, Indicates the first Feature representation of each target; Extract visual features for each target to obtain feature representation of all targets: in, The feature representation of all targets, Indicates the number of targets.

3. The visual relation extraction method based on attention intersection perception according to claim 2, characterized in that, The process of inputting the feature representations and common interest features of each pair of targets into the relationship prediction layer for relationship prediction includes the following steps: , and After concatenating the components, input them into the fully connected layer and calculate the... The first goal and the first The relational probability vector of the targets is expressed as: in, Indicates the first The first goal and the first A probability vector for predicting the relationship between targets. Indicates the number of relation categories. and These are the learnable parameters of the fully connected layer, and softmax is the activation function used to normalize a vector into a probability vector.

4. The visual relation extraction method based on attention intersection perception according to claim 3, characterized in that, The calculation of the relationship prediction loss function, the optimization of the visual relationship extraction model using the relationship prediction loss function, and the extraction of visual relationships using the optimized visual relationship extraction model include the following steps: Step 601, construct the first The first goal and the first Relationship tags of each target , It is a relation with a length equal to the number of relation categories. one-hot vector, It is a relation label, when the first The first goal and the first The first goal has a number When class relationships, ,otherwise ; Step 602, calculate the first... The first goal and the first The relationship prediction loss function of the target The expression is: in, The first one refers to the first one. The first goal and the first The first target's relation prediction probability vector One element, Indicates the first The first goal and the first The relationship prediction loss function for each objective; Step 603, calculate the relationship prediction loss function, the expression of which is: in, It is the relationship prediction loss function; Step 604: Predict the loss function based on the relationship. The Adam optimizer was used to optimize the visual relation extraction model until... Converge to obtain the final model parameters; use the optimized visual relation extraction model to extract visual relations.

Citation Information

Patent Citations

  • Visual relation detection method based on knowledge embedding

    CN116704202A