Character interaction detection method and device based on semantic perception
By adopting semantic perception-based methods in character interaction detection, image features are extracted and enhanced and semantic information is integrated, the problem of insufficient information extraction in character interaction detection is solved, and the accuracy of detection and recognition ability in complex scenarios is improved.
Patent Information
- Application Number
- CN202510246059.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art has the problem of insufficient extraction of environmental information and target information in character interaction detection, which leads to inaccurate prediction of interaction categories, and is difficult to accurately identify interaction behaviors in complex scenarios.
A character interaction detection method based on semantic perception is adopted to extract shallow, middle and deep features of the image through the feature extraction module, and feature enhancement module is used to enhance feature. At the same time, a semantic perception context module is introduced to integrate semantic information into the detection process, and visual and semantic features are fused through the GloVe model and the Transformer encoder to generate semantic perception context features.
It effectively improves the problem of low detection accuracy due to insufficient information extraction, and improves the accuracy of character interaction detection and recognition capabilities in complex scenarios.
Smart Images

Figure CN120182997A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human interaction detection, and in particular, to a method and device for human interaction detection based on semantic perception. Background Art
[0002] As an emerging direction in the field of computer vision, human interaction detection combines the object localization task of object detection and the behavior classification task of behavior recognition, that is, locates the behavior subjects (persons) and behavior objects (other persons or objects) in an image or video and classifies the behaviors exerted by the subjects on the objects, and finally obtains a triple <behavior subject, interaction category, behavior object>. At the same time, human interaction detection can cooperate with other tasks in computer vision to complete complex tasks, such as object detection, action retrieval, subtitle generation, etc., and has broad application prospects. With the gradual popularization of informatization and the booming development of the field of computer vision, human interaction detection has been widely applied in the fields of intelligent monitoring, intelligent security, etc., and has become a research hotspot.
[0003] With the continuous change of life scenarios, human interaction detection applications are used in more and more specific scenarios to help people save manpower and material resources. However, the continuous enrichment of scenarios also brings corresponding difficulties to the human interaction detection task. There are often some fine-grained targets and occluded targets in life scenarios, and the complex target types and backgrounds will affect the detection and judgment of the interaction detection network for the correct targets. In life scenarios, the interaction categories generated by people and objects are diverse, and it is difficult to accurately identify and detect them only relying on single visual information, which affects the detection accuracy.
[0004] In the prior art, the human interaction detection methods based on deep learning models can be divided into two categories: two-stage methods and single-stage methods. The two-stage methods split the human interaction detection task into two stages according to the characteristics of the task. In the first stage, the human body targets and object targets in the image are detected, and in the second stage, the interaction behaviors are predicted according to the detection results of the first stage. The advantage of the two-stage method is that it can decouple the object detection and interaction behavior classification tasks, so that the two stages can be optimized separately, and various information can also be fused for detection. However, the two-stage method has obvious disadvantages, the complexity of the model is relatively high, and the recognition speed is relatively slow. The single-stage method improves the disadvantages of the two-stage method. Different from the two-stage method that sequentially performs two subtasks, it converts the human interaction detection task into a parallel detection task and outputs a triple end-to-end based on the image, which significantly improves the recognition speed compared with the two-stage method. Most of the current single-stage models adopt the method of multi-task collaborative learning, and the features are shared among tasks. However, the features and optimization objectives concerned by object detection and relationship prediction may have large differences, and the shared features among tasks will cause interference with each other, thereby affecting the overall optimization effect and it is difficult to achieve the optimal performance.
[0005] With the continuous development of human society, the types of interactions between people and objects are increasing day by day, and the environment where person-object interactions occur is gradually becoming more complex. The complex and diverse scenarios and interaction types pose great challenges to the person-object interaction detection task. Therefore, at present, the person-object interaction detection task mainly faces the following two problems:
[0006] (1) The problem of insufficient extraction of environmental information and target information in images. In real-life scenarios, there are often some fine-grained targets and occluded targets, which are more difficult to detect compared to normal targets. In addition, the environment where people and objects are located is complex and changeable, and the interaction behaviors between people and objects are closely related to the surrounding environment. Therefore, how to fully extract and utilize the environmental information and fine-grained target information in images is an issue faced by current research.
[0007] (2) The problem of inaccurate prediction of interaction types due to insufficient information. In life scenarios, the interaction types between people and objects are diverse, and it is difficult to accurately identify them relying solely on visual information, which affects the detection accuracy. Therefore, how to incorporate more feature information into the detection process is another problem faced by current person-object interaction detection. Summary of the Invention
[0008] The present invention provides a person-object interaction detection method and device based on semantic perception to solve the technical problems existing in the above-mentioned prior art.
[0009] To achieve the above object, the present invention provides a person-object interaction detection method based on semantic perception, which includes:
[0010] S1: Extract features from the input image to obtain three sets of image features L, U, and H with different sizes. The image features L, U, and H correspond to the shallow features, middle features, and deep features of the image respectively;
[0011] S2: Adjust the sizes of the image features L and H to be the same as that of the image feature U to obtain the adjusted image features l, u, and h.
[0012] The image features l, u, and h are respectively evenly divided into 4 parts with equal sizes along the channel direction to obtain the evenly divided image features l i 、u i and h i , where i = 1, 2, 3, 4.
[0013] Calculate the weight value α using the activation function i , α i = sigmoid(u i ). The weight value α i is the weight for fusing shallow features and deep features.
[0014] Calculate the fused feature u' i , u' i = α i l i + (1 - α i ) h i ,
[0015] Perform channel - dimension concatenation on the fused feature u' i to obtain the enhanced feature F', u , F' u = [u'1, u'2, u'3, u'4],
[0016] Process the enhanced feature F' u to obtain the output feature where Conv(), B() and δ() represent convolution operation, batch normalization operation and rectified linear unit function respectively;
[0017] S3: Obtain the instances detected in the image from the output feature to get the class set and the bounding box set where N is the number of instances, c i is the class, b i is the bounding box,
[0018] Convert the class c i into text format respectively to get the text set of semantic context where t i is the text converted from the class c i ,
[0019] Use the text encoder of the GloVe model to represent the text set T of semantic context as word vectors, and use a fully - connected layer to convert the word vectors into dimensions matching the visual context features to obtain the semantic context feature W, W = MLP(TextEncoder(T)), where the function TextEncoder() represents the text encoder of the GloVe model, and the function MLP() represents the fully - connected layer for aligning semantic context features and visual context features.
[0020] Calculate the spatial feature P of the instance, P = MLP(B), B is the bounding box set,
[0021] Calculate the semantic context feature f of the region where the instance is located l , f l = MLP(Concat(W, P)), Concat is the string concatenation function, and the function MLP() represents the fully - connected layer,
[0022] According to the output feature Map to obtain visual context information Among them, N g is the number of sequence blocks after the image is flattened, and d h is the dimension size.
[0023] Express the semantic context feature f l as
[0024] According to the semantic context feature f l and the visual context information f v obtain the interaction feature o vl , o vl = SelfAttn(Concat(f v , f l ))), where SelfAttn() is the self-attention mechanism function.
[0025] Express o vl as Regarding the o among them v as the context feature;
[0026] S4: Perform self-attention on the output feature ,
[0027] Perform cross-attention mechanism with o v to obtain the output f out ,
[0028] According to f out perform interaction behavior prediction to obtain the confidence score S of the interaction action category.
[0029] Obtain the interaction category of the input image according to the confidence score S.
[0030] In an embodiment of the present invention, in step S1, feature extraction of the input image is performed through the feature extraction network DarkNet in the YOLO algorithm.
[0031] In an embodiment of the present invention, in step S1, the feature extraction network DarkNet includes a first convolutional module and a second convolutional module connected to each other. Among them, the convolutional kernel, stride, and padding of the first convolutional module are 3, 2, and 1 respectively. The second convolutional module includes 4 sub-convolutional modules connected in series. Each sub-convolutional module includes a convolutional sub-module and a c2f sub-module. The c2f sub-module includes n Bottleneck modules. When using the YOLOv8n model in the YOLO algorithm, d = 0.33. When using the YOLOv8m model in the YOLO algorithm, d = 0.67.
[0032] After uniformly preprocessing the input image into an image with a size of 640×640×3, it is input into the feature extraction network DarkNet for processing. After passing through the first convolution module, the input image becomes a feature map with a size of 320×320×64. Then, the feature map is input into the second convolution module. The convolution kernel, stride, and padding of the first convolution sub-module are 3, 2, and 1 respectively, the c2f parameter of the first c2f sub-module is 6d, the convolution kernel, stride, and padding of the second convolution sub-module are 3, 2, and 1 respectively, the c2f parameter of the second c2f sub-module is 6d, the convolution kernel, stride, and padding of the third convolution sub-module are 3, 2, and 1 respectively, the c2f parameter of the third c2f sub-module is 6d, the convolution kernel, stride, and padding of the fourth convolution sub-module are 3, 2, and 1 respectively, and the c2f parameter of the fourth c2f sub-module is 3d.
[0033] After being processed by the feature extraction network DarkNet, image features L, U, and H are obtained, with sizes of 80×80×256, 40×40×512, and 20×20×1024 respectively.
[0034] In an embodiment of the present invention, in step S2, the size of the image feature L is adjusted using the third convolution module. The convolution kernel, stride, and padding of the third convolution module are 3, 2, and 2 respectively. The image feature H is processed using the fourth convolution module and the bilinear interpolation module. The convolution kernel, stride, and padding of the fourth convolution module are 1, 1, and 0 respectively. The fourth convolution module is used to adjust the number of channels of the image feature H, and the bilinear interpolation module is used to adjust the height and width of the image feature adjusted by the fourth convolution module.
[0035] Convolution operations are performed to adjust its size. First, convolution operations are performed on the image feature U3 to adjust the number of channels, and then the height and width of the image feature U3 are adjusted through bilinear interpolation operations.
[0036] In an embodiment of the present invention, the combined attention mechanism function is integrated in the Transformer encoder.
[0037] In an embodiment of the present invention, step S4 is executed in the Transformer decoder.
[0038] In an embodiment of the present invention, the additive attention mechanism is executed in the Transformer encoder and the Transformer decoder. The additive attention mechanism is as follows:
[0039] Two learnable weight matrices W q 、W k are used to map the input vector to obtain the query Q and the key K.
[0040] Calculate the global attention query vector α. w a is a learnable parameter vector, d is 256,
[0041] Calculate the global query vector q, where n is the number of Qs obtained by mapping,
[0042] Calculate the output of the entire attention mechanism where, represents the normalized Q, and T1 and T2 represent linear transformations.
[0043] The present invention also provides a semantic perception-based human interaction detection device for performing the above method, which includes a feature extraction module, a feature enhancement module, a semantic perception context module, and an interaction behavior recognition module, where:
[0044] The feature extraction module is configured to execute step S1;
[0045] The feature enhancement module is configured to execute step S2;
[0046] The semantic perception context module is configured to execute step S3;
[0047] The interaction behavior recognition module is configured to execute step S4.
[0048] For the human interaction image, the semantic perception-based human interaction detection method provided by the present invention first uses the feature extraction module to extract features at different levels in the image, obtaining features at three levels: shallow, middle, and deep. Aiming at the problems of fine-grained targets and complex backgrounds, the present invention proposes a focus-diffusion feature enhancement module, which uses its special feature enhancement structure to adaptively enhance fine-grained features or context features according to the activation values of the activation function acting on the middle-level features. Aiming at the problem of inaccurate recognition of complex actions caused by insufficient context information in the image, the present invention proposes a semantic perception context module, which integrates semantic information into the detection process, improving the problem of blurred detection results caused by insufficient context information. The module constructs the semantic expression of the instance according to the instance information in the image, and then converts the semantic expression into semantic context features through the GloVe model. The Transformer encoder is used to fuse the semantic context features and visual context features, where the visual context features come from the output of the feature enhancement module, and finally the semantic perception-based visual context features are obtained, effectively improving the problem of insufficient context information. The features after feature enhancement are used as queries, and the semantic perception-based context features are used as key values and jointly input into the interaction behavior recognition module, and finally the confidence score S of the interaction action category is output through the self-attention and cross-attention mechanisms of the Transformer. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a schematic diagram of the complete implementation process of an embodiment of the present invention;
[0051] Figure 2 It is a schematic block diagram of the method for detecting human interaction based on semantic perception according to an embodiment of the present invention;
[0052] Figure 3 It is a schematic diagram of the DarkNet network according to an embodiment of the present invention;
[0053] Figure 4 It is a schematic diagram for explaining the feature enhancement module and step S2;
[0054] Figure 5 It is a schematic diagram of the semantic perception context module according to an embodiment of the present invention;
[0055] Figure 6 It is a working flowchart of the interaction behavior recognition module according to an embodiment of the present invention;
[0056] Figure 7a It is a schematic diagram of the process of multiplicative attention;
[0057] Figure 7b It is a schematic diagram of the process of additive attention. Detailed implementation manners
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0059] Figure 2 It is a schematic block diagram of the method for detecting human interaction based on semantic perception according to an embodiment of the present invention. As Figure 2 shown, the present invention provides a device for detecting human interaction based on semantic perception for executing the above method, which includes a feature extraction module, a feature enhancement module, a semantic perception context module, and an interaction behavior recognition module, wherein:
[0060] The feature extraction module is configured to perform the following step S1;
[0061] The feature enhancement module is configured to perform the following step S2;
[0062] The semantic-aware context module is configured to perform the following step S3;
[0063] The interaction behavior recognition module is configured to perform the following step S4.
[0064] The experimental platform of the present invention is as follows: a high-performance computer, a development platform with Python3 configured with Pytorch. The following provides a detailed description of each step and module of the present invention.
[0065] The method for detecting human interaction based on semantic awareness provided by the present invention, as Figure 1 shown in the complete implementation process schematic diagram of an embodiment of the present invention, includes:
[0066] S1: Extract features from the input image to obtain three sets of image features L, U, and H with different sizes. The image features L, U, and H correspond to the shallow features, middle-level features, and deep features of the image respectively;
[0067] Step S1 mainly solves the problem that the environmental information and target information in the image in the prior art are not extracted sufficiently. The present invention proposes a focus-diffusion feature enhancement network, which uses its unique structure to fully extract and enhance the fine-grained information and context information in the image, and improves the problem of low detection accuracy caused by insufficient information extraction.
[0068] Feature extraction refers to extracting useful information from an image, such as RGB, image structure, etc. information. Convolutional neural networks are usually used for feature extraction, and commonly used ones include ResNet, VGG, DenseNet, etc. Feature selection and dimensionality reduction help reduce the feature dimension and improve the algorithm efficiency. The image after feature extraction contains multi-scale feature information, which can better utilize the background information and small target information in the image compared with the existing single-scale features.
[0069] In this step, the feature extraction network DarkNet in the YOLO (You Only Look Once) algorithm is used for feature extraction work. This network is simple and efficient, with less computational complexity and fewer parameters, and has more advantages in the task of human interaction detection.
[0070] In this embodiment, feature extraction of the input image is performed through the feature extraction network DarkNet in the YOLO algorithm. In step S1, as Figure 3The following is a schematic diagram of the DarkNet network according to an embodiment of the present invention. The feature extraction network DarkNet includes a first convolutional module and a second convolutional module connected to each other. The convolution kernel (kernel size), stride, and padding of the first convolutional module are 3, 2, and 1 respectively. The second convolutional module includes 4 sub-convolutional modules connected in series. Each sub-convolutional module includes a convolutional sub-module and a c2f sub-module. The c2f sub-module includes n Bottleneck modules. n is the number of Bottlenecks in the figure. n depends on the size of the yolo model. When using the YOLOv8n model in the YOLO algorithm, d = 0.33. When using the YOLOv8m model in the YOLO algorithm, d = 0.67. This is because the d (depth) in the yolov8n model is 0.33, so n in c2f may be Or For another example, in the yolov8m model, d is 0.67, so n in the c2f module may be Or Where Is the ceiling symbol.
[0071] After uniformly preprocessing the input image into an image with a size of 640×640×3, it is input into the feature extraction network DarkNet for processing. After passing through the first convolutional module, the input image becomes a feature map with a size of 320×320×64. Then, the feature map is input into the second convolutional module. The convolution kernel, stride, and padding of the first convolutional sub-module are 3, 2, and 1 respectively. The c2f parameter of the first c2f sub-module is 6d. The convolution kernel, stride, and padding of the second convolutional sub-module are 3, 2, and 1 respectively. The c2f parameter of the second c2f sub-module is 6d. The convolution kernel, stride, and padding of the third convolutional sub-module are 3, 2, and 1 respectively. The c2f parameter of the third c2f sub-module is 6d. The convolution kernel, stride, and padding of the fourth convolutional sub-module are 3, 2, and 1 respectively. The c2f parameter of the fourth c2f sub-module is 3d.
[0072] After being processed by the feature extraction network DarkNet, image features L, U, and H are obtained, with sizes of 80×80×256, 40×40×512, and 20×20×1024 respectively.
[0073] The above is a detailed description of step S1 and the feature extraction module. The following is an explanation of step S2 and the feature enhancement module. The feature enhancement module enhances the features extracted by the feature extraction module to facilitate subsequent detection work, especially for small targets and background information that are difficult to be detected. The feature enhancement module can better extract the feature information therein. By Figure 2As shown in the figure, in the present invention, the input image is first subjected to feature extraction by the Darknet network, and the extracted features are enhanced by a feature enhancement module (focus-diffusion feature enhancement module) for fine-grained features and context features, and the obtained features are processed in three paths subsequently.
[0074] In the first path, object detection is performed based on the enhanced features, and semantic features are constructed according to the detected instance information and semantic context information is obtained; in the second path, context feature extraction is performed on the enhanced features to obtain visual context information, and the obtained semantic and visual context information are jointly input into the Transformer encoder to obtain semantic-aware context information. The semantic-aware context features are used as keys and values; in the third path, the enhanced features are used as queries and input into the Transformer decoder, and finally the confidence score of the interaction action is obtained.
[0075] As Figure 4 shown in the figure is the schematic diagram of the feature enhancement module and step S2, and step S2 is as follows:
[0076] S2: Adjust the sizes of the image features L and H to be the same as that of the image feature U to obtain the adjusted image features l, u, and h.
[0077] Due to the characteristics of the convolutional neural network, the receptive field of the network will become larger as the number of convolutional layers increases. Therefore, the shallow features contain more image detail information. Correspondingly, as the number of convolutional layers continues to increase, the receptive field expands continuously, and the deep features of the image extracted at this time will contain more image context information and global information.
[0078] Based on the above characteristics, the present invention designs a feature enhancement module. This module takes the three output features of the feature extraction module as inputs. The size of the shallow features is 80×80×256, the size of the middle-layer features is 40×40×512, and the size of the deep features is 20×20×1024. This module first unifies the sizes of these three groups of features with different sizes and unifies the three groups of features to the same dimension as the middle-layer features. For the shallow features, their sizes are adjusted through convolutional operations with k = 3, s = 2, and p = 2; for the deep features, the number of channels is first adjusted through convolutional operations with k = 1, s = 1, and p = 0, and then the height and width of the features are adjusted through bilinear interpolation operations.
[0079] The image features l, u, and h are respectively evenly divided into 4 parts with equal sizes along the channel direction to obtain the evenly divided image features l i 、u i and h i , where i = 1, 2, 3, 4.
[0080] As the characteristics of the convolutional neural network described above, the deep features may lose the information of small objects, while the shallow features may not provide sufficient background information. This feature enhancement module proposes an adaptive feature enhancement mechanism that applies an activation function to the values of the middle-level features to adaptively enhance the fine-grained features or context features. In the figure, α i is the value after the activation function sigmoid is applied to the middle-level features,
[0081] Calculate the weight value α using the activation function i , α i = sigmoid(u i ), and the weight value α i is the weight for fusing the shallow features and the deep features,
[0082] Determine the weight for fusing the shallow features and the deep features according to the obtained α i . If α i is greater than 0.5, it means that the middle-level features are more obvious in expressing the fine-grained features, so the fine-grained features account for a larger proportion in the fused features. Conversely, if α i is less than 0.5, it means that the middle-level features are more obvious in expressing the context features after passing through the activation function, so the context information accounts for a larger proportion in the fused features.
[0083] Calculate the fused feature u' i , u' i = α i l i + (1 - α i )h i ,
[0084] Concatenate the fused feature u' i in the channel dimension to obtain the enhanced feature F' u , F' u = [u'1, u'2, u'3, u'4], and F' u has the same size as the feature with unified size, which is convenient for subsequent feature processing.
[0085] Process the enhanced feature F' u to obtain the output feature where Conv(), B(), and δ() represent convolution operation, batch normalization operation, and rectified linear unit function respectively;
[0086] Among them, Conv() is a simple 1×1 convolutional module, mainly used to map the concatenated high-dimensional features back to the original dimension, enhance the feature expression ability, and keep the spatial size of the feature map unchanged. The normalization uses BatchNorm2d provided by Pytorch, which is applicable to the feature maps in the convolutional neural network, with the shape of [batch_size, channels, height, width]. After performing the convolutional operation on the fused features, the feature map is normalized by BatchNorm2d. The rectified linear unit (ReLU) is used as the activation function to perform non-linear transformation on the normalized features, and the formula is expressed as ReLU(x) = max(0, x).
[0087] The following is an explanation of step S3 and the semantic-aware context module.
[0088] Mining the context information in the image is crucial for accurately realizing human interaction detection. Some existing methods have extracted and utilized the context information in the image, but there are still some limitations. That is, only the visual features of the image are used as the context information, which will lead to the problem of insufficient context information, and then lead to the problem of low detection accuracy when facing complex interaction action categories. For example, in the interaction pair of "person playing tennis", in addition to the correct detection of "person" and "tennis ball", the features of "tennis racket" also play an important auxiliary role in understanding the whole scene. However, in real scenarios, practical problems such as occlusion of objects are inevitable, which will lead to the problem of blurred context information. At this time, the information of "there is a tennis racket" in the image will improve this ambiguity. Therefore, in order to solve the problem of insufficient context information, the present invention proposes a semantic-aware context module, which introduces semantic context information to visually perceive the semantic context information in the image and obtain the semantic-aware context features. The structure of this module is as Figure 5 shown. This module first constructs a context text description for the categories detected in the image, and then generates word vectors through the GloVe (Global Vectors for Word Representation) text encoder to represent the semantic context features of the region where the instance is located. Then, the image context features enhanced by the feature enhancement module and the encoded semantic context features are jointly input into the Transformer encoder to perform the attention operation, and the semantic-aware context features are obtained.
[0089] S3: Obtain the instances detected in the image from the output features to obtain the category set and the bounding box set where N is the number of instances, c i is the category, b i is the bounding box,
[0090] Convert the category c i into text format respectively to obtain a text set of semantic context where t i is the text converted from the category c i and
[0091] In this embodiment, the detected category (Object) is converted into a text description in the format of "A photo of a / an [Object]". In other embodiments, it can also be converted into other forms depending on actual needs.
[0092] Use the text encoder of the GloVe model to represent the text set T of semantic context as word vectors, and use a fully connected layer to convert the word vectors into dimensions matching the visual context features to obtain semantic context features W, W = MLP(TextEncoder(T)), where the function TextEncoder() represents the text encoder of the GloVe model for converting text descriptions into word vectors, and the function MLP() represents the fully connected layer for aligning semantic context features and visual context features.
[0093] In addition, to supplement spatial information, the present invention enhances the representation of semantic context features by splicing spatial features. For each instance, the language context features and spatial features corresponding to the instance are spliced to obtain f l which is used to represent the semantic context features of the region where the instance is located.
[0094] Calculate the spatial features P of the instance, P = MLP(B), where B is a set of bounding boxes
[0095] Calculate the semantic context features f l of the region where the instance is located, f l = MLP(Concat(W, P)), Concat is a string splicing function, and the function MLP() represents the fully connected layer
[0096] After obtaining the semantic context features, it is necessary to use a Transformer encoder to fuse semantic information and visual information. The present invention uses a merged attention mechanism to fuse visual information and semantic information. In the merged attention, text features and visual features are simply spliced together and then input into a single encoder to perform self-attention, thereby realizing feature fusion.
[0097] Map according to the output features to obtain visual context information where N g is the number of sequence blocks after flattening the picture, and d h is the dimension size
[0098] Represent the semantic context feature f l as
[0099] According to the semantic context feature f l and the visual context information f v obtain the interaction feature o vl , o vl = SelfAttn(Concat(f v , f l ))), where SelfAttn() is the self-attention mechanism function,
[0100] Represent o vl as Regarding the o among them v as the context feature. Here, o v is a part of o vl . Since the input and output of the Transformer self-attention have the same size, the input here is the semantic context feature f l concatenated with the visual context information f v . When outputting o vl , the corresponding-sized o v can be obtained by taking it out.
[0101] In this step, the semantic feature and the visual feature are used to construct the input feature by concatenation (Concat), and then the self-attention is performed so that the visual context feature can perceive the semantic context information of the picture, obtaining the output o vl The visual context feature o perceived by language v and the language context feature o perceived by vision l are combined. Encoding the semantic and context features can improve the problem of insufficient information basis caused by the existing single modality.
[0102] In the above text, N g is the number of sequence blocks after the picture is flattened, and d h is the dimension size. The input of the Vision Transformer is not the entire image, but the image cut into small pieces. Assuming the input image size is 224×224, the image is cut into fixed-size 16×16 squares, and each small square is a patch. Then the number of patches in each image is (224×224) / (16×16) = 196. After cutting, 196 patches of [16, 16, 3] are obtained, and the dimension of each token after flattening is 16×16×3 = 768, so the dimension is [196, 768]. The dimension at this time is represented as
[0103] The following describes step S4 and the interaction behavior recognition module. The interaction behavior recognition module is constructed based on the decoder structure of Transformer, and queries context information for the image information through the cross-attention mechanism of Transformer.
[0104] As Figure 6 shown, the input of the interaction behavior recognition module consists of two parts. The first part is the feature enhanced by the feature enhancement module This part serves as the query of the decoder; the second part is the semantically aware image context information, serving as the key and value of the decoder. In the interaction behavior recognition module, As the query of the decoder, self-attention is first performed, and then the cross-attention mechanism is performed with the semantically aware image context to obtain the output f out .
[0105] S4: Perform self-attention on the output feature and
[0106] perform the cross-attention mechanism with o v to obtain the output f out .
[0107] Based on f out perform interaction behavior prediction to obtain the confidence score S of the interaction action category,
[0108] and obtain the interaction category of the input image according to the confidence score S.
[0109] Finally, the output f out of the decoder is used in the input interaction detection head to perform the prediction of the interaction behavior, obtain the confidence score S of the interaction action category, and predict the interaction category according to the confidence score S output by the detection head. Figure 6The five blocks in it represent the inputs of the interaction behavior recognition module. Since this module is built based on the decoder structure of Transformer, the input of the Transformer decoder is in the form of small chunks (tokens), so it is represented in the figure as blocks. Add&norm consists of two parts. The first part is Add, which is essentially a residual connection. The input of this layer is added to the input of the previous layer to prevent the problem of degradation during the training of the deep neural network. The second part is Normalization. The normalization method used is Layer Normalization, which can accelerate the training speed and improve the training stability. The whole process can be expressed as LayerNorm(X + FeedFoward(X)), where X represents the input of Feed Forward, and FeedFoward(X) represents the output (the output has the same dimension as the input X, so they can be added). The parameters K and V represent Key and Value in Transformer, and here they represent the output o of the semantic perception context module. v .
[0110] In an embodiment of the present invention, in step S2, the image feature L is adjusted in size by the third convolution module. The convolution kernel, stride, and padding of the third convolution module are 3, 2, and 2 respectively. The image feature H is processed by the fourth convolution module and the bilinear interpolation module. The convolution kernel, stride, and padding of the fourth convolution module are 1, 1, and 0 respectively. The fourth convolution module is used to adjust the number of channels of the image feature H, and the bilinear interpolation module is used to adjust the height and width of the image feature adjusted by the fourth convolution module.
[0111] Perform a convolution operation to adjust its size. First, perform a convolution operation on the image feature U3 to adjust the number of channels, and then adjust the height and width of the image feature U3 through a bilinear interpolation operation.
[0112] In an embodiment of the present invention, the combined attention mechanism function is integrated in the Transformer encoder.
[0113] In an embodiment of the present invention, step S4 is executed in the Transformer decoder.
[0114] The present invention uses the encoder and decoder of Transformer in two modules. Due to the unique multiplicative multi-head attention mechanism of the Transformer encoder and decoder, the computational complexity of the whole model is relatively large. Therefore, the present invention proposes an efficient additive attention mechanism. The efficient additive attention solves the problem that the computational amount and memory occupation of the multi-head self-attention increase quadratically with the increase of the input length, and reduces the complexity without affecting the performance. As Figure 7bAs shown, first, the input attention query matrix is summarized into a global query vector, and then the interaction between the attention key and the global query vector is modeled through element-wise product to learn the globally context-aware key matrix, which is then summarized into a global key vector through additive attention. Next, element-wise product is used to aggregate the global key and the attention values, and then they are processed through a linear transformation to calculate the globally context-aware attention values. Finally, the original attention query and the globally context-aware attention values are added together to form the final output.
[0115] As Figure 7b shown, the additive attention mechanism is executed in the Transformer encoder and the Transformer decoder, and the additive attention mechanism is as follows:
[0116] Two learnable weight matrices W q and W k are used to map the input vectors to obtain the query Q and the key K.
[0117] Calculate the global attention query vector α. w a is a learnable parameter vector, and d is 256.
[0118] Calculate the global query vector q. where n is the number of Qs obtained by mapping.
[0119] Calculate the output of the entire attention mechanism. where represents the normalized Q, and T1 and T2 represent linear transformations.
[0120] The person interaction detection method based on semantic perception provided by the present invention, for person interaction images, first uses a feature extraction module to extract features at different levels in the image, obtaining features at three levels: shallow, middle, and deep. Aiming at the problems of fine-grained objects and complex backgrounds, the present invention proposes a focus-diffusion feature enhancement module, which uses its special feature enhancement structure to adaptively enhance fine-grained features or context features according to the activation values of the activation function acting on the middle-level features. Aiming at the problem of inaccurate recognition of complex actions caused by insufficient context information in the image, the present invention proposes a semantic perception context module, which integrates semantic information into the detection process and improves the problem of blurred detection results caused by insufficient context information. This module constructs a semantic expression of the instance according to the instance information in the image, and then converts the semantic expression into a semantic context feature through the GloVe model. The Transformer encoder is used to fuse the semantic context feature and the visual context feature, where the visual context feature comes from the output of the feature enhancement module, and finally the semantic perception visual context feature is obtained, effectively improving the problem of insufficient context information. The features after feature enhancement are used as queries, and the semantic perception context features are used as key values and jointly input into the interaction behavior recognition module. Finally, the confidence score S of the interaction action category is output through the self-attention and cross-attention mechanisms of the Transformer.
[0121] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0122] Those of ordinary skill in the art can understand that the modules in the device in the embodiment can be distributed in the device in the embodiment according to the description of the embodiment, or can be correspondingly changed and located in one or more devices different from this embodiment. The modules in the above embodiments can be combined into one module, or can be further split into multiple sub-modules.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting human interaction based on semantic perception, characterized in that: include: S1: Extract features from the input image to obtain three sets of image features L, U, and H of different sizes. The image features L, U, and H correspond to the shallow features, middle features, and deep features of the image, respectively. S2: Adjust the size of image features L and H to be the same as image feature U, and obtain the adjusted image features l, u and h. The image features l, u and h are divided into four parts of equal size along the channel direction to obtain the averaged image feature l i 、u i and h i , where i = 1, 2, 3, 4, Use the activation function to calculate the weight value α i , α i =sigmoid(u i ), weight value α i To fuse the weights of shallow features and deep features, Calculate the fused feature u i ′,u i ′=α i l i +(1-α i )h i , The fused feature u i ′ performs channel dimension splicing to obtain the enhanced feature F u ′,F u ′=[u1′,u2′,u3′,u4′], For the enhanced feature F u ′ is processed to obtain the output features Among them, Conv(), B() and δ() represent convolution operation, batch normalization operation and linear rectification function respectively; S3: From the output features Get the instances detected in the image and get the category set and bounding box collection Where N is the number of instances, c i For category, b i is the bounding box, Category c i Convert them into text format respectively to get the text collection of semantic context where t i For category c i The converted text, The text encoder of the GloVe model is used to represent the text set T of the semantic context as a word vector, and the fully connected layer is used to convert the word vector into a dimension that matches the visual context feature to obtain the semantic context feature W, W = MLP (TextEncoder (T)), where the function TextEncoder () represents the text encoder of the GloVe model, and the function MLP () represents the fully connected layer, which is used to align the semantic context features and the visual context features. Calculate the spatial feature P of the instance, P = MLP (B), B is the bounding box set, Calculate the semantic context feature f of the region where the instance is located l , f l =MLP(Concat(W,P)), Concat is a string concatenation function, and the function MLP() represents a fully connected layer. According to the output characteristics Mapping to obtain visual context information Among them, N g is the number of sequence blocks after the image is flattened, d h is the dimension size, The semantic context feature f l Expressed as According to the semantic context feature f l With visual context information f v Get the interaction feature o vl , o vl =SelfAttn(Concat(f v ,f l )), where SelfAttn() is the self-attention mechanism function, o vl Expressed as The o v as a contextual feature; S4: Output features Perform self-attention, With o v Execute the cross attention mechanism and get the output f out , According to f out Perform interactive behavior prediction and obtain the confidence score S of the interactive action category. The interaction category of the input image is obtained according to the confidence score S.
2. The method for detecting human interaction based on semantic perception according to claim 1, characterized in that: In step S1, feature extraction of the input image is performed by the feature extraction network DarkNet in the YOLO algorithm.
3. The method for detecting human interaction based on semantic perception according to claim 2, characterized in that: In step S1, the feature extraction network DarkNet includes a first convolution module and a second convolution module connected to each other, wherein the convolution kernel, step size and padding of the first convolution module are 3, 2 and 1 respectively, the second convolution module includes 4 sub-convolution modules connected in series, each sub-convolution module includes a convolution sub-module and a c2f sub-module, the c2f sub-module includes n Bottleneck modules, when the YOLOv8n model in the YOLO algorithm is used, d=0.33, when the YOLOv8m model in the YOLO algorithm is used, d=0.67, The input image is uniformly preprocessed into an image of size 640×640×3 and then input into the feature extraction network DarkNet for processing. After passing through the first convolution module, the input image becomes a feature map of size 320×320×64, and then the feature map is input into the second convolution module. The convolution kernel, step size and padding of the first convolution submodule are 3, 2 and 1 respectively, and the c2f parameter of the first c2f submodule is 6d. The convolution kernel, step size and padding of the second convolution submodule are 3, 2 and 1 respectively, and the c2f parameter of the second c2f submodule is 6d. The convolution kernel, step size and padding of the third convolution submodule are 3, 2 and 1 respectively, and the c2f parameter of the third c2f submodule is 6d. The convolution kernel, step size and padding of the fourth convolution submodule are 3, 2 and 1 respectively, and the c2f parameter of the fourth c2f submodule is 3d. After being processed by the feature extraction network DarkNet, the image features L, U, and H are obtained, with sizes of 80×80×256, 40×40×512, and 20×20×1024, respectively.
4. The method for detecting human interaction based on semantic perception according to claim 1, characterized in that: In step S2, the size of the image feature L is adjusted using the third convolution module, the convolution kernel, step size and padding of the third convolution module are 3, 2 and 2 respectively, and the image feature H is processed using the fourth convolution module and the bilinear interpolation module, the convolution kernel, step size and padding of the fourth convolution module are 1, 1 and 0 respectively, the fourth convolution module is used to adjust the number of channels of the image feature H, and the bilinear interpolation module is used to adjust the height and width of the image feature adjusted by the fourth convolution module. A convolution operation is performed to adjust its size. A convolution operation is first performed on the image feature U3 to adjust the number of channels, and then the height and width of the image feature U3 are adjusted through a bilinear interpolation operation.
5. The method for detecting human interaction based on semantic perception according to claim 1, characterized in that: The merged attention mechanism function is integrated into the Transformer encoder.
6. The method for detecting human interaction based on semantic perception according to claim 1, characterized in that: Step S4 is executed in the Transformer decoder.
7. The method for detecting human interaction based on semantic perception according to claim 5 or 6, characterized in that: The additive attention mechanism is implemented in the Transformer encoder and Transformer decoder. The additive attention mechanism is as follows: Using two learnable weight matrices W q , W k Map the input vector to obtain query Q and key K, Calculate the global attention query vector α, w a is the learnable parameter vector, d is 256, Calculate the global query vector q, Where n is the number of Q obtained by mapping, Calculate the output of the entire attention mechanism in, represents the normalized Q, and T1 and T2 represent linear transformations.
8. A person interaction detection device based on semantic perception, used to execute the method according to any one of claims 1 to 7, characterized in that: It includes feature extraction module, feature enhancement module, semantic perception context module and interactive behavior recognition module, among which: The feature extraction module is configured to perform step S1; The feature enhancement module is configured to perform step S2; The semantic-aware context module is configured to perform step S3; The interactive behavior recognition module is configured to execute step S4.