Visual intention analysis method and system based on cross-modal pyramid alignment

By combining a cross-modal pyramid alignment method with hierarchical relationship mining and the BERT model, the problem of insufficient multi-level information modeling in visual intent understanding is solved, and more efficient visual intent classification is achieved.

CN116434255BActive Publication Date: 2025-12-23WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310277551.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-17
Publication Date
2025-12-23
Estimated Expiration
2043-03-17

AI Technical Summary

Technical Problem

Existing visual intent understanding methods cannot effectively model the relationships between multi-level information globally, and directly applying image classification methods to visual intent understanding is inappropriate as they cannot capture the intent of the image. Insufficient connection between textual and visual information leads to discrepancies in intent understanding.

Method used

We employ a visual intent analysis method based on cross-modal pyramid alignment. By combining a hierarchical relationship mining network and a cross-modal pyramid alignment network with the BERT model, we extract visual and textual features and perform alignment at each level. We then use metric learning to establish the internal connections between textual and visual features and dynamically optimize visual intent understanding.

Benefits of technology

It improves the global understanding of visual intent, reduces modal discrepancies, and enhances the classification performance of visual intent. Experimental results show that it outperforms existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116434255B_ABST
    Figure CN116434255B_ABST
Patent Text Reader

Abstract

The application discloses a visual intention analysis method and system based on cross-modal pyramid alignment, and aims at the problem that simply modeling objects or backgrounds in image content can lead to intention understanding divergence. The application proposes a hierarchical relationship mining method, uses the hierarchical relationship between visual content and text intention labels, and improves global understanding of visual intention through hierarchical modeling. For the visual hierarchical structure, the application converts visual intention understanding into a hierarchical classification problem, captures multi-granularity features in different layers, and the features correspond to hierarchical intention labels. For the text hierarchical structure, the application directly extracts semantic representations from intention labels of different layers, supplements visual content modeling, and does not need additional manual labeling. Meanwhile, the application designs a cross-modal pyramid alignment module, further reduces the domain gap between the two modalities, and dynamically optimizes the visual intention understanding performance in a joint learning manner.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computer vision analysis, and relates to a visual intention analysis method and system, in particular to a visual intention analysis method and system based on cross-modal pyramid alignment. BACKGROUND

[0002] Intention classification (literatures 1-3) has been extensively studied in the field of natural language processing (NLP), which divides user-generated short texts into intention categories, reflecting human emotions or motivations. It plays a crucial role in intelligent response (literatures 4-5), intelligent dialogue (literatures 6-8), and other applications. In addition, some visual intention-related tasks are also applied to the real world (literatures 9-15). With the development of technology, visual content has become the main form of information transmission in social media. Compared with text intention classification, visual content cannot directly express personal opinions. Obviously, the classification method for text intention cannot solve this problem. Therefore, the present application studies the visual intention understanding problem, that is, analyzing the intention from the visual information captured by social media images.

[0003] Visual intention understanding is a subjective task involving visual information, which is inherent to human cognition and behavior. It involves the internal summary analysis of image content, rather than simply focusing on objects or backgrounds. For example, given a picture content of a girl sitting on the ground facing the sea with a tea beside her, the output of visual intention understanding is "enjoying life". In order to achieve this goal, the bright sea surface, warm sunlight, and other features in the image need to be globally associated, which requires global content modeling in visual images. That is, from low-level information (texture, color) to high-level semantic information (sitting posture) and the relationship between them, all information is essential for understanding visual intention. In contrast, the traditional image classification method (literatures 18-20) widely studied at present is a one-to-one correspondence identification task, which classifies specific objects in a complex scene as "tea" or "table". Therefore, it is not appropriate to directly apply image classification methods to visual intention understanding. In addition, the visual intention understanding task also includes common challenges (view difference, posture change, color change, etc.) in target classification (literatures 21-24). Therefore, how to globally model the relationship between multi-level information without being affected by irrelevant factors in visual images is the key to visual intention understanding.

[0004] Visual intent understanding is an emerging research topic, and a few works (16, 25-28) have been proposed. Most of these methods solve the problem from the perspective of emotion recognition. Among them, a similar work (16) considers the visual intent classification of both objects and backgrounds. Due to the richness and diversity of objects and environments in visual images, merely extracting the foreground or background is insufficient to capture the intent of the image. Therefore, work (16) introduces text information to help understand the intent. However, their text information focuses on describing objects (coffee, smartphone), actions (handheld, smile), and colors (writing, black) to supplement the diversity of specific intent categories. As the number of intent classes increases, collecting text descriptions of relative objects, actions, and colors requires a large amount of manpower. In addition, there is often some overlap in vocabulary between different intent categories, which can lead to ambiguity in intent understanding. At the same time, this method directly connects the text information with the visual information without interaction. It is difficult to update the text features adaptively during training to facilitate the understanding of visual intent. Therefore, how to obtain effective text information and establish the relationship between the two modalities is the key to visual intent understanding.

[0005] In summary, it is crucial to design a dynamic optimization-based cross-modal pyramid alignment algorithm for visual intent understanding to solve the above problems.

[0006] [1] Y. Ouyang, J. Ye, Y. Chen, X. Dai, S. Huang, and J. Chen, “Energy-based unknown intent detection with data manipulation,” in ACL-IJCNLP, 2021.

[0007] [2] Y. Hou, Y. Lai, C. Chen, W. Che, and T. Liu, “Learning to bridge metric spaces: Few-shot joint learning of intent detection and slot filling,” in ACL-IJCNLP, 2021.

[0008] [3] T. Dopierre, C. Gravier, and W. Logerais, “Protaugment: Unsupervised diverse short-texts paraphrasing for intent detection meta-learning,” in ACL-IJCNLP, 2021.

[0009] [Document 4] A. Kannan, K. Kurach, S. Ravi, T. Kaufmann, A. Tomkins, B. Miklos, G. Corrado, L. Lukacs, M. Ganea, P. Young et al., “Smart reply: Automated response suggestion for email,” in ACM SIGKDD, 2016.

[0010] [Document 5] K. C. Arnold, K. Chauncey, and K. Z. Gajos, “Predictive text encourages predictable writing,” in ACM IUI, 2020.

[0011] [Document 6] J. Worsham and J. Kalita, “Multi-task learning for natural language processing in the 2020s: where are we going?” Pattern Recognition Letters, 2020.

[0012] [Document 7] L. Qin, W. Che, Y. Li, H. Wen, and T. Liu, “A stack-propagation framework with token-level intent detection for spoken language understanding,” arXiv preprint arXiv:1909.02188, 2019.

[0013] [Document 8] C.-W. Goo, G. Gao, Y.-K. Hsu, C.-L. Huo, T.-C. Chen, K.-W. Hsu, and Y.-N. Chen, “Slot-gated modeling for joint slot filling and intent prediction,” in NAACL-HLT, 2018.

[0014] [Document 9] X. Tian, D. Tao, X.-S. Hua, and X. Wu, “Active reranking for web image search,” IEEE TIP, 2009.

[0015] [Ref. 10] S. Filipe, L. Itti, and L. A. Alexandre, “Bik-bus: biologically motivated 3d keypoint based on bottom-up saliency,” IEEE TIP, 2015.

[0016] [Ref. 11] L. Nie, F. Jiao, W. Wang, Y. Wang, and Q. Tian, “Conversational image search,” IEEE TIP, 2021.

[0017] [Ref. 12] X. Tang, K. Liu, J. Cui, F. Wen, and X. Wang, “Intent search: capturing user intention for one-click internet image search,” IEEE TPAMI, 2011.

[0018] [Ref. 13] Z. Ji, Y. Pang, and X. Li, “Relevance preserving projection and ranking for web image search reranking,” IEEE TIP, 2015.

[0019] [Ref. 14] S. Xiao, Z. Zhao, Z. Zhang, Z. Guan, and D. Cai, “Query-biased self attentive network for query-focused video summarization,” IEEE TIP, 2020.

[0020] [Ref. 15] X. Dong and M. J. Chantler, “Perceptually motivated image features using contours,”

[0021] IEEE TIP, 2016.

[0022] [Reference 16] M. Jia, Z. Wu, A. Reiter, C. Cardie, S. Belongie, and S.-N. Lim, “Intentonomy: a dataset and study towards human intent understanding,” in CVPR, 2021.

[0023] [Reference 17] J. R. Talevich, S. J. Read, D. A. Walsh, R. Iyer, and G. Chopra, “Toward a comprehensive taxonomy of human motives,” PloS one, 2017.

[0024] [Reference 18] M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in ICML, 2019.

[0025] [Reference 19] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in CVPR, 2017.

[0026] [Reference 20] C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi, “Inception-v4, inception-resnet and the impact of residual connections on learning,” in AAAI, 2017.

[0027] [Reference 21] F. Schroff, T. Treibitz, D. Kriegman, and S. Belongie, “Pose, illumination and expression invariant pairwise face-similarity measure via doppleganger list comparison,” in ·· ICCV, 2011.

[0028] [Ref. 22] R. Gross, I. Matthews, J. Cohn, T. Kanade, and S. Baker, “Multi-pie,” IVC, 2010.

[0029] [Ref. 23] M. Ye, C. Chen, J. Shen, and L. Shao, “Dynamic tri-level relation mining with attentive graph for visible infrared re-identification,” IEEE TIFS, 2021.

[0030] [Ref. 24] M. Ye, H. Li, B. Du, J. Shen, L. Shao, and S. C. Hoi, “Collaborative refining for person re-identification with label noise,” IEEE TIP, 2021.

[0031] [Ref. 25] X. Huang and A. Kovashka, “Inferring visual persuasion via body language, setting, and deep features,” in CVPR-Workshop, 2016.

[0032] [Ref. 26] J. Joo, F. F. Steen, and S.-C. Zhu, “Automated facial trait judgment and election outcome prediction: Social dimensions of face,” in ICCV, 2015.

[0033] [Ref. 27] B. Siddiquie, D. Chisholm, and A. Divakaran, “Exploiting multi-modal affect and semantics to identify politically persuasive web videos,” in ICMI, 2015.

[0034] [Document 28] C. Thomas and A. Kovashka, "Predicting the politics of an image using webly supervised data," in NeurIPS, 2019. SUMMARY

[0035] In view of the deficiencies of the prior art, the present application provides a visual intention analysis method and system based on cross-modal pyramid alignment.

[0036] The technical scheme adopted by the method of the present application is: a visual intention analysis method based on cross-modal pyramid alignment, comprising the following steps:

[0037] Step 1: constructing a hierarchical relationship mining network, a cross-modal pyramid alignment network and a BERT model;

[0038] The hierarchical relationship mining network is used to mine visual hierarchical information and comprises a residual network, a down-sampling layer, a pooling layer and a channel attention module.

[0039] The residual network comprises sequentially connected convolution layers, a max-pooling layer, a first module, a second module, a third module and a fourth module; the convolution kernel size of the convolution layers and the max-pooling layer is 7 and the step size is 2; the first module comprises 3 blocks, each block is composed of three convolution layers with convolution kernel sizes of 1x1, 3x3 and 1x1 and a step size of 1; the second module comprises 4 blocks, each block is composed of three convolution layers with convolution kernel sizes of 1x1, 3x3 and 1x1 and a step size of 2; the third module comprises 6 blocks, each block is composed of three convolution layers with convolution kernel sizes of 1x1, 3x3 and 1x1 and a step size of 2; the fourth module comprises 3 blocks, each block is composed of three convolution layers with convolution kernel sizes of 1x1, 3x3 and 1x1 and a step size of 2; a batch normalization layer and an activation layer are arranged after each convolution of the first module, the second module, the third module and the fourth module; and each block is connected through a residual connection.

[0040] The down-sampling layer comprises a convolution layer with a convolution kernel of 3x3 and a step size of 2 and a batch normalization layer.

[0041] The output of the first module is connected to the output of the second module after passing through the down-sampling layer, then connected to the output of the third module after passing through the down-sampling layer, and then connected to the output of the fourth module after passing through the down-sampling layer, and then input into the channel attention module.

[0042] The channel attention module is a convolution layer with a convolution kernel of 1x1 and a step size of 1.

[0043] The BERT model is used for extracting text features of coarse granularity, medium granularity and fine granularity.

[0044] The cross-modal pyramid alignment network is used for aligning the visual hierarchical information extracted by the hierarchical relationship mining network and the text hierarchical features extracted by the BERT model at each level, and comprises a PIENet and full connection layers corresponding to each level; the PIENet is composed of a multi-head attention module and a residual connection, and the multi-head attention module is composed of a full connection layer, an activation layer, a full connection layer and a softmax layer.

[0045] Step 2: input the visual picture to be processed into the hierarchical relationship mining network, extract the visual information of different levels of the picture, and aggregate the features extracted from different levels into aggregated features; the features output by the first module are connected to the features output by the second module through a down-sampling layer to obtain coarse-grained aggregated features; the features output by the second module are connected to the features output by the third module through a down-sampling layer to obtain medium-grained aggregated features; the features output by the third module are connected to the features output by the fourth module through a down-sampling layer to obtain fine-grained aggregated features; the coarse-grained and medium-grained features are subjected to a pooling layer and a linear layer to obtain corresponding classification losses, and the fine-grained features are subjected to a channel attention module, a pooling layer and a linear layer to obtain a corresponding fine-grained classification loss;

[0046] Step 3: extract text semantic information of different levels from the intent labels of different levels;

[0047] Step 4: input the three visual aggregated features of different levels obtained and the text semantic information of different levels extracted into the cross-modal pyramid alignment network;

[0048] Step 5: use metric learning to establish the internal relationship between the text features f t h and the visual features at the same level, and adopt layer normalization to obtain enhanced text features and visual features The semantic relationship between the enhanced text and visual features is measured by metric learning, and a corresponding modal loss is obtained to assist the final intent judgment.

[0049] The technical scheme adopted by the system of the application is: a visual intent analysis system based on cross-modal pyramid alignment, comprising the following modules:

[0050] Module 1 is used for constructing a hierarchical relationship mining network, a cross-modal pyramid alignment network and a BERT model.

[0051] The hierarchical relationship mining network is used for mining visual hierarchical information; and comprises a residual network, a down-sampling layer, a pooling layer and a channel attention module.

[0052] The residual network comprises sequentially connected convolution layers, a max-pooling layer, a first module, a second module, a third module and a fourth module; the convolution layers and the max-pooling layer each have a convolution kernel size of 7 and a step size of 2; the first module comprises three blocks, each block comprising three convolution layers with a kernel size of 1*1, 3*3 and 1*1 and a step size of 1; the second module comprises four blocks, each block comprising three convolution layers with a kernel size of 1*1, 3*3 and 1*1 and a step size of 2; the third module comprises six blocks, each block comprising three convolution layers with a kernel size of 1*1, 3*3 and 1*1 and a step size of 2; and the fourth module comprises three blocks, each block comprising three convolution layers with a kernel size of 1*1, 3*3 and 1*1 and a step size of 2; each convolution of the first module, the second module, the third module and the fourth module is followed by a batch normalization layer and an activation layer; and each block is connected through a residual connection.

[0053] The down-sampling layer comprises a convolution layer with a kernel size of 3*3 and a step size of 2 and a batch normalization layer.

[0054] The output of the first module is connected to the output of the second module after being input into the down-sampling layer, then connected to the output of the third module after being input into the down-sampling layer, and then connected to the output of the fourth module after being input into the down-sampling layer, and finally input into the channel attention module.

[0055] The channel attention module is a convolution layer with a kernel size of 1*1 and a step size of 1.

[0056] The BERT model is used for extracting text features of coarse granularity, medium granularity and fine granularity.

[0057] The cross-modal pyramid alignment network is used for aligning the visual hierarchical information extracted by the hierarchical relationship mining network and the text hierarchical features extracted by the BERT model at each level; and comprises a PIENet and full connection layers corresponding to each level; the PIENet is composed of a multi-head attention module and a residual connection; and the multi-head attention module comprises a full connection layer, an activation layer, a full connection layer and a softmax layer.

[0058] Module 2 is used for inputting the visual picture to be processed into the hierarchical relationship mining network, extracting different levels of visual information of the picture, and aggregating the features extracted from different levels into aggregated features; the features output by the first module are connected to the features output by the second module through a down-sampling layer to obtain coarse-grained aggregated features; the features output by the second module are connected to the features output by the third module through a down-sampling layer to obtain medium-grained aggregated features; the features output by the third module are connected to the features output by the fourth module through a down-sampling layer to obtain fine-grained aggregated features; the coarse-grained and medium-grained features pass through a pooling layer and a linear layer to obtain corresponding classification losses, and the fine-grained features pass through a channel attention module, a pooling layer and a linear layer to obtain a corresponding fine-grained classification loss;

[0059] Module 3 is used for extracting different levels of text semantic information from different levels of intent labels;

[0060] Module 4 is used for putting the three different levels of visual aggregated features obtained and the different levels of text semantic information extracted into the cross-modal pyramid alignment network;

[0061] Module 5 is used for establishing the internal connection between the text features f t h and the visual features in the same level by using metric learning, and obtaining enhanced text features and visual features by using layer normalization.

[0062] The present application has the following advantages:

[0063] (1) The present application proposes a hierarchical relationship mining scheme to adaptively aggregate multi-level visual representations, and improves the global understanding of visual intent by hierarchical modeling.

[0064] (2) The present application directly uses the intent labels with hierarchical structure as text information to serve the visual intent understanding. In addition, the text as auxiliary semantic information does not need additional labeling.

[0065] (3) The present application proposes a cross-modal pyramid alignment scheme between visual and text information and a dynamic joint training strategy, which narrows the domain gap between the two modalities in visual intent understanding, and dynamically optimizes the visual intent understanding performance in a joint learning manner.

[0066] (4) The method proposed in the present application is evaluated on a dataset, and the comprehensive experiments directly prove the superiority of the method, which is better than the existing visual intent understanding method. BRIEF DESCRIPTION OF DRAWINGS

[0067] Figure 1 A hierarchical relationship mining network structure diagram of an embodiment of the present application;

[0068] Figure 2 A cross-modal pyramid alignment network structure diagram of an embodiment of the present application. DETAILED DESCRIPTION

[0069] In order to facilitate those skilled in the art to understand and implement the present application, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0070] The present application comprises three main components. First, hierarchical visual information is captured by hierarchical relationship mining (HRM), and hierarchical text information is captured by a pre-trained model. Among them, both visual and text information are learned from coarse to fine granularity. Then, visual information and text information with the same hierarchical structure are simultaneously input into a cross-modal pyramid alignment module, which is to fuse and promote the contribution of each modality information to visual intent understanding. Finally, a dynamic joint training strategy suppresses the excessive influence of text information, effectively balancing the performance of different modalities in visual intent understanding.

[0071] See Figure 1 and Figure 2 The present application provides a visual intent analysis method based on cross-modal pyramid alignment, comprising the following steps:

[0072] Step 1: Construct a hierarchical relationship mining network, a cross-modal pyramid alignment network and a BERT model;

[0073] See Figure 1 The hierarchical relationship mining network of the present embodiment is used to mine visual hierarchical information; comprising a residual network, a down-sampling layer, a pooling layer and a channel attention module;

[0074] The residual network of the embodiment comprises sequentially connected convolutional layers, a maximum pooling layer, a first module, a second module, a third module and a fourth module; the convolutional kernel size of the convolutional layers and the maximum pooling layer is 7, and the step size is 2; the first module comprises three identical blocks, each block comprising three convolutional layers, the convolutional kernel being 1*1, 3*3 and 1*1 respectively, and the step size being 1; the second module comprises four identical blocks, each block comprising three convolutional layers, the convolutional kernel being 1*1, 3*3 and 1*1 respectively, and the step size being 2; the third module comprises six identical blocks, each block comprising three convolutional layers, the convolutional kernel being 1*1, 3*3 and 1*1 respectively, and the step size being 2; the fourth module comprises three identical blocks, each block comprising three convolutional layers, the convolutional kernel being 1*1, 3*3 and 1*1 respectively, and the step size being 2; a batch normalization layer and an activation layer are arranged after each convolution of the first module, the second module, the third module and the fourth module; each block is connected through a residual connection;

[0075] The down-sampling layer of the embodiment comprises a convolutional layer with a convolutional kernel of 3*3 and a step size of 2, and a batch normalization layer;

[0076] The output of the first module of the embodiment is connected to the output of the second module after passing through the down-sampling layer, then connected to the output of the third module after passing through the down-sampling layer, then connected to the output of the fourth module after passing through the down-sampling layer, and then input into the channel attention module;

[0077] The channel attention module of the embodiment is a convolutional layer with a convolutional kernel of 1*1 and a step size of 1;

[0078] The BERT model of the embodiment is used to extract text features of coarse granularity, medium granularity and fine granularity;

[0079] See Figure 2 The cross-modal pyramid alignment network of the embodiment is used to align the visual hierarchical information extracted by the hierarchical relationship mining network and the text hierarchical features extracted by the BERT model at each hierarchical level (coarse granularity, medium granularity and fine granularity); the cross-modal pyramid alignment network comprises a PIENet and a full connection layer corresponding to each hierarchical level; the PIENet is composed of a multi-head attention module and a residual connection, and the multi-head attention module comprises a full connection layer, an activation layer, a full connection layer and a softmax layer;

[0080] Step 2: input the visual picture to be processed into the hierarchical relationship mining network, extract the visual information of different levels of the picture, and aggregate the features extracted from different levels into aggregated features; the features output by the first module are connected to the features output by the second module through a down-sampling layer to obtain coarse-grained aggregated features; the features output by the second module are connected to the features output by the third module through a down-sampling layer to obtain medium-grained aggregated features; the features output by the third module are connected to the features output by the fourth module through a down-sampling layer to obtain fine-grained aggregated features; the coarse-grained and medium-grained features pass through a pooling layer and a linear layer to obtain corresponding classification losses, and the fine-grained features pass through a channel attention module, a pooling layer and a linear layer to obtain a corresponding fine-grained classification loss;

[0081] In order to obtain a complete feature representation of visual content, the embodiment adopts a hierarchical structure to capture multi-level features from low to high. The features captured by the current convolutional layer (denoted as v h ) are spliced with the aggregated features obtained before (denoted as D(μ h-1 )) to obtain the current aggregated features μ h ∈R n+m , where n and m represent the channel dimensions of v h and D(μ h-1 ) respectively, D(·) represents down-sampling, and the calculation formula of μ h is:

[0082] μ h =concate[ν h ,D(μ h-1 )]

[0083] where n and m represent the channel dimensions of v h and D(μ h-1 ) respectively, D(·) represents down-sampling, and concate[] represents splicing operation.

[0084] Step 3: extract text semantic information of different levels from different levels of intent labels;

[0085] Step 4: put the three different levels of visual aggregated features obtained and the different levels of text semantic information extracted into the cross-modal pyramid alignment network;

[0086] Step 5: use metric learning to establish the internal relationship between the text features f t h and the visual features at the same level, and adopt layer normalization to obtain enhanced text features and visual features Through metric learning, the semantic relationship between the enhanced text and visual features is measured to obtain corresponding modal losses to assist the final intent judgment.

[0087] The hierarchical relationship mining network of the present application utilizes hierarchical classification to capture information of different granularities in visual content; the cross-modal pyramid alignment network utilizes the information to assist visual intention understanding and narrow the differences between different modal information;

[0088] The hierarchical relationship mining network used in the present application is a trained hierarchical relationship mining network, in the training process, first divide a plurality of original pictures into a training set, a test set and a validation set; input the original image to the hierarchical relationship mining network and extract visual hierarchical information, and the extracted text features are dynamically jointly trained by the cross-modal pyramid alignment network; the network parameters are optimized and updated by using forward propagation and back propagation;

[0089] In this embodiment, each aggregated feature μ h Through an average pooling layer and a fully connected layer, it is projected to the corresponding hierarchical intention classification. Since the characteristics of different levels have unique contributions to fine-grained classification, we add a channel attention before the average pool of fine-grained classification, which reweights according to the contribution of different channels to fine-grained visual intention understanding. For visual intention classification, this embodiment calculates the loss of coarse-to-fine-grained classification according to its hierarchical features.

[0090] In the hierarchical relationship mining, the hierarchical classification loss function is:

[0091]

[0092] In the above formula, represents the feature extracted by the last fully connected layer when the ith sample is classified in the hth layer; f j h represents the class score of the jth element in the vector f, and ε h represents the hierarchical weight of the hth layer, is the true value of the current sample i in the hth layer classification, H is the number of levels, and N is the number of the current sample;

[0093] This embodiment extracts different levels of text semantic information from different levels of intent labels. Coarse-grained and medium-grained text information is a combination of the description of fine-grained labels belonging to the same level label. In order to obtain effective initial hierarchical text features, we will describe the input BERT (Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "Bert: Pretraining of deep bidirectional transformers for language understanding," arXiv:1810.04805, 2018.) and use the pre-trained model to capture the initial hierarchical text representation from the coarse to fine level categories on a large English data corpus.

[0094] Due to the domain gap between the visual hierarchy extracted from the pre-trained visual model and the text hierarchy captured from the pre-trained language model, this embodiment refers to PIENet (Y. Song and M. Soleymani, "Polysemous visual-semantic embedding for cross-modal retrieval," in CVPR, 2019.) and combines the metric learning function (K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, "Momentum contrast for unsupervised visual representation learning," in CVPR, 2020. H. Diao, Y. Zhang, L. Ma, and H. Lu, "Similarity reasoning and filtration for image-text matching." in AAAI, 2021.) to establish the internal relationship between text and image at the same level. Due to the particularity of visual intent understanding, images contain rich background and object types. At the same time, as the description of the intent, the text cannot specify a specific object or person. In order to avoid the correspondence between the nouns in the text and the objects in the image, this embodiment uses the self-attention mechanism (Z. Lin, M. Feng, C. N. d. Santos, M. Yu, B. Xiang, B. Zhou, and Y. Bengio, "A structured self-attentive sentence embedding," in ICLR, 2017.) to focus on the important part of each modality, and to link the key area with the important words. This embodiment uses the visual features extracted by the hierarchical relationship mining as the initial visual representation The text features captured by BERT are taken as initial text representations f t h ,

[0095] The self-attention mechanism is as follows:

[0096]

[0097] Here, the initial visual representations and the initial text representations are put into the self-attention mechanism at each different level, d k represents the corresponding feature dimension. When the mechanism is used for the original visual features, Q, K, and V all represent the corresponding When it is used for the original text features, Q, K, and V all represent the corresponding f t h .

[0098] After self-attention, the obtained features are input into a fully connected layer, and then a sigmoid activation function is used to obtain local guiding features. Then, the remaining structure in the network is used to add the newly obtained local features to the initial features. Finally, the instance is normalized by using the layer normalization method (J.L. Ba, J.R. Kiros, and G.E. Hinton, “Layer normalization,” arXiv: 1607.06450, 2016.) to obtain enhanced features and

[0099] During the training process, the self-attention and the metric learning are used to establish the internal connection between the text features f t h and the visual features at the same level, and the layer normalization is used to obtain enhanced text features and visual features

[0100] In the metric learning of the present embodiment, the cosine similarity is first used to calculate the similarity between the modalities, and the calculation formula of the cosine similarity is as follows:

[0101]

[0102] where Dim is the dimension of the current feature, which is used to calculate the distance between the visual information and the text information; then the bidirectional ranking loss is used to train at the hth level, and the formula is as follows:

[0103]

[0104] where m is a boundary parameter, and [x]+≡max(x, 0), representing the similarity between image-text pairs; representing the similarity between image and corresponding negative text, representing the similarity between text and corresponding negative image;

[0105] The embodiment of the cross-modal pyramid alignment network considers all corresponding levels of the two modes to enhance the understanding of the intention; the optimization strategy is as follows:

[0106]

[0107] wherein, δ h representing The level weight of the corresponding h layer;

[0108] Since the initial text features extracted by the pre-training model contain limited intention information, the embodiment adaptively integrates hierarchical relationship mining (HRM) and cross-modal pyramid alignment (CPA) in a single framework by designing a dynamic joint training strategy, so as to balance the contribution of text and image pairs to visual intention understanding. The basic idea is to integrate the hierarchical classification loss Λ hrm as the dominant loss, and gradually integrate the cross-modal pyramid alignment loss Λ cpa for optimization, the formula is as follows:

[0109]

[0110] wherein, E(·) is the average value of the hierarchical classification of the last period, and α is a constraint parameter, which controls the contribution of the entire training process.

[0111] The following further illustrates the present application through experiments. The deep learning framework used in the experiment is Pytorch. The hardware environment of the experiment is NVIDIA GeForce RTX 3090*8 graphics card, and the processor is Intel(R) Xeon(R) Gold 6240.

[0112] The experiment evaluates the method of the embodiment on the Intentonomy dataset. Intentonomy divides all the intention images into 9 coarse classes and 28 fine classes, of which 12740, 498 and 1217 images are used for training, verification and testing respectively. In particular, the segmentation of Intentonomy is different between the published github and the paper. The experiment uses the partition dataset available in the github of Intentonomy. The objects and background content in the Intentonomy dataset are diverse, covering a wide range of daily life scenes such as parties, vacations and work.

[0113] The experiment takes ResNet50 as the backbone network to extract effective features, and the backbone network is pre-trained on a large-scale dataset ImageNet. Since the image is high resolution, the experiment first adjusts the longest side of the image to 1280 in proportion. Standard data augmentation is applied on the image, including random horizontal flipping and random cropping size adjustment to 224x224 and training with zero padding. The experiment uses stochastic gradient descent (SGD) as the optimizer, with a momentum value of 0.9 and a batch size of 128. For the learning rate, the experiment increases from 0 to 1e-3 through linear warm-up in the first 5 epochs. The HRM different level weights ε h and the CPA different weights δ h are set to 1, and α is set to 0.1. In addition, the experiment directly assigns each fine-grained label to the corresponding medium / coarse label. Therefore, the model of the experiment does not introduce additional labels, and the corresponding medium / coarse label does not have additional multi-level label annotation costs.

[0114] In order to verify the effectiveness of the present application, the classification results of the present application are compared with existing advanced visual intent classification methods, which mainly include:

[0115] (1) Visual: K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016.

[0116] (2) Jia: M. Jia, Z. Wu, A. Reiter, C. Cardie, S. Belongie, and S.-N. Lim, “Intentonomy: a dataset and study towards human intent understanding,” in CVPR, 2021.

[0117] In particular, since the train / val / test segmentation in the Jia paper is different from that in the released dataset, the experiment uses the segmentation method released on github.

[0118] Different visual intent classification methods are compared in Macro F1, Micro F1, and samples F1, and (+ / +) represents the improvement compared with the Visual method. The results are shown in Table 1.

[0119] Table 1

[0120]

[0121]

[0122] In addition, the Intentonomy dataset is divided into different sub-classes according to the key content (object-class, content-class, others-class) and the classification difficulty (Easy, Medium, Hard).

[0123] Different visual intent classification methods are compared according to different visual content. The intent categories are divided into object-dependent, context-dependent and others. The results are shown in Table 2.

[0124] Table 2

[0125]

[0126] Different visual intent classification methods are compared according to the classification difficulty. According to the difficulty, the intent categories are divided into three categories: easy, medium and difficult. The results are shown in Table 3.

[0127] Table 3

[0128]

[0129]

[0130] From Table 1 and Table 2, it can be seen that compared with Jia, the hierarchical relationship mining (HRM) of the embodiment has more improvement without text information. At the same time, compared with the method of adding text information, the cross-modal pyramid alignment of the embodiment significantly improves the recognition performance (Macro F1 value: 25.07%→27.37%, Micro F1 value: 32.94%→41.77%, Samples F1 value: 33.61%→42.68%).

[0131] From Table 2 and Table 3, it can be seen that the average score of the content and the difficulty of the embodiment is the highest. In addition, compared with Jia which collects label supplemented visual information (content +1.79% vs +5.03%, difficulty +0.93% vs +2.21%), the un-labeled intent label as text information in the method of the embodiment is more effective.

[0132] It should be understood that the above description of the preferred embodiments is more detailed, and therefore should not be considered as limiting the scope of protection of the present application. Those skilled in the art can make substitutions or modifications without departing from the scope of protection of the present application, and all fall within the scope of protection of the present application. The scope of protection of the present application should be subject to the appended claims.

Claims

1. A visual intent analysis method based on cross-modal pyramid alignment, characterized in that, The method comprises the following steps: Step 1: constructing a hierarchical relationship mining network, a cross-modal pyramid alignment network and a BERT model; The hierarchical relationship mining network is used for mining visual hierarchical information and comprises a residual network, a down-sampling layer, a pooling layer and a channel attention module; The residual network comprises sequentially connected convolution layers, a max-pooling layer, a first module, a second module, a third module and a fourth module; the convolution layers and the max-pooling layer have a convolution kernel size of 7 and a step size of 2; the first module comprises three blocks, each block comprising three convolution layers with a convolution kernel of 1x1, 3x3 and 1x1 and a step size of 1; the second module comprises four blocks, each block comprising three convolution layers with a convolution kernel of 1x1, 3x3 and 1x1 and a step size of 2; the third module comprises six blocks, each block comprising three convolution layers with a convolution kernel of 1x1, 3x3 and 1x1 and a step size of 2; the fourth module comprises three blocks, each block comprising three convolution layers with a convolution kernel of 1x1, 3x3 and 1x1 and a step size of 2; each convolution of the first module, the second module, the third module and the fourth module is provided with a batch normalization layer and an activation layer; and each block is connected through a residual connection; The down-sampling layer comprises a convolution layer with a convolution kernel of 3x3 and a step size of 2 and a batch normalization layer; The output of the first module is connected to the output of the second module after passing through the down-sampling layer, then connected to the output of the third module after passing through the down-sampling layer, then connected to the output of the fourth module after passing through the down-sampling layer, and then input into the channel attention module; The channel attention module is a convolution layer with a convolution kernel of 1x1 and a step size of 1; The BERT model is used for extracting coarse-grained, medium-grained and fine-grained text features; The cross-modal pyramid alignment network is used for aligning the visual hierarchical information extracted by the hierarchical relationship mining network and the text hierarchical features extracted by the BERT model at each level; The cross-modal pyramid alignment network comprises a PIENet and a full connection layer corresponding to each level; the PIENet is composed of a multi-head attention module and a residual connection, and the multi-head attention module comprises a full connection layer, an activation layer, a full connection layer and a softmax layer; Step 2: inputting a visual picture to be processed into the hierarchical relationship mining network, extracting visual information of different levels of the picture, and aggregating the features extracted from different levels into aggregated features; The features output by the first module are connected to the features output by the second module after passing through the down-sampling layer to obtain coarse-grained aggregated features; The features output by the second module are connected to the features output by the third module after passing through the down-sampling layer to obtain medium-grained aggregated features; The features output by the third module are connected to the features output by the fourth module after passing through the down-sampling layer to obtain fine-grained aggregated features; the coarse-grained and medium-grained features pass through a pooling layer and a linear layer to obtain corresponding classification losses, and the fine-grained features pass through a channel attention module, a pooling layer and a linear layer to obtain a corresponding fine-grained classification loss; Step 3: extracting different levels of text semantic information from different levels of intention labels; Step 4: putting the obtained three different levels of visual aggregation features and the extracted different levels of text semantic information into the cross-modal pyramid alignment network; Step 5: Establishing the internal connection between the text features f and the visual features v at the same level by using metric learning, and obtaining enhanced text features f and visual features v by layer normalization t h and visual features and visual features By measuring the semantic relationship between the enhanced text and visual features through metric learning, the corresponding modal loss is obtained to assist the final intent judgment.​ 2. The visual intent analysis method based on cross-modal pyramid alignment according to claim 1, characterized in that: In step 2, the features ν h captured by the current convolutional layer are concatenated with the previously obtained aggregated features D(μ h-1 ) to obtain the current aggregated features μ h ∈R n+m ; μ h = concate[ν h , D(μ h-1 )] where n and m denote the channel dimensions of X h and D(μ h-1 ), respectively, and D(·) denotes downsampling and concate[] denotes concatenation operation.

3. The visual intent analysis method based on cross-modal pyramid alignment according to claim 1 or 2, characterized in that: The hierarchical relationship mining network captures information of different granularities in visual content through hierarchical classification; The cross-modal pyramid alignment network uses the text information to assist visual intention understanding and narrow the difference between different modal information; During the training process, first, a plurality of original pictures are divided into a training set, a test set and a validation set; the original image is input into the hierarchical relationship mining network to extract visual hierarchical information, and the extracted text features are dynamically jointly trained through the cross-modal pyramid alignment network; the network parameters are optimized and updated through forward propagation and back propagation; In the hierarchical relationship mining, the hierarchical classification loss function is: In the above formula, represents the feature extracted by the last fully connected layer when the ith sample is classified at the hth layer; f j h represents the class score of the jth element in the vector f, ε h represents the hierarchical weight of the hth layer, is the true value of the current sample i in the hth layer classification, H is the number of layers, and N is the number of current samples; During training, self-attention and metric learning are used to establish the internal connection between text features f t h and visual features at the same level, and layer normalization is used to obtain enhanced text features and visual features In the metric learning, the similarity between the modes is calculated using the cosine similarity, and the calculation formula of the cosine similarity is: where Dim is the dimension of the current feature, used to calculate the distance between visual information and text information; then, the bidirectional ranking loss is used to train the h layer, and the formula is: where m is a margin parameter, [x] +≡ max(x, 0), denotes the similarity between an image-text pair; denotes the similarity between an image and a corresponding negative text, denotes the similarity between a text and a corresponding negative image; The cross-modal pyramid alignment network considers all corresponding levels of the two modes to enhance intention understanding; the optimization strategy is as follows: where δ h represents hierarchy weight corresponding to the h layer; The hierarchical classification loss Λ hrm As the dominant loss, the cross-modal pyramid alignment loss is gradually integrated Optimization is performed, as follows: where E(·) is the average of the previous epoch's hierarchical classification, and a is a constraint parameter that controls contribution to the overall training process.

4. A visual intent analysis system based on cross-modal pyramid alignment, characterized by, The method comprises the following modules: Module 1, used for constructing a hierarchical relationship mining network, a cross-modal pyramid alignment network and a BERT model; The hierarchical relationship mining network is used for mining visual hierarchical information; and comprises a residual network, a down-sampling layer, a pooling layer and a channel attention module; The residual network comprises sequentially connected convolutional layers, a maximum pooling layer, a first module, a second module, a third module and a fourth module; the convolutional layers and the maximum pooling layer have a convolution kernel size of 7 and a step size of 2; the first module comprises 3 blocks, each block comprising three convolutional layers with a convolution kernel size of 1x1, 3x3 and 1x1 and a step size of 1; the second module comprises 4 blocks, each block comprising three convolutional layers with a convolution kernel size of 1x1, 3x3 and 1x1 and a step size of 2; the third module comprises 6 blocks, each block comprising three convolutional layers with a convolution kernel size of 1x1, 3x3 and 1x1 and a step size of 2; and the fourth module comprises 3 blocks, each block comprising three convolutional layers with a convolution kernel size of 1x1, 3x3 and 1x1 and a step size of 2; each convolutional layer of the first module, the second module, the third module and the fourth module is provided with a batch normalization layer and an activation layer after the end of the convolution; and each block is connected through a residual connection; The down-sampling layer comprises a convolutional layer with a convolution kernel size of 3x3 and a step size of 2 and a batch normalization layer; The output of the first module is connected to the output of the second module after passing through the down-sampling layer, then connected to the output of the third module after passing through the down-sampling layer, then connected to the output of the fourth module after passing through the down-sampling layer, and finally input into the channel attention module; The channel attention module is a convolution layer with a convolution kernel of 1*1 and a step length of 1; The BERT model is used to extract text features of coarse granularity, medium granularity and fine granularity; The cross-modal pyramid alignment network is used to align the visual hierarchical information extracted by the hierarchical relationship mining network and the text hierarchical features extracted by the BERT model at each level; The PIENet is composed of a multi-head attention module and a residual connection, and the multi-head attention module is composed of a full connection layer, an activation layer, a full connection layer and a softmax layer; Module 2 is used to input the visual picture to be processed into the hierarchical relationship mining network, extract visual information of different levels of the picture, and aggregate the features extracted from different layers into aggregated features; The features output by the first module are connected to the features output by the second module through a down-sampling layer to obtain coarse-grained aggregated features; The features output by the second module are connected to the features output by the third module through a down-sampling layer to obtain medium-grained aggregated features; The features output by the third module are connected to the features output by the fourth module through a down-sampling layer to obtain fine-grained aggregated features; the coarse-grained and medium-grained features are subjected to a pooling layer and a linear layer to obtain corresponding classification losses, and the fine-grained features are subjected to a channel attention module, a pooling layer and a linear layer to obtain a corresponding fine-grained classification loss; Module 3 is used to extract text semantic information of different levels from different levels of intent labels; Module 4 is used to put the three different levels of visual aggregated features obtained and the different levels of text semantic information extracted into the cross-modal pyramid alignment network; Module 5, for establishing the internal connection between the text features f and the visual features v at the same level by using metric learning, and obtaining enhanced text features by using layer normalization t h and visual features and visual features The semantic relationship between the enhanced text and visual features is measured by metric learning, and the corresponding modal loss is obtained to assist the final intent judgment.​ 5. The visual intent analysis system based on cross-modal pyramid alignment according to claim 4, characterized in that: In module 2, the features ν h captured by the current convolutional layer are concatenated with the previously obtained aggregated features D(μ h-1 ), resulting in the current aggregated features μ h ∈R n+m ; μ h = concate [v h , D(μ h-1 )] where n and m represent the channel dimensions of X h and D(μ h-1 ), respectively, D(·) denotes downsampling, and concate[] denotes concatenation operation.

6. The visual intent analysis system based on cross-modal pyramid alignment according to claim 4 or 5, characterized in that: The hierarchical relationship mining network captures information of different granularities in the visual content by using hierarchical classification; The cross-modal pyramid alignment network uses the text information to assist visual intent understanding and reduces the difference between different modal information; In the training process, first, a plurality of original pictures are divided into a training set, a test set and a validation set; the original image is input into the hierarchical relationship mining network and visual hierarchical information is extracted, and the extracted text features are subjected to dynamic joint training through the cross-modal pyramid alignment network; the network parameters are optimized and updated by using forward propagation and back propagation; In the hierarchical relationship mining, the hierarchical classification loss function is: In the above formula, represents the feature extracted by the last fully connected layer when the ith sample is classified at the hth layer; f j h represents the class score of the jth element in the vector f, ε h represents the hierarchical weight of the hth layer, is the true value of the current sample i in the hth layer classification, H is the number of hierarchies, and N is the number of current samples; During training, self-attention and metric learning are used to establish the internal connection between text features f t h and visual features at the same level, and layer normalization is used to obtain enhanced text features and visual features In the metric learning, the similarity between the modes is calculated using the cosine similarity, and the calculation formula of the cosine similarity is: Wherein, Dim is the dimension of the current feature, which is used to calculate the distance between the visual information and the text information; then, the bidirectional ranking loss is used to train the h layer, and the formula is: where m is a margin parameter, [x] + ≡ max(x, 0), denotes the similarity between an image-text pair; denotes the similarity between an image and a corresponding negative text, denotes the similarity between a text and a corresponding negative image; The cross-modal pyramid alignment network considers all corresponding levels of the two modes to enhance intent understanding; the optimization strategy is as follows: where δ h represents hierarchy weight corresponding to the h layer; The hierarchical classification loss Λ hrm As the dominant loss, the cross-modal pyramid alignment loss is gradually integrated Optimization is performed, as follows: where E(·) is the average of the previous epoch's hierarchical classification, and a is a constraint parameter that controls the contribution to the overall training process.

Citation Information

Patent Citations

  • Semantic image segmentation method and system based on edge enhancement

    CN111462126A

  • Depth map super-resolution method and system based on uncertainty perception feature transmission

    CN115511708A