Multi-task Learning Image Memorability Prediction Method Based on Time-Space Relationship

By introducing multi-task learning and multi-head attention mechanisms into the image memory prediction model, the memory prediction of image features over time is simulated, which solves the problem of insufficient prediction of existing models on data sets that can better reflect human consistency, and significantly improves prediction accuracy.

CN115272829BActive Publication Date: 2025-06-27TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211004849.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-22
Publication Date
2025-06-27
Estimated Expiration
2042-08-22

AI Technical Summary

Technical Problem

Existing image memory prediction models, such as AMnet, are difficult to adapt to datasets that better reflect human consistency, such as SUNMemorability dataset, with low prediction values, indicating that the model ignores certain features when simulating image memory.

Method used

The multi-task learning image memory prediction method based on time-space relationship is adopted, and the surface visual features and high-level semantic features are extracted through the Resnet101 residual network, and the memory prediction of the changes in features over time is simulated using LSTM. Combined with the multi-head attention mechanism, strengthen the contribution and relationship of different characteristics to memory and improve the accuracy of model prediction.

Benefits of technology

The prediction accuracy of image memory on different data sets is significantly improved, especially on the SUNMemorability data set. The predicted value of the model is close to the theoretical value of human consistency, proving the effectiveness of the model in simulating the impact of different characteristics on memory in the human brain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272829B_ABST
    Figure CN115272829B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for predicting image memorability based on time - space relationship, comprising the following steps: constructing a training set and a test set; passing an image through a Resnet101 residual network to obtain three different spatial component feature information; inputting the three feature information into three LSTM registers with the same scale and 1024 channels to simulate the process of memory changing over time; the core of the LSTM register is a recurrent neural network RNN; obtaining enhanced spatial relationship; obtaining short - term memory, first intermediate memory, second intermediate memory, and long - term memory y<supgt;4< / supgt>; obtaining the final state of surface visual feature information related to enhanced short - term memory and the final state of high - level semantic feature information related to long - term memory; recombining the processed three - type spatial feature information and linearly combining them with three weight parameters to obtain the final memorability prediction value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for predicting image memorability in multi-task learning. Background Art

[0002] Different images leave different impressions on people. Although the results vary from person to person, the results of large-scale population statistics show a certain consistency. Image memorability (between 0 and 1) is a quantitative representation of the impression left by such an image. Whether it is person-to-person, machine-to-person communication or storing visual content, visual information is an important cognitive measure to be considered when processing visual content. Being able to find or design a memorable memorability can help make some very memorable presentations and data visualizations, improve the memorability of specific parts of the graphical user interface (GUI), or help treat some diseases with memory decline, such as Alzheimer's disease and Parkinson's disease. Commercially, image memorability is applied to advertising production to achieve better promotion and publicity effects.

[0003] In the field of deep learning, it has been proved by P. Isola, J. Xiao, A. Torralba, and A. Oliva that for a group, image memorability has a stable property [1], that is, individuals in the group tend to remember the same images with the same probability, regardless of the delay, and it can be quantified and measured. Therefore, A. Khosla, A. S. Raju, A. Torralba, and A. Oliva et al. introduced the LaMem dataset and proved the above view through experiments, and at the same time called this property human consistency [2]. A. Krizhevsky, I. Sutskever, and G. E. Hinton et al. trained this data and named it the AlexNet network [3]. The predicted value of human consistency for the LaMem dataset on this network is 0.64, which is obviously related to the measured value of 0.68 obtained by A. Khosla in the experiment. Subsequently, Jiri Fajtl1, Vasileios Argyriou1 et al. found that the image regions that immediately attract people's attention in the image seem to be associated with highly memorable visual content, and then introduced the attention mechanism into the training and prediction of image memorability [4]. The resulting AMnet reached a predicted value of 0.678 for human consistency in the LaMem dataset, which is very close to the theoretical value of 0.68.

[0004] Although AMnet has achieved good prediction results on the LaMem dataset, it is difficult to adapt to datasets that better reflect human consistency. For example, the theoretical value of human consistency for the SUNMemorability dataset is 0.76, while the predicted value of AMnet is only 0.649, indicating that it cannot simulate the human consistency of datasets with more consistent image memorability.

[0005] Considering that the predicted values of AMnet are relatively low on different datasets, we believe that this is due to its inherent defects, which prevent it from considering various information of images like a crowd, leading to the neglect of certain features. According to the research of related patent [5], the prediction of image memorability is closely related to high-level semantic features. According to another related patent [6], there are mutual influences among the low-level features, image attribute features, and image memorability of images.

[0006] References

[0007] [1] P. Isola, J. Xiao, D. Parikh, A. Torralba, and A. Oliva. What makes a photograph memorable? IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7): 1469–1482, July 2014.

[0008] [2] A. Khosla, A. S. Raju, A. Torralba, and A. Oliva. Understanding and predicting image memorability at a large scale. In Proceedings of the IEEE International Conference on Computer Vision, pages 2390–2398, 2015.

[0009] [3] A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, volume 25, 2012.

[0010] [4] Fajtl, J., et al. "AMNet: Memorability Estimation with Attention." 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition IEEE, 2018.

[0011] [5] Chu Jinghui, Gu Huimin, Su Yuting, Jing Peiguang. A method for predicting image memorability based on sparse low-rank regression model [P]. Tianjin: CN107909091A, 2018-04-13.

[0012] [6] Liu An'an, Shi Yingdi, Su Yuting. A multi-view learning method combining low-rank representation and sparse regression [P]. Tianjin: CN107545276A, 2018-01-05.

[0013] [7] Liu An'an, Shi Yingdi, Su Yuting. A learning method combining low-rank representation and sparse regression [P]. Tianjin: CN107590505A, 2018-01-16.

[0014] [8] M. Mancas and O. Le Meur. Memorability of natural scenes: The role of attention. In International Conference on Image Processing, pages 196–200. IEEE, 2013.

[0015] [9]. Khosla, A. Das Sarma, and R. Hamid. What makes an image popular? In Proceedings of the international conference on Worldwide Web, pages 867–876. ACM, 2014. Summary of the Invention

[0016] The object of the present invention is to provide a multi-task learning method for predicting image memorability, which simulates the influence of different features on memorability by the human brain. The technical solution is as follows:

[0017] A multi-task learning method for predicting image memorability based on time-space relationship, comprising the following steps:

[0018] The first step is to use the publicly available SUN Memorability dataset to obtain a training set and a test set;

[0019] In the second step, the image is passed through the Resnet101 residual network to obtain three different spatial component feature information: high-level semantic feature information X0 (1024×14×14), intermediate feature information X1 (1024×28×28), and surface visual feature information X2 (1024×56×56);

[0020] In the third step, the three types of feature information are input into three LSTM registers with the same scale and 1024 channels. The method is as follows: First, obtain the overall information. Use the view() function to reduce the dimensionality of the three different spatial component feature information from three dimensions to two dimensions, which are 1024×196, 1024×784, and 1024×3136 respectively. Use the average value of the second dimension of each feature information to construct three new 1024×1 vectors. Through 2 different Linear linear convolution functions, the initial state and initial parameters of the LSTM register are generated, and the initial state h of the LSTM register with a dimension of 1024×1 is obtained 00 ,h 01 ,h 02 and the initial parameters cs0, cs1, cs2; Use the initial state h i corresponding to each feature information X 0i and the initial parameter cs i to construct three LSTM registers with 1024 channels. Then use the view() function to restore the feature information to 1024×14×14, 1024×28×28, and 1024×56×56 respectively, and input them into the corresponding LSTM registers to simulate the process of memory changing over time;

[0021] In the third step, the core of the LSTM register is the recurrent neural network RNN. The output of the RNN is related not only to the current input but also to the previous output. In any LSTM register, a new RNN and a new register state h are generated after each time iteration. Set the number of iteration steps to 4; After obtaining the new register state h for each iteration, through continuous dimensionality reduction, activation, removal of randomness, and then dimensionality reduction of h, the memory degree prediction component out_step at each iteration step is obtained; Each type of feature information iterates 4 steps in the corresponding LSTM register, and 4 LSTM register states will be generated; The 4 register states generated by the surface visual feature information are combined into H0, the 4 register states generated by the intermediate feature information are combined into H1, and the 4 register states generated by the high-level semantic feature information are combined into H2; The 4 memory degree prediction components generated by the i-th feature among the three spatial component features are: out_step1i, out_step2i, out_step3i, out_step4i;

[0022] Step 4: Take H0, H1, and H2 as three inputs and input them into the first multi-head attention network to obtain H i and H i its own relationship Hatten, where the relationship Hatten represents the connection between features in space; superimpose the obtained relationship Hatten on the original H i to get H′ i = H i + Hatten, thus obtaining H′0, H′1, and H′2 that strengthen the spatial relationship;

[0023] Step 5: Combine the 4-step memory prediction components out_step of the three LSTM registers with 1024 channels obtained in Step 3 into Y (3×4 dimensions), and input three identical Ys into the second multi-head attention network to obtain the relationship Yatten between Y and Y itself, where the relationship Yatten represents the connection between tasks in time; add the relationship Yatten to the original Y, i.e., Y′ = Y + Yatten, to get Y′ that strengthens the spatial relationship, and Y′ is also 3×4 dimensions, thus obtaining short-term memory y1, first intermediate memory y2, second intermediate memory y3, and long-term memory y4;

[0024] (5) Stretch the long-term memory y4 and short-term memory y1 to 3×1024 dimensions using linear transformation to obtain Ky4 and Ky1 respectively; the stretched short-term memory Ky1 and H′0 are input into the third multi-head attention mechanism network to obtain the relationship XYatten1; the stretched long-term memory Ky4 and H′2 are input into the fourth multi-head attention mechanism network to obtain the relationship XYatten2; add the relationship XYatten1 and the relationship XYatten2 to H′0 and H′2 respectively to obtain the final state H″0 of the surface visual feature information strengthened related to short-term memory and the final state H″2 of the high-level semantic feature information strengthened related to long-term memory;

[0025] (6) Recombine the three processed spatial feature information H″0, H′1, and H″2, and after the first convolution for dimensionality reduction, dropout(0.5) function, relu() function, and the second convolution for dimensionality reduction, obtain the final three memory prediction components p0, p1, and p2, and then linearly combine them with three weight parameters to obtain the final memory prediction value p = p0×a + p1×b + p2×c.

[0026] The substantial features and beneficial effects of the present invention are as follows:

[0027] (1) Use the shallow layer of the residual network to extract surface visual features, and use the deep layer of the residual network to extract high-level semantic features. Simulate the influence of different features on the memory degree by the human brain.

[0028] (2) Introduce multi-task learning into the research of memorability. Regard the results output by different feature extraction layers as the results of different task learning regarding the memorability of images. And use the self-attention mechanism to strengthen the relationship between time and space caused by different features.

[0029] (3) Adopt the multi-head attention mechanism to strengthen the mutual influence between the shallow and deep networks, strengthen the contribution of surface visual features to short-term memorability and high-level semantic features to long-term memorability, and further improve the model prediction accuracy. Description of the Drawings

[0030] Figure 1 It is a general view of the model, including multi-layer feature extraction, LSTM network memorability prediction, and multi-task attention mechanism.

[0031] Figure 2 It is to extract low-level features and high-level features of images using a multi-layer network.

[0032] Figure 3 It is for multi-task learning to separate the influence of different features on memorability and achieve prediction.

[0033] Figure 4 It is for the convergence of the human consistency prediction value RC in the test set. The light color is the experimental data, and the dark color is the simulated data.

[0034] Figure 5 It is for the convergence of the mean square error MSE between the predicted value and the actual value in the test set. The light color is the experimental data, and the dark color is the simulated data.

[0035] Table 1 shows the comparison of experimental results of whether to introduce the attention mechanism in the SUN Memorability dataset.

[0036] Table 1 Comparison of experimental results of whether to introduce the attention mechanism in the sun dataset

[0037] Detailed Implementation Manner

[0038] The present invention proposes a multi-task learning image memorability prediction based on time-space relationship on AMnet. The surface visual features and high-level semantic features of an image are extracted through a CNN network, and then an LSTM is used to simulate the change of each feature's memorability over time. A self-attention mechanism is used to strengthen the influence of the surface visual features and high-level semantic features on the memorability contribution. A self-attention mechanism is used to strengthen the mutual relationship between the surface visual features and high-level semantic features. A T-F attention mechanism is used to strengthen the influence of high-level semantic features on long-term memorability and strengthen the influence of surface visual features on short-term memorability. By extending AMnet in the time dimension and space dimension and introducing a transformer to study multiple attention mechanisms, better prediction results are achieved.

[0039] The following will describe the implementation manner in detail with reference to the drawings:

[0040] The first step is to download the dataset and perform preprocessing: Use the publicly available SUN Memorability dataset, which contains 2,222 images. Each image corresponds to an experimental value of image memorability (i.e., the so-called true value). The size of each image is 16×16. Five groups of data are randomly divided according to the ratio of test set: training set of 1:1 for training respectively to ensure the randomness of model training. The images are further preprocessed by resizing them to 16×16, randomly flipping them horizontally, and finally performing a normalization operation.

[0041] The second step is that the entire network includes obtaining different spatial feature information through a residual network; obtaining different time memorability prediction components through an LSTM register; strengthening the relationship between the spatial feature information and the time memorability prediction components with each other and themselves through an attention mechanism; and regressing to the final memorability prediction value. The following will introduce these four parts one by one:

[0042] (1) Pass the image through a Resnet101 residual network to obtain different features. Among them, the output of layer1 is the surface visual features with a size of 256×56×56, the output of layer2 is the intermediate features with a size of 512×28×28, and the output of layer3 is the high-level semantic features with a size of 1024×14×14. In order to pass through the same LSTM register later and keep the same number of channels, the three types of features are convolved, and the number of channels after convolution is 1024. To accelerate training, standardize the data, eliminate gradient explosion, and then let these three features pass through the corresponding scale of convolution-inconv layer, batch normalization - BN layer, relu layer, and drop layer. The high-level semantic feature information X0 (1024×14×14), intermediate feature information X1 (1024×28×28), and surface visual feature information X2 (1024×56×56) are obtained.

[0043] (2) In order to input the three types of feature information into three LSTM registers of the same scale with 1024 channels. The specific approach is as follows: First, obtain the overall information. Use the view() function to reduce the dimensionality of the three types of feature information from three dimensions to two dimensions (1024×196, 1024×784, 1024×3136 respectively). Utilize the average value of the second dimension of each type of feature information and 1024 to form three new 1024×1 vectors. Then, generate the initial state and initial parameters of the LSTM register through 2 different Linear linear convolution functions, obtaining the initial state h of the LSTM register with a dimension of 1024×1 00 ,h 01 ,h 02 and the initial parameters cs0, cs1, cs2. Use the initial state h i corresponding to each feature information X 0i and the parameter cs i to construct three LSTM registers with 1024 channels. Then use the view() function to restore the feature information to 1024×14×14, 1024×28×28, 1024×56×56 and input it into the corresponding LSTM register to simulate the process of memory changing over time.

[0044] (3) The core of the LSTM register is a special recurrent neural network RNN. The output result of this RNN is related not only to the current input but also to the previous output. In any LSTM register, a new RNN and a new register state h are generated after each time iteration. Set the number of iteration steps to 4. After obtaining h for each iteration, through continuous dimensionality reduction, activation, removal of randomness, and further dimensionality reduction of h, the memory degree prediction component out_step at the current iteration step is obtained. To sum up, an image corresponds to three spatial component feature information, namely high-level semantic feature information, intermediate feature information, and surface visual feature information. Each feature information iterates 4 steps in the corresponding LSTM register, generating 4 LSTM register states. The 4 register states generated by the surface visual feature information are combined as H0, the 4 register states generated by the intermediate feature information are combined as H1, and the 4 register states generated by the high-level semantic feature information are combined as H2. The 4 memory degree prediction components generated by the i-th feature among the three spatial component features are: out_step1i, out_step2i, out_step3i, out_step4i.

[0045] (3) Use H0, H1, and H2 as the three inputs and input them into the first multi-head attention network to obtain H i and H iThe relationship with itself, which is the connection between features in space. Hatten = self.Multiheadattention(H0, H1, H2). Add this connection to the original H i to get H′ i = H i + Hatten to obtain H′0, H′1, H′2 with enhanced spatial relationships.

[0046] (4) Combine the 4-step prediction results out_step of the LSTM registers of the 3 spatial states in (2) into Y (3×4 dimensions), and input three identical Ys into the second multi-head attention network to obtain the relationship between Y and Y itself, which is the connection between tasks in time. Yatten = self.Multiheadattention(Y, Y, Y). Add this connection to the original Y, and Y′ = Y + Yatten to obtain Y′ with enhanced spatial relationships. Y′ is also 3×4 dimensions and can be divided into short-term memory y1, first intermediate memory y2, second intermediate memory y3, and long-term memory y4 according to the second dimension.

[0047] (5) Stretch the long-term memory y4 and short-term memory y1 to 3×1024 dimensions using linear transformation to obtain Ky4 and Ky1 respectively. The stretched short-term memory Ky1 and H′0 are input into the third multi-head attention mechanism network together; the stretched long-term memory Ky4 and H′2 are input into the fourth multi-head attention mechanism network together. Obtain the spatio-temporal relationship, and add this relationship to H′0 and H′2 to obtain the final state H″0 of the surface visual feature information related to short-term memory and the final state H″2 of the high-level semantic feature information related to long-term memory.

[0048] XYatten1 = self.Multiheadattention(H′2, Ky4, Ky4), H″2 = H′2 + XYatten1.

[0049] XYatten2 = self.Multiheadattention(H′0, Ky1, Ky1), H″0 = H′0 + XYatten2.

[0050] (6) Recombine the three processed spatial feature information H″0, H′1, H″2. After the first convolution for dimensionality reduction, dropout(0.5) function, relu() function, and the second convolution for dimensionality reduction, obtain the final three memory degree prediction components p0, p1, p2, and then linearly combine them with three weight parameters to obtain the final memory degree prediction value p = p0×a + p1×b + p2×c.

[0051] Step 3: Model training. Calculate the human consistency parameter RC and the mean squared error MSE between the final memory prediction value and the actual value. Set LOSS equal to MSE, and perform backpropagation layer by layer from the output layer to the hidden layer to update the network parameters. Use the Adam optimizer to continuously feedback and optimize until the error no longer decreases. The initial learning rate is set to 0.00001. It is decreased to 1 / 10 of the original every 10 rounds. The batch_size is set to 16, and the number of training times is set to 50. According to Figure 4 , Figure 5 As shown, after training for 25 epochs, the two metrics converge and the model tends to be stable.

[0052] Step 4: Input the five randomly grouped test sets into the model trained for the training set respectively to obtain the corresponding experimental results, and average and record the results.

[0053] Step 5: This model uses the simulated value RC (Spearman’s rank correlation) of human consistency and the mean squared error MSE to measure the model prediction effect. Calculate according to the memory degree of the actual images in the test set. As shown in Table 1, compared with not adding any attention mechanism, introducing the spatial attention mechanism, the temporal attention mechanism, and the spatio-temporal attention mechanism can all improve RC and reduce MSE. The model that introduces all attention mechanisms simultaneously has the best prediction effect, and RC = 0.704516 and MSE = 0.009969 are calculated, which is the closest to the true value.

Claims

1. A multi-task learning image memorability prediction method based on time-space relationship, comprising the following steps: In the first step, use the publicly available SUN Memorability dataset to obtain a training set and a test set; Step 2: Pass the image through the Resnet101 residual network to obtain three different spatial component feature information: high-level semantic feature information X0, intermediate feature information X1, and surface visual feature information X2; among them, The dimension of X0 is 1024×14×14, the dimension of the intermediate feature information X1 is 1024×28×28, and the dimension of the surface visual feature information X2 is 1024×56×56; Step 3: Input the three types of feature information into three LSTM registers with the same scale and 1024 channels as follows: First, obtain the overall information. Use the view() function to reduce the dimensionality of the three different spatial component feature information from three dimensions to two dimensions, which are 1024×196, 1024×784, and 1024×3136 respectively. Construct three new 1024×1 vectors using the average value of the second dimension of each type of feature information. Generate the initial state and initial parameters of the LSTM register through 2 different Linear linear convolution functions, and obtain the initial state h of the LSTM register with a dimension of 1024×1 00 , h 01 , h 02 and the initial parameters cs0, cs1, cs2; Use each feature information X i , i = 0, 1, 2, the corresponding initial state h 0i and the initial parameter cs i to construct three LSTM registers with 1024 channels. Then use the view() function to restore the feature information to 1024×14×14, 1024×28×28, and 1024×56×56 respectively, and input them into the corresponding LSTM registers to simulate the process of memory changing over time; In the third step, the core of the LSTM register is the recurrent neural network RNN. The output of the RNN is related not only to the current input but also to the previous output. In any LSTM register, a new RNN and a new register state h are generated after each time iteration. Set the number of iteration steps to 4; After obtaining a new register state h in each iteration, through continuous dimensionality reduction, activation, removal of randomness, and then dimensionality reduction of h, the memorability prediction component out_step at each iteration step is obtained; Each type of feature information iterates 4 steps in the corresponding LSTM register, generating 4 LSTM register states; the 4 register states generated by the surface visual feature information are combined into H0, the 4 register states generated by the intermediate feature information are combined into H1, and the 4 register states generated by the high-level semantic feature information are combined into H2; the 4 memorability prediction components generated by the i-th feature among the three spatial component features are respectively: out_step1i, out_step2i, out_step3i, out_step4i; Step 4: Take H0, H1, and H2 as the three inputs and input them into the first multi-head attention network to obtain H i The relationship between H i itself, Hatten, where Hatten represents the connection between features in space; superimpose the obtained relationship Hatten on the original H i to get H′ i = H i + Hatten, thereby obtaining H′0, H′1, and H′2 that strengthen the spatial relationship; In the fifth step, combine the 4-step memorability prediction components out_step of the three LSTM registers with 1024 channels obtained in the third step into Y. Y is 3×4-dimensional, and input three identical Ys into the second multi-head attention network to obtain the relationship Yatten between Y and itself. The relationship Yatten represents the connection between tasks in time; add the relationship Yatten to the original Y, that is, Y′ = Y + Yatten, to obtain Y′ with enhanced spatial relationship. Y′ is also 3×4-dimensional, thereby obtaining short-term memory y1, first intermediate memory y2, second intermediate memory y3, and long-term memory y4; Linearly stretch the long-term memory y4 and the short-term memory y1 to 3×1024 dimensions to obtain Ky4 and Ky1 respectively; input the stretched short-term memory Ky1 and H′0 into the third multi-head attention mechanism network to obtain the relationship XYatten1; input the stretched long-term memory Ky4 and H′2 into the fourth multi-head attention mechanism network to obtain the relationship XYatten2; add the relationship XYatten1 and the relationship XYatten2 to H′0 and H′2 respectively to obtain the final state H″0 of the surface visual feature information strengthened related to short-term memory and the final state H″2 of the high-level semantic feature information strengthened related to long-term memory; After recombining and processing the three spatial feature information H″0, H′1, H″2, through the first convolution for dimensionality reduction, the dropout() function, the relu() function, and the second convolution for dimensionality reduction, the final three memory degree prediction components p0, p1, p2 are obtained. Then, the final memory degree prediction value p = p0×a + p1×b + p2×c is obtained by linearly combining with three weight parameters.

Citation Information

Patent Citations

  • Image memorizing degree predication method based on sparse low-rank regression model

    CN107909091A

  • Deep attention mechanism-based image description generation method

    CN108052512A