An Image Aesthetic Quality Evaluation Method Based on Brain-Inspired Multimodal Interaction Network
Through the image aesthetic quality evaluation method based on brain-inspired multimodal interactive network, we learn the association representation between images and text, and use SMF for feature fusion, which solves the problem of existing methods ignoring modal feature association and higher-order information, and achieves stronger aesthetic prediction and nonlinear expression capabilities.
Patent Information
- Application Number
- CN202310256229.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-03-16
AI Technical Summary
The existing multimodal image aesthetic quality evaluation method ignores the correlation information between modal features and the higher-order information of features, resulting in poor semantic interpretation of the fusion process.
A method of image aesthetic quality evaluation based on brain-inspired multimodal interaction network is proposed. By establishing a brain-inspired multimodal interaction network model including image and text perception module, recognition module and evaluation module, the interaction relationship between image perception features and text perception features is learned, and the feature fusion can be used to obtain aesthetic distribution.
By learning the correlation representation between images and text, modeling higher-order information of the modeling features improves the nonlinear expression ability of the aesthetic model, provides more powerful aesthetic predictions, and shows good generalization ability in emotional classification tasks.
Smart Images

Figure CN116152221B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of digital image processing and pattern recognition, and particularly relates to an image aesthetic quality evaluation method based on a brain-inspired multi-modal interaction network. Background Art
[0002] The widespread use of digital devices enables us to take a set of photos to capture momentary memories and preserve the most precious moments. With the rapid development of social media, a large number of photos are now shared via the Internet, such as on Facebook and Flickr. In the fields of art and photography, image aesthetics conveys beauty through images. The Image Aesthetic Assessment (IAA) task aims to enable computers to automatically predict the aesthetic level of images from an aesthetic perspective, just as they are imitating the ability of humans to perceive and understand beauty. IAA has a wide range of applications, such as image enhancement, automatic cropping, text generation, and photo retrieval. There are also some practical application requirements, such as intelligent advertising design (Alibaba's "Luban" AI designer), smartphone-guided photography (Samsung), etc. Therefore, in recent years, IAA has received increasing attention in the fields of computer vision and multimedia, attracting many researchers to explore this field.
[0003] Existing image aesthetic quality evaluation methods can generally be divided into two categories: Generalized Image Aesthetic Assessment (GIAA) and Personalized Image Aesthetic Assessment (PIAA). GIAA is more inclined to the aesthetic commonalities of the public and requires a large dataset to offset the inconsistencies of individuality. While PIAA tends to effectively describe the aesthetic perception results of different people for images, mainly using the prior knowledge obtained by GIAA for personalized aesthetic transfer learning to achieve a PIAA model for specific users. Aesthetic experience is a perceptual experience that evaluates and induces emotions and involves a process of understanding. This experience refers to a subjective feeling (such as pleasure or beauty) caused by coming into contact with things, as well as a judgment of like or attraction generated by them. The comments uploaded by users on the Internet are subjective feeling information given by users observing images and integrating their own memory information, thus triggering the emergence of multi-modal IAA methods that incorporate text input. Compared with single-modal IAA that uses only a single image as input, multi-modal IAA methods integrate the text modality, adding knowledge and thus improving the model performance. However, existing multi-modal IAA often ignores the correlation information between modal features and the high-order information of features, resulting in weak semantic interpretability in the fusion process. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes an image aesthetic quality evaluation method based on a brain-inspired multi-modal interaction network, including:
[0005] S1. Establish a brain-inspired multi-modal interaction network model;
[0006] The brain-inspired multimodal interaction network model includes: an image and text perception module, an identification module, and an evaluation module;
[0007] S2. Input the image data into the image perception module, and obtain the image perception features through the improved VGG16 backbone network pre-trained on ImageNet and a three-level convolutional network structure;
[0008] The improved VGG16 backbone network is the VGG16 backbone network obtained by deleting the classification layer in the original VGG16 network;
[0009] S3. Input the text data into the text perception module to extract the text perception features;
[0010] The text data is the subjective evaluation text of the image data;
[0011] S4. Through the identification module, learn the interaction relationship between the image perception features and the text perception features to obtain the correlation representation between the image and the text;
[0012] S5. The evaluation module uses the SMF that can be low-rank decomposed to fuse the image perception features, the text perception features, and the correlation representation between the image and the text. After fusion, power normalization and L2 regularization are performed to obtain the aesthetic distribution.
[0013] Preferably, obtaining the visual features of the image data includes:
[0014] The image data extracts 512-dimensional visual context features through the improved VGG16 backbone network. According to the extracted 512-dimensional visual context features, use a three-layer 1×1 convolutional network to extract hierarchical features. Perform mean processing on the width and height dimensions of each hierarchical feature to obtain the global context information of the image, and fuse the global context information of the image to obtain the final image perception features.
[0015] Preferably, obtaining the text perception features of the text data includes:
[0016] The text data is encoded through the 300-dimensional word embedding tool GloVe. Each word is encoded into a real-valued vector that captures the semantic characteristics between words. According to the real-valued vectors of the text data, use Bi-LSTM to extract text features. Use residual connection to add the real-valued vectors and the extracted text features, and perform layer normalization operation to obtain the abstract text context features. Use three types of filtering window heights (r∈{2, 3, 4}) to extract the context semantic information from abstract to specific, obtain features at three scales, and splice the features at the three scales to obtain the final text perception features.
[0017] Preferably, learning the interaction relationship between the image perception features and the text perception features to obtain the correlation representation between the image and the text includes:
[0018] Input the image perception features and text perception features into the recognition module. Through the KI-LSTM model that strengthens image memory, learn the interaction relationship between the perception features. In the KI-LSTM model, pay attention to the residual network structure of the image to avoid forgetting the image features. Then, complete the integration of the subjective evaluation text knowledge related to the image, and output the association representation between the image and the text.
[0019] Preferably, use the SMF that can be low-rank decomposed to fuse the image perception features, text perception features, and the association representation between the image and the text, including:
[0020]
[0021] Among them, represents the output after the SMF fuses the image perception features, text perception features, and the association representation between the image and the text. R represents R decomposition factors, and f text represents the text perception features, f cogn represents the association representation between the image and the text, f img represents the image perception features, w text,r 、w cogn,r 、w img,r respectively represent the weight coefficients of f text 、f cogn 、f img . represents the input vector, represents the outer product, represents the element-wise product.
[0022] Advantages of the present invention:
[0023] The present invention proposes to integrate the user's implicit memory by KI-LSTM to learn the association relationship representation between the image and the text, model the high-order information of the features, and at the same time improve the non-linear expression ability of the aesthetics model;
[0024] The present invention proposes a general SMF to fuse multi-modal features to utilize the complementarity of heterogeneous data and provide more powerful aesthetics prediction. The SMF uses the low-rank matrix decomposition method to reduce the parameters, and the increase in the number of parameters is only linearly related to the number of modalities;
[0025] The present invention can be used for downstream task sentiment classification, and it is verified that the present invention has good generalization ability on the evaluation indexes of this task. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a flowchart of an image aesthetic quality evaluation method based on a brain-inspired multi-modal interaction network of the present invention;
[0027] Figure 2 Schematic diagram of the process for the KI-LSTM of the present invention to learn the correlation relationship between images and texts;
[0028] Figure 3 Schematic diagram of the process for the SMF of the present invention to fuse multimodal features. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] An image aesthetic quality evaluation method based on a brain-inspired multimodal interaction network, as Figure 1 shown, includes:
[0031] S1. Establish a brain-inspired multimodal interaction network model;
[0032] The brain-inspired multimodal interaction network model includes an image and text perception module, a recognition module, and an evaluation module;
[0033] S2. Input the image data into the image perception module, and obtain image perception features through an improved VGG16 backbone network pre-trained on ImageNet and a three-level convolutional network structure;
[0034] The improved VGG16 backbone network is the VGG16 backbone network obtained by deleting the classification layer in the original VGG16 network;
[0035] S3. Input the text data into the text perception module to extract text perception features;
[0036] The text data is the subjective evaluation text of the image data;
[0037] S4. Through the recognition module, learn the interaction relationship between the image perception features and the text perception features to obtain the correlation representation between the image and the text;
[0038] S5. The evaluation module uses the SMF that can be low-rank decomposed to fuse the image perception features, the text perception features, and the correlation representation between the image and the text. After fusion, power normalization and L2 regularization are performed to obtain the aesthetic distribution.
[0039] Obtain the visual features of the image data, including:
[0040] The image data extracts 512-dimensional visual context features through an improved VGG16 backbone network, extracts hierarchical features using a 3-layer 1×1 convolutional network based on the extracted 512-dimensional visual context features, performs mean processing on the width and height dimensions of each level of features to obtain the global context information of the image, and fuses the global context information of the image to obtain the final image perception features.
[0041] Extract 512-dimensional visual context features V through an improved VGG16 backbone network I , including:
[0042] V I = VGG16(I)
[0043] where I represents the input image.
[0044] Extract hierarchical features using a 3-layer 1×1 convolutional network based on the extracted 512-dimensional visual context features including:
[0045]
[0046] where V I represents the visual context features, ReLU(·) represents the activation function, Avgpool(·) represents the average pooling operation, and Conv(·) represents the convolutional operation.
[0047] Obtain the text perception features of the text data, including:
[0048] The text data is encoded by the 300-dimensional word embedding tool GloVe, and each word is encoded into a real-valued vector that captures the semantic characteristics between words. The text features are extracted from the real-valued vectors of the text data using Bi-LSTM. The residual connection is used to add the real-valued vectors and the extracted text features, and the abstract text context features are obtained through layer normalization operations. Filtering windows of three heights (r ∈ {2, 3, 4}) are used to extract the context semantic information from the abstract to the specific, and features at three scales are obtained. The features at the three scales are concatenated to obtain the final text perception features.
[0049] Extract text features using Bi-LSTM, and the formula is as follows:
[0050]
[0051] where H t represents the text features extracted from the input word embedding through Bi-LSTM, and represent the forward and backward hidden states at time t, respectively.
[0052] and Obtained respectively by the following formulas:
[0053]
[0054]
[0055] where x t represents encoding each word from GloVe into a real-valued vector that captures some semantic characteristics between words, and represent the forward and backward LSTMs respectively.
[0056] To reduce the problem of vanishing gradients, a residual connection is used to add the input and output in the network. Then, the text context feature V T is obtained through layer normalization operation and is defined as follows:
[0057]
[0058] where x t represents encoding each word from GloVe into a real-valued vector that captures the semantic characteristics between words, and H t represents the text features extracted from the input word embeddings by the Bi-LSTM, Norm(·) represents the regularization operation, represents addition.
[0059] Finally, three heights of filter windows (r ∈ {2, 3, 4}) are used to extract the context semantic information from abstract to concrete, which is defined as follows:
[0060]
[0061] where represents the context semantic feature of the text, W t r represents the weight, V T represents the text context feature, b represents the bias, and ReLU(·) represents the activation function.
[0062] Learn the interaction relationship between the image perception feature and the text perception feature to obtain the correlation representation between the image and the text, including:
[0063] As Figure 2 shown, input the image perception feature and the text perception feature into the recognition module, learn the interaction relationship between the perception features through the KI-LSTM model that strengthens the image memory, pay attention to the residual network structure of the image in the KI-LSTM model to avoid forgetting the image features, and then complete the integration of the subjective evaluation text knowledge related to the image, and output the correlation representation between the image and the text.
[0064] The KI-LSTM model that enhances image memory learns the interaction relationships between modalities, which is defined as follows:
[0065]
[0066]
[0067]
[0068]
[0069] Among them, represents the input of text features at time t, represents the hidden layer state at time t - 1, represents the output at time t - 1, respectively represent the weight coefficients in the forget gate, input gate, and output gate among them, respectively represent the weight coefficients in the forget gate, input gate, and output gate among them, respectively represent the weight coefficients in the forget gate, input gate, and output gate among them, respectively represent the biases of the forget gate, input gate, and output gate, represents the output at the current time t, ⊙ represents the Hadamard product in the formula, and * represents the convolution symbol.
[0070] Finally, the semantic information of the subjective evaluation text of the image is extracted.
[0071] Then, in the KI-LSTM, pay attention to the residual network structure of the image to avoid forgetting image features, and the formula is defined as follows:
[0072]
[0073]
[0074]
[0075]
[0076] Among them, represents the input of image features at time t, respectively represent the weight coefficients in the forget gate, input gate, and output gate among them, represents the output at the current time t.
[0077] After that, the integration of the subjective evaluation text knowledge related to the image is completed, which is expressed as follows:
[0078]
[0079]
[0080] Among them, represents the integration result of subjective evaluation text knowledge related to images, respectively represent the weight coefficients of represents the bias, represents the hidden state at the current time t, represents and the weight coefficient connecting to
[0081] For example, Figure 3 As shown, the evaluation module uses the SMF that can be low-rank decomposed to fuse the image modality, text modality, and the interactive modality that integrates user memory information, and outputs the aesthetic distribution.
[0082] A common tensor representation for multimodal fusion in this embodiment is expressed as:
[0083]
[0084] Among them, and b ∈ d o respectively represent the weight and the bias. is calculated from the unimodal representation, where is a set of tensor outer products on the vectors indexed by m, represents the unimodal input with 1 appended. Using low-rank decomposition, it can be simplified to:
[0085]
[0086] Among them, f img and f text are two modal features extracted by the image perception module and the text perception module respectively. f cogn As the third modality is the feature of the interaction between the image and the text extracted by the cognitive module. Using and parallel decomposition simplifies the calculation. This process decouples different modalities, so the proposed SMF can be easily extended to any number of modalities.
[0087] In this embodiment, a new multimodal fusion method is proposed: The three modalities of fusing images, texts, and interactive correlation representations are expressed by the formula as follows:
[0088]
[0089] Among them, R represents R decomposition factors. f text represents the text perception feature, f cogn represents the association representation between the image and the text, f img represents the image perception feature, w text,r 、w cogn,r 、w img,r respectively represent the weight coefficients of f text 、f cogn 、f img . represents the input vector, represents the output of the SMF. represents the outer product, represents the element-wise product.
[0090] f cogn As the third modality, it is the interaction feature between the image and the text extracted by the cognitive module. By using and parallel decomposition, the calculation of is simplified. After the SMF output, power normalization and L2 regularization are performed, and finally the aesthetic distribution is obtained.
[0091] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An image aesthetic quality evaluation method based on a brain-inspired multimodal interaction network, characterized in that, it includes: S1. Establish a brain-inspired multimodal interaction network model; The brain-inspired multimodal interaction network model includes an image and text perception module, a recognition module, and an evaluation module; S2. Input the image data into the image perception module, and obtain image perception features through an improved VGG16 backbone network pre-trained on ImageNet and a three-layer convolutional network structure; The improved VGG16 backbone network is the VGG16 backbone network obtained by deleting the classification layer in the original VGG16 network; S3. Input the text data into the text perception module to extract text perception features; The text data is the subjective evaluation text of the image data; S4. Through the recognition module, learn the interaction relationship between the image perception features and the text perception features to obtain the correlation representation between the image and the text; Input the image perception features and the text perception features into the recognition module, learn the interaction relationship between the perception features through the KI-LSTM model that strengthens image memory, and pay attention to the residual network structure of the image in the KI-LSTM model to avoid forgetting the image features. After that, complete the integration of the subjective evaluation text knowledge related to the image, and output the correlation representation between the image and the text; S5. The evaluation module uses the SMF that can be low-rank decomposed to fuse the image perception features, the text perception features, and the correlation representation between the image and the text. After fusion, perform power normalization and L2 regularization to obtain the aesthetic distribution; The evaluation module uses the SMF that can be low-rank decomposed to fuse the image perception features, the text perception features, and the correlation representation between the image and the text, including: Among them, represents the output after fusing the image perception feature, text perception feature, and the association representation between the image and the text. R represents R decomposition factors, and f text represents the text perception feature, and f cogn represents the association representation between the image and the text, and f img represents the image perception feature, and w text,r 、w cogn,r 、w img,r respectively represent the weight coefficients of f text 、f cogn 、f img . represents the input vector, represents the outer product, represents the element-wise product.
2. The image aesthetic quality evaluation method based on a brain-inspired multimodal interaction network according to claim 1, characterized in that, obtaining the visual features of the image data, including: The image data extracts 512-dimensional visual context features through the improved VGG16 backbone network, extracts hierarchical features using a three-layer 1×1 convolutional network according to the extracted 512-dimensional visual context features, performs mean processing on the width and height dimensions of each hierarchical feature, obtains the global context information of the image, and fuses the global context information of the image to obtain the final image perception features.
3. The image aesthetic quality evaluation method based on a brain-inspired multimodal interaction network according to claim 1, characterized in that, obtaining the text perception features of the text data, including: The text data is encoded by the 300-dimensional word embedding tool GloVe, and each word is encoded into a real-valued vector that captures the semantic characteristics between words. According to the real-valued vectors of the text data, use Bi-LSTM to extract text features, add the real-valued vectors and the extracted text features using residual connections, and perform layer normalization operations to obtain the abstract text context features. Use three heights of filtering windows (r∈{2, 3, 4}) to extract the context semantic information from abstract to specific, obtain features at three scales, and splice the features at the three scales to obtain the final text perception features.