Image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion

Through the multimodal emotional semantic adaptive fusion method, the multi-scale feature pyramid and cross-modal dynamic weights are used to solve the limitations of single-scale feature extraction in the existing technology, and a comprehensive and flexible aesthetic feature representation of image aesthetic evaluation is achieved.

CN120374621BActive Publication Date: 2025-09-02GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510864388.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-02
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The existing image aesthetic evaluation methods rely on visual feature extraction at a single scale, making it difficult to capture the local details and global structure of the image at the same time, and the fixed cross-modal interaction paradigm cannot adapt to the complex relationship between emotions and aesthetics in different scenarios.

Method used

The multimodal emotional semantic adaptive fusion method is adopted to extract multi-scale features of vision and text through the multi-scale feature pyramid structure, and the multi-head attention mechanism and gating mechanism are used to integrate the visual and text aesthetic features enhanced by emotional semantics, and integrate fusion features of different scales based on cross-modal dynamic weights.

Benefits of technology

A more comprehensive aesthetic feature representation of image aesthetic evaluation tasks is achieved, which improves the flexibility and accuracy of multimodal interaction in complex scenarios, and dynamically adapts to different emotions and aesthetic relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374621B_ABST
    Figure CN120374621B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for image aesthetic quality assessment based on multimodal emotional semantic adaptive fusion. The method comprises the following steps: obtaining an aesthetic image and corresponding comment text; extracting visual features at different levels from the aesthetic image to construct a multi-scale feature sequence of the aesthetic image; extracting semantic features at the word, phrase, and sentence levels from the comment text to construct a multi-scale feature sequence of the text; obtaining visual and textual aesthetic features enhanced with emotional semantics based on a multi-head attention mechanism, fusing these features to obtain fused features at different scales, constructing cross-modal dynamic weights, and integrating the fused features at different scales into an aesthetic representation based on a scale attention mechanism; predicting the aesthetic quality of the aesthetic image based on the aesthetic representation, and outputting an aesthetic evaluation prediction result. The present invention can effectively and dynamically utilize cross-modal emotional semantic information to improve the flexibility and accuracy of multimodal interaction in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image aesthetic quality evaluation, and in particular to an image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion. Background Art

[0002] Image aesthetic evaluation aims to enable computers to mimic human aesthetic perception and automatically judge the aesthetic quality of images. It has widespread applications in areas such as image enhancement, album organization, and image retrieval. Early image aesthetic evaluation methods focused on designing handcrafted features based on photographic rules, such as the rule of thirds, depth of field, and color harmony. In recent years, driven by the development of deep learning, researchers have employed convolutional neural networks to address the current mainstream image aesthetic evaluation tasks, namely aesthetic distribution prediction, aesthetic score regression, and aesthetic binary classification. As a uniquely human ability, aesthetic perception is highly abstract and complex, making image aesthetic evaluation research a challenging task.

[0003] Existing image aesthetic evaluation methods rely on single-scale visual feature extraction, making it difficult to simultaneously capture both local image details (such as texture and color) and global structure (such as composition and scene), resulting in limited representation of complex aesthetic elements. Furthermore, while researchers have begun to attempt to integrate additional auxiliary information, such as text modality and emotional semantics, to enhance the representation of image aesthetic features, their module designs rely on a fixed cross-modal interaction paradigm, making it difficult to dynamically adjust the fusion strategy based on the emotional semantics of the image-text pair, making it difficult to adapt to the complex relationship between emotion and aesthetics in different scenarios. Summary of the Invention

[0004] In order to overcome the defects and shortcomings of the existing technology, the present invention provides an image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion. The present invention utilizes the emotional information and comment text generated in the visual aesthetic process, and hierarchically extracts the multi-scale features of aesthetic images and texts through a multi-scale feature pyramid structure, effectively capturing visual semantics and emotional information from details to the overall situation, and generates emotion-guided multi-scale visual and text aesthetic features through intra-modal emotional semantic enhancement. The visual aesthetic features and text aesthetic features enhanced by emotional semantics are fused based on a gating mechanism, and the fused features of different scales are integrated into aesthetic representations based on cross-modal dynamic weights, thereby realizing hierarchical modeling of multi-scale visual features and dynamic allocation and efficient utilization of emotional information.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] The present invention provides an image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion, comprising the following steps:

[0007] Obtain aesthetic images and corresponding comment texts;

[0008] Extract visual features at different levels from aesthetic images and construct multi-scale feature sequences of aesthetic images;

[0009] Extract word-level, phrase-level, and sentence-level semantic features from the review text and construct a multi-scale feature sequence of the text;

[0010] The multi-scale feature sequence of the aesthetic image is processed by the image sentiment classifier to output the image sentiment features at each scale, and the visual aesthetic features with sentiment semantic enhancement are obtained based on the multi-head attention mechanism.

[0011] The multi-scale feature sequence of the text is processed by the text sentiment classifier to output the text sentiment features at each scale, and the text aesthetic features with sentiment semantic enhancement are obtained based on the multi-head attention mechanism;

[0012] Based on the gating mechanism, the visual aesthetic features enhanced by emotional semantics and the text aesthetic features are integrated to obtain fusion features of different scales.

[0013] Construct cross-modal dynamic weights and integrate fusion features of different scales into aesthetic representation based on the scale attention mechanism;

[0014] The aesthetic quality of the aesthetic image is predicted based on the aesthetic representation, and the aesthetic evaluation prediction result is output.

[0015] As a preferred technical solution, different levels of visual features are extracted from aesthetic images, specifically including:

[0016] Based on the ResNet50 network, feature maps of different levels are extracted, and a global average pooling operation is used to map them to a unified dimension through a fully connected layer to obtain visual features, which can be specifically expressed as follows:

[0017] ;

[0018] in, represents global average pooling, represents the fully connected layer, Represents visual features, Represents an aesthetic image.

[0019] As a preferred technical solution, semantic features at the word, phrase, and sentence levels are extracted from the review text, specifically including:

[0020] Based on the BERT pre-trained model, word-level semantic features are extracted from the comment text. The forward and backward hidden states are generated based on the bidirectional gated recurrent unit as the encoder. The average of the forward and backward hidden states is calculated to obtain phrase-level semantic features containing contextual semantics. All phrase-level features are averaged and pooled to obtain sentence-level semantic features.

[0021] As a preferred technical solution, we extract word-level semantic features from review text based on the BERT pre-trained model, and use a bidirectional gated recurrent unit as an encoder to generate forward and backward hidden states, including:

[0022] ;

[0023] ;

[0024] in, 、 denote the forward and backward hidden states respectively, represents the hidden state of the previous time step, represents the hidden state of the next time step, represents a forward gated recurrent unit, represents a backward gated recurrent unit, represents word-level features, Indicates the number of words in the comment text.

[0025] As a preferred technical solution, the multi-scale feature sequence of the aesthetic image is output through the image emotion classifier to output the image emotion features at each scale. Based on the multi-head attention mechanism, the visual aesthetic features with enhanced emotion semantics are obtained, specifically including:

[0026] For each scale, the image sentiment feature is used as the query and the visual feature is used as both the key and the value to calculate the first attention score:

[0027] ;

[0028] in, represents the first attention score, Represents the emotional characteristics of the image, Represents visual features, Representing visual features Dimensions, express function, 、 、 is a learnable parameter during training;

[0029] The first attention score is multiplied by the visual feature to obtain the weighted feature, which is then residually connected with the original visual feature. The layer normalization operation is performed to obtain the visual aesthetic feature with enhanced emotional semantics, which is specifically expressed as:

[0030] ;

[0031] in, Representation layer normalization operation, Visual aesthetic features representing emotional semantic enhancement.

[0032] As a preferred technical solution, the multi-scale feature sequence of the text is output by the text sentiment classifier at each scale. Based on the multi-head attention mechanism, the text aesthetic features with enhanced sentiment semantics are obtained, specifically including:

[0033] For each scale, the text sentiment feature is used as the query and the semantic feature is used as both the key and the value to calculate the second attention score:

[0034] ;

[0035] in, represents the second attention score, express function, Represents the sentiment characteristics of the text, Represents semantic features, 、 、 represents the learnable parameters during training, Representing semantic features Dimensions;

[0036] The second attention score is multiplied by the semantic feature to obtain the weighted feature, which is then residually connected with the original semantic feature. The text aesthetic feature enhanced with emotional semantics is obtained through layer normalization operation, which is specifically expressed as:

[0037] ;

[0038] in, Representation layer normalization operation, Text aesthetic features representing sentiment semantic enhancement.

[0039] As a preferred technical solution, the visual aesthetic features and text aesthetic features enhanced by emotional semantics are fused based on a gating mechanism to obtain fusion features of different scales, including:

[0040] The visual aesthetic features enhanced by emotional semantics and the text aesthetic features are spliced ​​together to obtain the spliced ​​feature vector, which is expressed as:

[0041] ;

[0042] in, represents the concatenated feature vector, Representation layer normalization operation, represents a multilayer perceptron, Represents a splicing operation, Visual aesthetic features that represent emotional semantic enhancement, Text aesthetic features that represent emotional semantic enhancement;

[0043] The concatenated feature vector Input the fully connected layer, map the output to the [0,1] interval based on the Sigmoid activation function, and obtain the gating signal , expressed as:

[0044] ;

[0045] in 、 are learnable parameters during training, represents the sigmoid function, Indicates in The degree of gate opening and closing under the scale;

[0046] Based on the gated signal The fusion features of different scales are obtained, which are expressed as:

[0047] ;

[0048] in, Indicates the Fusion features at different scales.

[0049] As a preferred technical solution, a cross-modal dynamic weight is constructed, which is expressed as:

[0050] ;

[0051] in, represents the cross-modal dynamic weight, express function, represents the fully connected layer, Indicates the Fusion features under different scales;

[0052] Based on the scale attention mechanism, the fusion features of different scales are integrated into the aesthetic representation, which is expressed as:

[0053] ;

[0054] in, Represents aesthetic representation, express activation function, 、 are learnable parameters during training, Represents a splicing operation.

[0055] As a preferred technical solution, the aesthetic quality of the aesthetic image is predicted based on the aesthetic representation, and the aesthetic evaluation prediction result is output, which specifically includes:

[0056] Mapping aesthetic features to aesthetic distribution predictions is expressed as:

[0057] ;

[0058] in, represents the aesthetic distribution prediction, express function, represents the fully connected layer, represents aesthetic representation;

[0059] Based on the aesthetic distribution prediction, the aesthetic score and aesthetic binary classification prediction results are obtained, which are expressed as:

[0060] ;

[0061] ;

[0062] in, represents the aesthetic score, represents the number of score labels, Represents the aesthetic binary classification prediction result, Represents the classification threshold.

[0063] As a preferred technical solution, the aesthetic quality of the aesthetic image is predicted based on the aesthetic representation, and the EMD loss is used as the objective function to calculate the distance between the predicted distribution and the true distribution, which is expressed as:

[0064] ;

[0065] ;

[0066] ;

[0067] in, represents the loss function, represents the number of score labels, The cumulative distribution function representing the true aesthetic distribution, represents the true aesthetic distribution, represents the predicted aesthetic distribution.

[0068] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0069] (1) This paper uses multimodal emotional semantic adaptive fusion to effectively and dynamically utilize cross-modal emotional semantic information, thereby obtaining a more comprehensive aesthetic feature representation for image aesthetic evaluation tasks.

[0070] (2) This paper uses a multi-scale feature pyramid structure to hierarchically extract multi-scale features of images and texts, effectively capturing visual semantics and emotional information from details to the global picture, solving the problem of insufficient representation capabilities of traditional single-scale features. It also generates emotion-guided multi-scale visual and text aesthetic features through intra-modal emotional semantic enhancement, further exploring the potential impact of emotional semantics on aesthetics.

[0071] (3) The present invention dynamically generates cross-modal fusion weights based on emotional semantics, integrates fusion features of different scales into aesthetic representations based on cross-modal dynamic weights, realizes adaptive information integration of image-text pairs, and improves the flexibility and accuracy of multimodal interaction in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 Schematic diagram of the overall implementation architecture of the image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion of the present invention;

[0073] Figure 2 Schematic diagram of the network architecture for extracting visual features of different levels of aesthetic images in the present invention;

[0074] Figure 3 This is a schematic diagram of the network architecture for using emotional information to perform emotional semantic enhancement on multi-scale features in the present invention;

[0075] Figure 4 Schematic diagram of the network architecture of the present invention that integrates fusion features of different scales into aesthetic representation based on cross-modal dynamic weights. DETAILED DESCRIPTION

[0076] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0077] like Figure 1 As shown, this embodiment provides an image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion, comprising the following steps:

[0078] First, an aesthetic image-text pair is taken as input (an aesthetic image and its corresponding commentary). Visual features at different levels (such as shallow texture features and deep semantic features) are extracted from the input aesthetic image to construct a multi-scale feature sequence of the aesthetic image. Semantic features at the word, phrase, and sentence levels are extracted from the commentary text corresponding to the aesthetic image to construct a multi-scale feature sequence of the text. On this basis, image and text sentiment features are extracted at each scale, and a multi-head attention mechanism is used to obtain visual and text aesthetic features with enhanced intra-modal sentiment semantics. Then, through cross-modal dynamic weight fusion based on a gating mechanism, the fusion weights of the image and text modalities are adaptively adjusted. The fused features at different scales are integrated into an aesthetic representation based on a scale attention mechanism to achieve fine-grained cross-modal interaction. Finally, the aesthetic representation is input into a fully connected layer and a softmax function to predict the aesthetic quality of the aesthetic image, and the final aesthetic evaluation prediction result is output. For the image aesthetic evaluation task, an aesthetic feature representation with enhanced cross-modal sentiment semantics is obtained.

[0079] Specifically, based on the image-text emotion dataset, this embodiment uses the ResNet network and the BERT model for pre-training to obtain the image emotion encoder pre-training model and text sentiment encoder pre-training model ,like Figure 1 As shown in the upper part, 、 Represents the image emotion encoder pre-training model and text emotion encoder pre-training model parameters, frozen parameters 、 , use it as a pre-trained image sentiment classifier and text sentiment classifier;

[0080] In this embodiment, given an input aesthetic image ,First, by applying ResNet50 pre-trained on ImageNet as the backbone network, ,the visual features are extracted from different convolutional layers of ResNet50, and ,multi-scale feature sequences of aesthetic images are constructed, such as Figure 2 As shown, the output of the shallow convolutional layer (Conv2 layer) is extracted as the first-level feature , which mainly contains local detail information of the image, such as texture, edge, etc.; the output of the middle convolutional layer (Conv3 layer) is extracted as the second-level feature , this feature contains certain semantic information, such as the local shape of the object; extract the output of the deep convolutional layer (Conv4 layer) as the third-level feature , which represents the global structural information of the image, such as scene layout, overall composition, etc. The feature maps of different levels extracted are subjected to global average pooling operation, and then mapped to a unified dimension through a fully connected layer (FC) to obtain visual features. , its formula is as follows:

[0081] ;

[0082] in, represents global average pooling;

[0083] In this embodiment, it is assumed that the comment text is entered have words, and use the BERT pre-training model to extract word-level features from the input comment text to obtain word-level features. , , and then use the bidirectional gated recurrent unit (Bidirectional GRU) as the encoder to generate phrase-level hidden states. The bidirectional GRU can simultaneously consider the contextual information of the text and better model the semantics of the phrase. Specifically expressed as:

[0084] ;

[0085] ;

[0086] in, and Represent the forward and backward hidden states respectively, and take the average of the two Used to represent phrase-level features containing contextual semantics, represents the hidden state of the previous time step (i.e., the i-1th position), which is used to capture the previous information of the sequence. Represents the hidden state of the next time step (i.e., the i+1th position), capturing the subsequent information, represents a forward gated recurrent unit, Represents the backward gated recurrent unit; then, all phrase-level features are averaged and pooled to obtain sentence-level features, and finally, a multi-scale feature sequence of the text is constructed. ,in is the word-level semantic feature, is the phrase-level semantic feature, are sentence-level semantic features. Subsequently, the dimensions of these features are unified to the same dimension as the image features through the fully connected layer.

[0087] like Figure 3As shown in the figure, in the image and text modalities, the emotional information is used to enhance the multi-scale features, thereby highlighting the semantic information related to emotions. Taking the visual modality as an example, the multi-scale feature sequence of the aesthetic image is transformed into Input them into the pre-trained image sentiment classifier respectively, and output the image sentiment features at each scale In order to explore the potential influence of emotional semantics on the aesthetic process, a multi-head attention mechanism is used to enhance the multi-scale visual features of the intra-modal emotion. Specifically, for each scale, the image emotional features are As a query, visual features As both the key and the value, the first attention score is calculated using the following formula:

[0088] ;

[0089] in, 、 、 are learnable parameters during training, Representing visual features Then, the first attention score is multiplied by the visual feature to obtain the weighted feature, which is then added to the original visual feature and residual connection is performed. Finally, the layer normalization operation is performed to obtain the emotionally enhanced visual aesthetic feature. , the formula is:

[0090] ;

[0091] in, Representation layer normalization operation;

[0092] Similarly, for text modality, in order to realize the enhancement effect of emotional semantics on multi-scale text features, the emotionally enhanced text aesthetic features are also obtained in the following way: :

[0093] ;

[0094] ;

[0095] in, Representing semantic features Dimensions;

[0096] like Figure 4As shown in the figure, in order to achieve the adaptive fusion of cross-modal emotional semantics of images and texts, a cross-modal dynamic weight fusion unit is designed to adjust the fusion strategy according to different emotional semantic information to adaptively utilize auxiliary information; the gating mechanism helps to focus on the most relevant text semantics and filter out redundant information. Specifically, the visual aesthetic features of the emotional enhancement are first and text aesthetic characteristics Perform splicing to obtain the spliced ​​feature vector , expressed as:

[0097] ;

[0098] Among them, the multi-layer perceptron MLP contains two layers of ReLU, Represents the splicing operation, the spliced ​​feature vector Input into a fully connected layer, use the Sigmoid activation function to map the output to the [0, 1] interval to obtain the gating signal :

[0099] ;

[0100] in , are learnable parameters during training, represents the sigmoid function, Indicates in The degree of gate opening and closing under the scale is used to control the transmission of relevant information. For highly relevant text semantics, A higher value means that more information is allowed to pass through, and the most important semantic information is utilized to a large extent. The lower the value, the more irrelevant information can be suppressed;

[0101] Based on the gated signal The fusion features of different scales are obtained, which are expressed as:

[0102] ;

[0103] in, Indicates the Based on the above operations, the fusion features under different scales not only dynamically and adaptively fuse cross-modal related emotional semantics, but also strengthen the cross-modal interaction between aesthetic features, which helps to learn more discriminative image aesthetic representations. Features at different levels contain aesthetic information of different granularity (such as shallow details, mid-level structures, and deep semantics). Directly using features at a single level may lead to information loss. Therefore, in order to retain the complementary information of details and semantics, and taking into account the different sensitivities of different images to different scales, the fusion features of different scales are dynamically integrated by using the scale attention mechanism. Integrate into a unified aesthetic representation :

[0104] ;

[0105] ;

[0106] Among them, cross-modal dynamic weight Dynamically assign the importance of features at different scales, The higher it is, the more important the scale feature is to the aesthetic judgment of the current image. 、 is a learnable parameter during training. Therefore, based on the above operations, the model can capture complex aesthetic attributes more comprehensively.

[0107] Finally, according to the aesthetic representation after fusion The aesthetic quality of the image is predicted and the aesthetic distribution is output. By applying a fully connected layer and softmax activation function, the aesthetic features are transformed into Mapping to aesthetic distribution predictions , the formula is:

[0108] ;

[0109] In this embodiment, the true distribution is expressed as ,in , Indicates the number of score labels. To optimize the model, Earth Mover's Distance (EMD) loss is used as the objective function. EMD loss measures the difference between two probability distributions. For aesthetic distribution prediction, it can better reflect the distance between the predicted distribution and the true distribution. The specific calculation formula is as follows:

[0110] ;

[0111] in, , Is the real aesthetic distribution and predicting aesthetic distribution By minimizing the EMD loss, the model's prediction results are made as close as possible to the true aesthetic distribution.

[0112] In the inference phase, the proposed model can directly output the predicted aesthetic distribution, based on which the aesthetic score can be inferred. And aesthetic binary classification prediction results Specifically, set a classification threshold In this embodiment, , according to the prediction score The comparison with the threshold is used to classify the image into “high aesthetic quality” (class 1) or “low aesthetic quality” (class 0), which is calculated as follows:

[0113] ;

[0114] ;

[0115] The present invention constructs a multi-scale feature pyramid to extract hierarchical features of images and texts, uses an intra-modal emotion enhancement module to strengthen the emotional semantic association between vision and text, and designs a cross-modal dynamic weight fusion unit to achieve adaptive information integration based on emotion. It effectively improves the representation ability and fusion flexibility of multimodal features, and its aesthetic evaluation performance in complex scenes is significantly better than that of existing technologies. It provides a solution for image aesthetic evaluation that is closer to the human aesthetic mechanism, which is impossible for previous methods to achieve.

[0116] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A method for image aesthetic quality evaluation based on multimodal emotional semantic adaptive fusion, characterized by: The steps include: Obtain aesthetic images and corresponding comment texts; Extract visual features at different levels from aesthetic images and construct multi-scale feature sequences of aesthetic images; Extract word-level, phrase-level, and sentence-level semantic features from the review text and construct a multi-scale feature sequence of the text; The multi-scale feature sequence of the aesthetic image is processed by the image sentiment classifier to output the image sentiment features at each scale. Based on the multi-head attention mechanism, the visual aesthetic features with enhanced sentiment semantics are obtained, specifically including: For each scale, the image sentiment feature is used as the query and the visual feature is used as both the key and the value to calculate the first attention score: ; in, represents the first attention score, Represents the emotional characteristics of the image, Represents visual features, Representing visual features Dimensions, express function, 、 、 is a learnable parameter during training; The first attention score is multiplied by the visual feature to obtain the weighted feature, which is then residually connected with the original visual feature. The layer normalization operation is performed to obtain the visual aesthetic feature with enhanced emotional semantics, which is specifically expressed as: ; in, Representation layer normalization operation, Visual aesthetic features that represent emotional semantic enhancement; The multi-scale feature sequence of the text is processed by the text sentiment classifier to output the text sentiment features at each scale. Based on the multi-head attention mechanism, the text aesthetic features with enhanced sentiment semantics are obtained, specifically including: For each scale, the text sentiment feature is used as the query and the semantic feature is used as both the key and the value to calculate the second attention score: ; in, represents the second attention score, express function, Represents the sentiment characteristics of the text, Represents semantic features, 、 、 represents the learnable parameters during training, Representing semantic features Dimensions; The second attention score is multiplied by the semantic feature to obtain the weighted feature, which is then residually connected with the original semantic feature. The text aesthetic feature enhanced with emotional semantics is obtained through layer normalization operation, which is specifically expressed as: ; in, Representation layer normalization operation, Text aesthetic features that represent emotional semantic enhancement; Based on the gating mechanism, the visual aesthetic features and text aesthetic features enhanced by emotional semantics are integrated to obtain fusion features of different scales, including: The visual aesthetic features enhanced by emotional semantics and the text aesthetic features are spliced ​​together to obtain the spliced ​​feature vector, which is expressed as: ; in, represents the concatenated feature vector, Representation layer normalization operation, represents a multilayer perceptron, Represents a splicing operation; The concatenated feature vector Input the fully connected layer, map the output to the [0, 1] interval based on the Sigmoid activation function, and obtain the gating signal , expressed as: ; in 、 are learnable parameters during training, represents the sigmoid function, Indicates in The degree of gate opening and closing under the scale; Based on the gated signal The fusion features of different scales are obtained, which are expressed as: ; in, Indicates the Fusion features under different scales; Construct cross-modal dynamic weights and integrate fusion features of different scales into aesthetic representation based on the scale attention mechanism; The aesthetic quality of the aesthetic image is predicted based on the aesthetic representation, and the aesthetic evaluation prediction result is output.

2. The image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion according to claim 1 is characterized in that: Extract different levels of visual features from aesthetic images, including: Based on the ResNet50 network, feature maps of different levels are extracted, and a global average pooling operation is used to map them to a unified dimension through a fully connected layer to obtain visual features, which can be specifically expressed as follows: ; in, represents global average pooling, represents the fully connected layer, Represents visual features, Represents an aesthetic image.

3. The image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion according to claim 1 is characterized in that: Extract word-level, phrase-level, and sentence-level semantic features from the review text, including: Based on the BERT pre-trained model, word-level semantic features are extracted from the comment text. The forward and backward hidden states are generated based on the bidirectional gated recurrent unit as the encoder. The average of the forward and backward hidden states is calculated to obtain phrase-level semantic features containing contextual semantics. All phrase-level features are averaged and pooled to obtain sentence-level semantic features.

4. The image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion according to claim 3 is characterized in that: The word-level semantic features of the review text are extracted based on the BERT pre-trained model, and the forward and backward hidden states are generated based on the bidirectional gated recurrent unit as the encoder, including: ; ; in, 、 denote the forward and backward hidden states respectively, represents the hidden state of the previous time step, represents the hidden state of the next time step, represents a forward gated recurrent unit, represents a backward gated recurrent unit, represents word-level features, Indicates the number of words in the comment text.

5. The image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion according to claim 1 is characterized in that: Construct cross-modal dynamic weights, expressed as: ; in, represents the cross-modal dynamic weight, express function, represents the fully connected layer, Indicates the Fusion features under different scales; Based on the scale attention mechanism, the fusion features of different scales are integrated into the aesthetic representation, which is expressed as: ; in, Represents aesthetic representation, express activation function, 、 are learnable parameters during training, Represents a splicing operation.

6. The image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion according to claim 1 is characterized in that: The aesthetic quality of the aesthetic image is predicted based on the aesthetic representation, and the aesthetic evaluation prediction results are output, specifically including: Mapping aesthetic features to aesthetic distribution predictions is expressed as: ; in, represents the aesthetic distribution prediction, express function, represents the fully connected layer, represents aesthetic representation; Based on the aesthetic distribution prediction, the aesthetic score and aesthetic binary classification prediction results are obtained, which are expressed as: ; ; in, represents the aesthetic score, represents the number of score labels, Represents the aesthetic binary classification prediction result, Represents the classification threshold.

7. The image aesthetic quality evaluation method based on multimodal emotional semantic adaptive fusion according to claim 1 is characterized in that: The aesthetic quality of the aesthetic image is predicted based on the aesthetic representation, and the EMD loss is used as the objective function to calculate the distance between the predicted distribution and the true distribution, which is expressed as: ; ; ; in, represents the loss function, represents the number of score labels, The cumulative distribution function representing the true aesthetic distribution, represents the true aesthetic distribution, represents the predicted aesthetic distribution.

Citation Information

Patent Citations

  • Multi-modal feature fusion method and system for sentiment analysis

    CN116644385A

  • Transform-based multi-modal aesthetic quality evaluation method

    CN117635964A