Visual emotion analysis method and system based on attribute guidance, terminal and storage medium

Through the alignment and association analysis of visual and text attribute representations, an attribute emotion map is constructed, which solves the problem of unanalyzed attribute associations in the existing technology, and achieves more accurate sentiment analysis.

CN120236152AActive Publication Date: 2025-07-01SHENZHEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510718555.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

The prior art does not use a dedicated attribute network to obtain attribute representations in visual sentiment analysis, and the correlation between attributes is not analyzed, resulting in inaccurate sentiment analysis results.

Method used

By obtaining the emotional image to be analyzed by the target user, visual and text attribute representation extraction, alignment processing and association analysis are performed, attribute emotion map is constructed for emotional prediction, and visual attribute representation extraction is optimized by using multi-level attribute expert module and text-guided multi-level alignment module, dynamically adjusting attribute representation weights, and constructing attribute emotion maps for emotional reasoning.

Benefits of technology

The accuracy and interpretability of emotion analysis results are improved, and the interpretability of emotion recognition is enhanced by analyzing the multiple interactions between emotions and attributes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236152A_ABST
    Figure CN120236152A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, and discloses a visual sentiment analysis method and system based on attribute guidance, a terminal and a storage medium, and the method comprises the steps: obtaining a to-be-analyzed sentiment image of a target user, carrying out the representation extraction to obtain a plurality of visual attribute representations, and carrying out the representation to obtain a plurality of text attribute representations; performing alignment processing on all the visual attribute representations and the text attribute representations to obtain a plurality of target visual attribute representations, and performing attribute association analysis to obtain attribute association information; and constructing an attribute emotion map, obtaining a target attribute emotion map after optimization, and performing emotion prediction according to the target attribute emotion map to obtain an emotion prediction result. According to the method, extraction of the visual characterization is optimized by guiding the multiple visual attribute characterization under the text attribute characterization, the weight of each attribute characterization is dynamically adjusted through attribute association analysis, the attribute emotion graph is constructed to deeply analyze the relation between the emotion and the attribute, and the accuracy of the emotion analysis result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular, to a visual emotion analysis method, system, terminal and computer-readable storage medium based on attribute guidance. Background Art

[0002] Emotion is a unique and indispensable feature of humans, permeating all aspects of daily life. With the rise of social media, more and more people express their emotional experiences by sharing pictures. Therefore, visual emotion analysis has received extensive attention. Its goal is to deeply understand an individual's emotional response to various visual stimuli and provide explanations for these predictions. In the research of visual emotion analysis, an inherent challenge is called the "emotion gap". Emotion is a complex reaction of the brain to external stimuli and internal states, while images contain different forms of visual elements. Although the semantic information at the pixel level is relatively easy to identify, the perception of highly complex emotions is more challenging. This so-called emotion gap mainly stems from the representational differences between visual elements and human emotional experiences.

[0003] Although deep learning methods have made significant progress in visual representation extraction, they ignore the subtle differences in emotional expression. For example, low-level visual representations such as color, lighting, and composition can provide certain emotional cues. To solve this problem, existing technologies narrow this emotion gap through multimodal learning, attention mechanisms, and large-scale data-driven methods. For example, using language descriptions to enhance emotional semantic understanding, or constructing more emotion-perceptive visual representations through generative models. However, existing technologies use neural networks to extract global representational information, without mentioning using a dedicated attribute network to obtain attribute representations, but obtaining corresponding attribute representations through the proposed network; secondly, in emotion analysis, the associations between attributes are not analyzed, and emotion reasoning analysis is directly carried out, and the classification prediction of emotion categories is directly performed on the representations, resulting in inaccurate emotion analysis results.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a visual emotion analysis method, system, terminal and storage medium based on attribute guidance, aiming to solve the problem that existing technologies do not use a dedicated attribute network to obtain specific attribute representations in emotion analysis, do not analyze the associations between attributes, and perform emotion reasoning analysis, resulting in inaccurate emotion analysis results.

[0006] To achieve the above purpose, the present invention provides a visual emotion analysis method based on attribute guidance. The visual emotion analysis method based on attribute guidance includes the following steps: Obtain the emotional image to be analyzed of the target user, perform feature extraction on the emotional image to be analyzed to obtain multiple visual attribute features, and perform text feature extraction on the emotional image to be analyzed to obtain multiple text attribute features; Align all the visual attribute features and all the text attribute features to obtain multiple target visual attribute features, and perform attribute correlation analysis on all the target visual attribute features to obtain attribute correlation information; Construct an attribute emotional graph according to the attribute correlation information, perform optimization processing on the attribute emotional graph to obtain a target attribute emotional graph, and perform emotional prediction according to the target attribute emotional graph to obtain an emotional prediction result.

[0007] Optionally, in the above-mentioned attribute-guided visual emotion analysis method, the step of obtaining the emotional image to be analyzed of the target user, performing feature extraction on the emotional image to be analyzed to obtain multiple visual attribute features, and performing text feature calculation on the emotional image to be analyzed to obtain multiple text attribute features specifically includes: Obtain the emotional image to be analyzed of the target user and construct a multi-level attribute expert module; Perform feature extraction on the emotional image to be analyzed through the multi-level attribute expert module to obtain multi-level attribute features, where the multi-level attribute features include low-level attribute features, intermediate-level attribute features, and high-level attribute features; Perform non-linear transformation on the multi-level attribute features to obtain multiple visual attribute features; Perform text feature extraction on the emotional image to be analyzed to obtain multiple text attribute features.

[0008] Optionally, in the above-mentioned attribute-guided visual emotion analysis method, the specific operation of performing non-linear transformation on the multi-level attribute features is: ; The specific operation of performing text feature extraction on the emotional image to be analyzed is: ; Among them, is the visual attribute feature set, is the activation function, is the low-level attribute feature, is the intermediate-level attribute feature, is the high-level attribute feature, and are both learnable parameters, is the number of attribute modules, is the text attribute feature set, is the text adapter, is a text encoder, is an attribute embedding encoding.

[0009] Optionally, in the attribute-guided visual sentiment analysis method, the process of aligning all the visual attribute representations and all the text attribute representations to obtain multiple target visual attribute representations specifically includes: Performing scale assimilation on all the visual attribute representations and all the text attribute representations, and aligning all the visually and textually attribute representations after scale assimilation to obtain multiple aligned visual attribute representations; Calculating a loss for all the visually and textually attribute representations after scale assimilation to obtain an attribute alignment loss, and optimizing all the aligned visual attribute representations according to the attribute alignment loss to obtain multiple target visual attribute representations.

[0010] Optionally, in the attribute-guided visual sentiment analysis method, the process of performing attribute correlation analysis on all the target visual attribute representations to obtain attribute correlation information specifically includes: Calculating scores for the attribute correlations of all the target visual attribute representations to obtain multiple correlation scores; Performing attribute correlation analysis on all the attributes according to all the correlation scores to obtain attribute correlation information.

[0011] Optionally, in the attribute-guided visual sentiment analysis method, the process of constructing an attribute sentiment map according to the attribute correlation information, optimizing the attribute sentiment map to obtain a target attribute sentiment map, and performing sentiment prediction according to the target attribute sentiment map to obtain a sentiment prediction result specifically includes: Performing spatial mapping on the attribute correlation information to obtain an attribute mapping result, and constructing an attribute sentiment map according to the attribute mapping result; Calculating a sentiment classification loss for all the visual attribute representations, and calculating a total loss for the sentiment classification loss and the attribute alignment loss to obtain a target total loss; Optimizing the attribute sentiment map according to the target total loss to obtain a target attribute sentiment map, and performing sentiment prediction according to the target attribute sentiment map to obtain a sentiment prediction result, where the sentiment prediction result includes sentiment information and a sentiment explanation.

[0012] Optionally, in the attribute-guided visual sentiment analysis method, the calculation of the loss for all the visual attribute representations specifically is: ; The total loss calculation for the emotion classification loss and the attribute alignment loss is specifically as follows: ; Among them, is the emotion classification loss, is the total number of emotion categories, is the number of attribute modules, is the th true label, is the non - linear mapping, F is the set composed of visual attribute representations, is the correlation analysis between attributes, is the number of emotion categories, is the target total loss, is the hyperparameter, is the attribute alignment loss.

[0013] Optionally, in the above - mentioned attribute - guided visual emotion analysis method, the attribute - guided visual emotion analysis system includes: A representation processing module, configured to obtain the emotion image to be analyzed of the target user, perform representation extraction on the emotion image to be analyzed to obtain multiple visual attribute representations, and perform text representation extraction on the emotion image to be analyzed to obtain multiple text attribute representations; An attribute analysis module, configured to align all the visual attribute representations and all the text attribute representations to obtain multiple target visual attribute representations, and perform attribute correlation analysis on all the target visual attribute representations to obtain attribute correlation information; An emotion prediction module, configured to construct an attribute - emotion graph based on the attribute correlation information, perform optimization processing on the attribute - emotion graph to obtain a target attribute - emotion graph, and perform emotion prediction based on the target attribute - emotion graph to obtain an emotion prediction result.

[0014] In addition, to achieve the above object, the present invention also provides a terminal, where the terminal includes: a memory, a processor, and an attribute - guided visual emotion analysis program stored on the memory and executable on the processor. When the attribute - guided visual emotion analysis program is executed by the processor, the steps of the above - mentioned attribute - guided visual emotion analysis method are implemented.

[0015] In addition, to achieve the above object, the present invention also provides a computer - readable storage medium, where the computer - readable storage medium stores an attribute - guided visual emotion analysis program. When the attribute - guided visual emotion analysis program is executed by a processor, the steps of the above - mentioned attribute - guided visual emotion analysis method are implemented.

[0016] In the present invention, a to-be-analyzed emotional image of a target user is obtained, characterization extraction is performed on the to-be-analyzed emotional image to obtain a plurality of visual attribute characterizations, and text characterization extraction is performed on the to-be-analyzed emotional image to obtain a plurality of text attribute characterizations; all the visual attribute characterizations and all the text attribute characterizations are subjected to alignment processing to obtain a plurality of target visual attribute characterizations, and attribute correlation analysis is performed on all the target visual attribute characterizations to obtain attribute correlation information; an attribute emotional map is constructed according to the attribute correlation information, the attribute emotional map is optimized to obtain a target attribute emotional map, and emotional prediction is performed according to the target attribute emotional map to obtain an emotional prediction result. The present invention adopts a multi-level attribute expert-level feature extraction method, introduces a text-guided multi-level alignment module, and at the same time optimizes the extraction of multiple visual attribute characterizations under the guidance of text attribute characterizations; attribute correlation analysis is proposed in the emotional reasoning task to analyze all attributes and dynamically adjust the weight of each attribute characterization, and the attribute correlation graph is mapped to the emotional space for emotional reasoning. By constructing an attribute emotional map, the relationship between emotions and attributes can be deeply analyzed, not only considering single characterizations, but also integrating the interaction between multiple attributes, enhancing the interpretability of emotion recognition, and thus improving the accuracy of emotional analysis results. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic diagram of an existing visual emotion analysis method and a preferred embodiment of the present invention; Figure 2 is a flowchart of a preferred embodiment of the visual emotion analysis method based on attribute guidance of the present invention; Figure 3 is an overall schematic diagram of the visual emotion analysis method based on attribute guidance of the present invention; Figure 4 is a schematic diagram of a single-dimensional attribute correlation module in a preferred embodiment of the present invention; Figure 5 is a structural diagram of a preferred embodiment of the visual emotion analysis system based on attribute guidance of the present invention; Figure 6 is a schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0019] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship, movement conditions, etc. between components in a specific posture (as shown in the attached drawings). If the specific posture changes, the directional indications will also change accordingly.

[0020] In addition, if there are descriptions such as "first", "second", etc. involved in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0021] The prior art constructs visual elements based on art and psychological principles and attempts to analyze the relationship with human emotions through the constructed visual elements. However, due to the complexity of emotions and the diversity of visual elements that can prompt emotions, it is difficult for artificially constructed visual elements to cover all and may cause errors in real-world recognition. With the continuous development of deep learning networks, convolutional neural networks are superior in tasks such as classification and use neural networks to extract visual representations from images for emotion prediction. Convolutional neural networks can effectively extract global visual representations from images and perform emotion classification. As shown in (a) of Figure 1 , relying only on global information for emotion prediction may ignore important emotion clues embedded in local representations. Therefore, local information is also considered crucial in emotion prediction. To improve the accuracy of prediction, local representation extraction is combined with global information, as shown in (b) of Figure 1 . However, these methods still have limitations in visual emotion prediction because they oversimplify the complexity of emotion arousal. Simply combining global and local representations is still not sufficient to establish a meaningful connection with emotions. To solve this problem, the present invention proposes attribute-guided visual emotion analysis, which analyzes the relationship between different-level attribute representations and constructs an emotion attribute graph for reasoning, as shown in (c) of Figure 1 .

[0022] The attribute-guided visual emotion analysis method described in the preferred embodiment of the present invention, as shown in Figure 2 , the attribute-guided visual emotion analysis method includes the following steps: Step S10: Obtain the emotional image to be analyzed of the target user, perform feature extraction on the emotional image to be analyzed to obtain multiple visual attribute features, and perform text feature extraction on the emotional image to be analyzed to obtain multiple text attribute features.

[0023] Specifically, in the embodiments of the present invention, as Figure 3 shown, a visual emotional analysis method based on attribute guidance is proposed, which is used to analyze the relationship between attribute features in an image and explain the influence of attributes on emotions; a multi-level attribute guidance module, which uses multiple expert networks to capture visual features related to attributes based on text feature guidance, and maps these features to a latent feature space through an adaptive layer and a non-linear mapping, so as to map the image and text features to the same space and extract accurate visual attribute features; secondly, a correlation reasoning module analyzes the relationship between attributes; finally, an emotional reasoning module constructs an attribute emotion map using these attribute features, which can explain the importance change of different attributes for emotions in the emotional reasoning process. The specific analysis process is that in a deep neural network, as the network depth increases, its processing of features gradually biases towards semantic-level information. Therefore, the present invention proposes a multi-level extraction network, where attributes at different levels correspond to different parts of the network, and the higher the attribute level, the deeper the corresponding network depth. In addition, how to accurately extract visual attribute features remains a key challenge. The present invention uses a proprietary sub-module, that is, each type of attribute has a corresponding network module. For example, two low-level attributes include colorfulness and brightness, two intermediate-level attributes include objects and scenes, and two high-level attributes include facial expressions and human actions.

[0024] Obtain the emotional image to be analyzed of the target user, denoted by ; and construct a multi-level attribute expert module; perform feature extraction on the emotional image to be analyzed through the multi-level attribute expert module according to the feature extraction formula to obtain multi-level attribute features, where the multi-level attribute features include low-level attribute features, intermediate-level attribute features, and high-level attribute features; the feature extraction formula is: ; where is the low-level attribute feature, is the intermediate-level attribute feature, is the high-level attribute feature, is an image encoder. In the embodiments of the present invention, the proprietary sub-module of each attribute is responsible for extracting the corresponding attribute representation. Since there are differences in multiple attribute representations, an attribute space is constructed for each attribute. Secondly, different attribute-specific representations are also different, and a more fine-grained representation of each attribute can be learned. The sub-module uses a multi-layer perceptron to extract the unique fine-grained representation of each attribute and performs a non-linear transformation, that is, performs a non-linear transformation on the multi-level attribute representation to obtain multiple visual attribute representations. The specific operation of performing a non-linear transformation on the multi-level attribute representation is as follows: ; wherein, is a set of visual attribute representations, is an activation function, and are both learnable parameters, is the number of attribute modules, is a set composed of visual attribute representations, including six attribute representations, wherein, are respectively 6 visual attribute representations, namely brightness, chromaticity, scene type, object category, facial expression, and human action.

[0025] After that, CLIP (Contrastive Language-Image Pre-training) is a powerful language-image contrast learning model that can establish an association between text and semantic representations in the CLIP space. In the present invention, the CLIP text encoder is used to encode the labels of six attributes to obtain accurate visual representations under the guidance of the text representations of the attributes. However, in order to enable the text representation to match the visual representation in the attribute space, a text adapter is proposed to perform text representation extraction, that is, perform text representation extraction on the to-be-analyzed emotional image to obtain multiple text attribute representations. The specific operation of performing text representation extraction on the to-be-analyzed emotional image is as follows: ; wherein, is a set of text attribute representations, is a text adapter, is a text encoder, is an attribute embedding encoding. wherein, are respectively 6 text attribute representations, namely brightness, chromaticity, scene type, object category, facial expression, and human action.

[0026] Step S20: Align all the visual attribute representations and all the text attribute representations to obtain a target representation alignment result, and perform attribute correlation analysis based on the target representation alignment result to obtain attribute correlation information.

[0027] Specifically, after obtaining the visual attribute representations and text attribute representations, it is necessary to ensure that the visual attribute representations and text attribute representations are on the same scale to ensure the reliability and accuracy of alignment. That is, perform scale assimilation on all the visual attribute representations and all the text attribute representations to obtain multiple scale-assimilated visual attribute representations (denoted by ), and multiple scale-assimilated text attribute representations (denoted by ). Multi-level text-guided alignment can achieve visual and text representation alignment at different levels, thereby enhancing the accuracy of visual representations; and align all the scale-assimilated visual attribute representations and text attribute representations to obtain multiple aligned visual attribute representations. Under the guidance of text representations, an attribute loss is formulated to ensure the accuracy and diversity of the visual attribute representations extracted from the image. That is, perform loss calculation on all the scale-assimilated visual attribute representations and text attribute representations to obtain an attribute alignment loss, denoted by , and the corresponding expression is: ; where, is the total number of attribute categories, is the total number of attributes in a batch, is the th visual attribute feature, is the th text attribute feature, is the similarity calculation. Among them, the formula for similarity calculation is: ; As the loss of attributes converges, the visual semantic representations will become more and more accurate. After that, optimize all the aligned visual attribute representations according to the attribute alignment loss to obtain multiple target visual attribute representations, that is, multiple accurate visual attribute representations.

[0028] After that, the attribute correlation analysis can analyze all attributes, dynamically adjust the weights represented by each attribute, amplify important attribute categories, and weaken irrelevant attribute categories. The correlation between attributes can reveal potential emotion prediction to help the model better interpret complex emotion expressions in images. For example, subtle changes among color, brightness, and facial expressions can change the expressed emotion; while the attribute correlation module analyzes the relationship between attributes by calculating the correlation scores between attributes. Specifically, it calculates the attribute correlation scores for all the target visual attribute representations to obtain multiple correlation scores; and performs attribute correlation analysis on all attributes based on all the correlation scores to obtain attribute correlation information. As Figure 4 shown, where "brightness" is associated with other attributes, self-attention can dynamically adjust the association between embeddings and automatically update the importance among the embeddings. The multi-head attention mechanism obtains the relationship between attributes from multiple dimensions to form a relationship graph of attributes. Among them, Figure 4 in 、 、 、 、 、 and are all visual attribute representations, , , , , , , , , and are all weight matrices, X is matrix multiplication, + is matrix addition, and softmax is the normalized exponential function.

[0029] Step S30: Construct an attribute emotion graph based on the attribute correlation information, perform optimization processing on the attribute emotion graph to obtain a target attribute emotion graph, and perform emotion prediction based on the target attribute emotion graph to obtain an emotion prediction result.

[0030] Specifically, in the emotion reasoning task, it not only involves emotion recognition, but also requires the model to understand the causes of emotions and their important triggering factors. Emotions are high-dimensional and complex representations. The attribute correlation information is mapped to the emotion space, so that in the emotion space, an attribute-emotion map is constructed. The constructed emotion map is the key factor of the attributes of highly refined emotions. The knowledge of the emotion map can be applied to any field and data analysis of visual emotions, which is beneficial to the understanding and analysis of emotions. Nonlinear mapping can map the low-level attribute correlation information to the high-dimensional emotion space to realize the construction of the attribute-emotion map. Specifically, the attribute correlation information is subjected to spatial mapping to obtain an attribute mapping result, and an attribute-emotion map is constructed according to the attribute mapping result. Then, loss calculation is performed on all the visual attribute representations to obtain an emotion classification loss. Among them, the loss calculation for all the visual attribute representations is specifically as follows: ; Among them, is the emotion classification loss, is the total number of emotion categories, is the number of attribute modules, is the th true label, is the nonlinear mapping, F is the set composed of visual attribute representations, is the correlation analysis between attributes, is the number of emotion categories; as the emotion classification loss function converges, the construction of the attribute-emotion map is gradually completed.

[0031] In the visual emotion analysis task, if both need to perform optimally, it is necessary to balance between the attribute alignment loss and the emotion classification loss. In the embodiments of the present invention, hyperparameters are introduced to dynamically adjust the influence of the optimization between the two losses on the construction of the emotion map; specifically, the total loss calculation is performed on the emotion classification loss and the attribute alignment loss to obtain the target total loss; among them, the total loss calculation for the emotion classification loss and the attribute alignment loss is specifically as follows: ; Among them, is the target total loss, is the hyperparameter, is the attribute alignment loss; then, emotion prediction is performed according to the target attribute-emotion map to obtain an emotion prediction result, where the emotion prediction result includes emotion information and emotion explanation.

[0032] The present invention uses a multi-level attribute expert feature extraction method and introduces a text-guided multi-level alignment module. At the same time, when extracting multiple visual attribute representations, under the guidance of text attribute representations, the extraction of visual attribute representations is optimized. In the emotion reasoning task, attribute correlation analysis is proposed to analyze all attributes and dynamically adjust the weights of each attribute representation, and map the attribute correlation graph to the emotion space for emotion reasoning. By constructing an attribute-emotion graph, the relationship between emotions and attributes can be deeply analyzed, considering not only single representations but also the interaction of multiple attributes, enhancing the interpretability of emotion recognition, and thus improving the accuracy of emotion analysis results.

[0033] Furthermore, as Figure 5 shown, based on the above-mentioned attribute-guided visual emotion analysis method, the present invention also correspondingly provides an attribute-guided visual emotion analysis system, wherein the attribute-guided visual emotion analysis system includes: A representation processing module 51, configured to obtain a to-be-analyzed emotion image of a target user, perform representation extraction on the to-be-analyzed emotion image to obtain multiple visual attribute representations, and perform text representation extraction on the to-be-analyzed emotion image to obtain multiple text attribute representations; An attribute analysis module 52, configured to perform alignment processing on all the visual attribute representations and all the text attribute representations to obtain multiple target visual attribute representations, and perform attribute correlation analysis on all the target visual attribute representations to obtain attribute correlation information; An emotion prediction module 53, configured to construct an attribute-emotion graph according to the attribute correlation information, perform optimization processing on the attribute-emotion graph to obtain a target attribute-emotion graph, and perform emotion prediction according to the target attribute-emotion graph to obtain an emotion prediction result.

[0034] Furthermore, as Figure 6 shown, based on the above-mentioned attribute-guided visual emotion analysis method, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 6 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0035] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In some other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a visual emotion analysis program 40 based on attribute guidance is stored on the memory 20, and the visual emotion analysis program 40 based on attribute guidance can be executed by the processor 10, so as to implement the visual emotion analysis method based on attribute guidance in this application.

[0036] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chips, and is used to run program codes stored in the memory 20 or process data, such as executing the visual emotion analysis method based on attribute guidance, etc.

[0037] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 30 is used to display information on the terminal and to display a visual user interface. Components of the terminal communicate with each other through a system bus.

[0038] In one embodiment, when the processor 10 executes the visual emotion analysis program 40 based on attribute guidance in the memory 20, the following steps are implemented: Obtain the emotion image to be analyzed of the target user, perform feature extraction on the emotion image to be analyzed to obtain a plurality of visual attribute features, and perform text feature extraction on the emotion image to be analyzed to obtain a plurality of text attribute features; Align all the visual attribute features and all the text attribute features to obtain a plurality of target visual attribute features, and perform attribute correlation analysis on all the target visual attribute features to obtain attribute correlation information; Construct an attribute emotion map according to the attribute correlation information, perform optimization processing on the attribute emotion map to obtain a target attribute emotion map, and perform emotion prediction according to the target attribute emotion map to obtain an emotion prediction result.

[0039] Among them, obtaining the emotional image to be analyzed of the target user, extracting representations from the emotional image to be analyzed to obtain multiple visual attribute representations, and extracting text representations from the emotional image to be analyzed to obtain multiple text attribute representations specifically include: Obtain the emotional image to be analyzed of the target user and construct a multi-level attribute expert module; Extract representations from the emotional image to be analyzed through the multi-level attribute expert module to obtain multi-level attribute representations, where the multi-level attribute representations include low-level attribute representations, intermediate-level attribute representations, and high-level attribute representations; Perform a non-linear transformation on the multi-level attribute representations to obtain multiple visual attribute representations; Extract text representations from the emotional image to be analyzed to obtain multiple text attribute representations.

[0040] Among them, the specific operation of performing a non-linear transformation on the multi-level attribute representations is: ; The specific operation of extracting text representations from the emotional image to be analyzed is: ; Among them, is the visual attribute representation set, is the activation function, is the low-level attribute representation, is the intermediate-level attribute representation, is the high-level attribute representation, and are both learnable parameters, is the number of attribute modules, is the text attribute representation set, is the text adapter, is the text encoder, is the attribute embedding code.

[0041] Among them, the specific operation of aligning all the visual attribute representations and all the text attribute representations to obtain multiple target visual attribute representations includes: Perform scale assimilation on all the visual attribute representations and all the text attribute representations, and align the scale-assimilated visual attribute representations and text attribute representations to obtain multiple aligned visual attribute representations; Calculate the attribute alignment loss for all the scale-assimilated visual attribute representations and text attribute representations, and optimize all the aligned visual attribute representations according to the attribute alignment loss to obtain multiple target visual attribute representations.

[0042] Among them, the attribute correlation analysis of all the target visual attribute representations to obtain attribute correlation information specifically includes: Calculating scores for the attribute correlations of all the target visual attribute representations to obtain multiple correlation scores; Performing attribute correlation analysis on all attributes according to all the correlation scores to obtain attribute correlation information.

[0043] Among them, constructing an attribute sentiment graph according to the attribute correlation information, performing optimization processing on the attribute sentiment graph to obtain a target attribute sentiment graph, and performing sentiment prediction according to the target attribute sentiment graph to obtain a sentiment prediction result specifically includes: Performing spatial mapping on the attribute correlation information to obtain an attribute mapping result, and constructing an attribute sentiment graph according to the attribute mapping result; Calculating a sentiment classification loss for all the visual attribute representations, and calculating a total loss for the sentiment classification loss and the attribute alignment loss to obtain a target total loss; Performing optimization processing on the attribute sentiment graph according to the target total loss to obtain a target attribute sentiment graph, and performing sentiment prediction according to the target attribute sentiment graph to obtain a sentiment prediction result, where the sentiment prediction result includes sentiment information and a sentiment explanation.

[0044] Among them, the calculation of the loss for all the visual attribute representations is specifically: ; The calculation of the total loss for the sentiment classification loss and the attribute alignment loss is specifically: ; Among them, is the sentiment classification loss, is the total number of sentiment categories, is the number of attribute modules, is the th true label, is a non-linear mapping, F is a set composed of visual attribute representations, is the correlation analysis between attributes, is the number of sentiment categories, is the target total loss, is a hyperparameter, is the attribute alignment loss.

[0045] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a visual emotion analysis program guided by attributes, and when the visual emotion analysis program guided by attributes is executed by a processor, the steps of the visual emotion analysis method guided by attributes as described above are implemented.

[0046] In summary, the present invention provides a visual emotion analysis method, system, terminal and storage medium guided by attributes. The method includes: obtaining an emotion image to be analyzed of a target user, performing feature extraction on the emotion image to be analyzed to obtain a plurality of visual attribute features, and performing text feature extraction on the emotion image to be analyzed to obtain a plurality of text attribute features; aligning all the visual attribute features and all the text attribute features to obtain a plurality of target visual attribute features, and performing attribute correlation analysis on all the target visual attribute features to obtain attribute correlation information; constructing an attribute emotion map according to the attribute correlation information, performing optimization processing on the attribute emotion map to obtain a target attribute emotion map, and performing emotion prediction according to the target attribute emotion map to obtain an emotion prediction result. The present invention uses a multi-level attribute expert-level feature extraction method, introduces a text-guided multi-level alignment module, and optimizes the extraction of visual attribute features by extracting a plurality of visual attribute features under the guidance of text attribute features; in the emotion reasoning task, attribute correlation analysis is proposed to analyze all attributes and dynamically adjust the weight of each attribute feature, and map the attribute correlation graph to the emotion space for emotion reasoning. By constructing an attribute emotion map, the relationship between emotions and attributes can be deeply analyzed, not only considering single representations, but also integrating the interaction between multiple attributes, enhancing the interpretability of emotion recognition, and thus improving the accuracy of emotion analysis results.

[0047] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including the element.

[0048] Certainly, those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disc, etc.

[0049] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or modifications can be made according to the above description, and all such improvements and modifications shall fall within the protection scope of the appended claims of the present invention.

Claims

1. A visual sentiment analysis method based on attribute guidance, characterized in that, The attribute-guided visual sentiment analysis method includes: Obtain the sentiment image to be analyzed of the target user, perform feature extraction on the sentiment image to be analyzed to obtain multiple visual attribute features, and perform text feature extraction on the sentiment image to be analyzed to obtain multiple text attribute features; Align all the visual attribute features and all the text attribute features to obtain multiple target visual attribute features, and perform attribute correlation analysis on all the target visual attribute features to obtain attribute correlation information; Construct an attribute sentiment graph according to the attribute correlation information, perform optimization processing on the attribute sentiment graph to obtain a target attribute sentiment graph, and perform sentiment prediction according to the target attribute sentiment graph to obtain a sentiment prediction result.

2. The visual emotion analysis method based on attribute guidance according to claim 1, wherein The step of obtaining the sentiment image to be analyzed of the target user, performing feature extraction on the sentiment image to be analyzed to obtain multiple visual attribute features, and performing text feature extraction on the sentiment image to be analyzed to obtain multiple text attribute features specifically includes: Obtain the sentiment image to be analyzed of the target user and construct a multi-level attribute expert module; Perform feature extraction on the sentiment image to be analyzed through the multi-level attribute expert module to obtain multi-level attribute features, where the multi-level attribute features include low-level attribute features, intermediate-level attribute features, and high-level attribute features; Perform non-linear transformation on the multi-level attribute features to obtain multiple visual attribute features; Perform text feature extraction on the sentiment image to be analyzed to obtain multiple text attribute features.

3. The visual emotion analysis method based on attribute guidance according to claim 2, wherein The specific method for performing non-linear transformation on the multi-level attribute features is: ; The specific method for performing text feature extraction on the sentiment image to be analyzed is: ; Among them, is for visual attribute table solicitation, is an activation function, is a low-level attribute representation, is a mid-level attribute representation, is a high-level attribute representation, and are both learnable parameters, is the number of attribute modules, is for text attribute table solicitation, is a text adapter, is a text encoder, is an attribute embedding encoding.

4. The method for visual sentiment analysis based on attribute guidance according to claim 1, wherein The step of aligning all the visual attribute features and all the text attribute features to obtain multiple target visual attribute features specifically includes: Perform scale assimilation on all the visual attribute features and all the text attribute features, and align all the visually and textually scale-assimilated features to obtain multiple aligned visual attribute features; Calculate the attribute alignment loss for all the visually and textually scale-assimilated features, and optimize all the aligned visual attribute features according to the attribute alignment loss to obtain multiple target visual attribute features.

5. The visual emotion analysis method based on attribute guidance according to claim 1, wherein The step of performing attribute correlation analysis on all the target visual attribute features to obtain attribute correlation information specifically includes: Calculate the score of the attribute correlation of all the target visual attribute features to obtain multiple correlation scores; Perform attribute correlation analysis on all the attributes according to all the correlation scores to obtain attribute correlation information.

6. The method for visual emotion analysis based on attribute guidance according to claim 4, wherein The step of constructing an attribute sentiment graph according to the attribute correlation information, performing optimization processing on the attribute sentiment graph to obtain a target attribute sentiment graph, and performing sentiment prediction according to the target attribute sentiment graph to obtain a sentiment prediction result specifically includes: Perform spatial mapping on the attribute correlation information to obtain an attribute mapping result, and construct an attribute sentiment graph according to the attribute mapping result; Calculate the loss for all the visual attribute representations to obtain the sentiment classification loss, and calculate the total loss for the sentiment classification loss and the attribute alignment loss to obtain the target total loss; Optimize the attribute sentiment map according to the target total loss to obtain the target attribute sentiment map, and perform sentiment prediction according to the target attribute sentiment map to obtain the sentiment prediction result, where the sentiment prediction result includes sentiment information and sentiment explanation.

7. The method for attribute-guided visual sentiment analysis according to claim 6, wherein The calculation of the loss for all the visual attribute representations is specifically as follows: ; The calculation of the total loss for the sentiment classification loss and the attribute alignment loss is specifically as follows: ; Among them, is the emotion classification loss, is the total number of emotion categories, is the number of attribute modules, is the th true label, is the non-linear mapping, F is the set composed of visual attribute representations, is the correlation analysis between attributes, is the number of emotion categories, is the total target loss, is the hyperparameter, is the attribute alignment loss.

8. A visual emotion analysis system based on attribute guidance, characterized in that The attribute-guided visual sentiment analysis system includes: A representation processing module, configured to obtain a sentiment image to be analyzed of a target user, perform representation extraction on the sentiment image to be analyzed to obtain a plurality of visual attribute representations, and perform text representation extraction on the sentiment image to be analyzed to obtain a plurality of text attribute representations; An attribute analysis module, configured to perform alignment processing on all the visual attribute representations and all the text attribute representations to obtain a plurality of target visual attribute representations, and perform attribute correlation analysis on all the target visual attribute representations to obtain attribute correlation information; A sentiment prediction module, configured to construct an attribute sentiment map according to the attribute correlation information, optimize the attribute sentiment map to obtain the target attribute sentiment map, and perform sentiment prediction according to the target attribute sentiment map to obtain the sentiment prediction result.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and an attribute-guided visual sentiment analysis program stored on the memory and executable on the processor. When the attribute-guided visual sentiment analysis program is executed by the processor, the steps of the attribute-guided visual sentiment analysis method according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an attribute-guided visual sentiment analysis program. When the attribute-guided visual sentiment analysis program is executed by the processor, the steps of the attribute-guided visual sentiment analysis method according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Multi-modal emotion guiding method and system based on emotion map, and storage medium

    CN112133406A

  • Event description text generation method, device and equipment based on video data

    CN119904786A

  • Text-based sentiment classification method and apparatus, and computer device and storage medium

    WO2023134083A1